AI Training Enters a New Era: Frontline Labs Signal a Shift Toward Recursive Self-Improvement

Stock News
Yesterday

The artificial intelligence sector is witnessing a transformative trend that investors should closely monitor. Over the past two years, market discussions about AI have centered on several key questions: whether model capabilities can continue advancing, if inference costs can drop, whether new trillion-dollar applications will emerge beyond coding, and if data center construction is at risk of overcapacity. However, with Alphabet's release of Gemini 3.8 Flash, SSI's announcement of a 10x compute expansion over the next 12 months, and OpenAI's public discussions on recursive self-improvement (RSI) and training resource allocation, frontier labs are shifting the competitive focus in a new direction: AI is no longer just serving users—it is also beginning to participate in training, evaluating, and optimizing the next generation of AI.

RSI, or Recursive Self-Improvement, is commonly understood not as a model suddenly gaining "self-awareness," but rather as involving models in more stages of the research and development process: generating algorithms, writing code, designing experiments, invoking training tools, evaluating outcomes, fixing errors, and feeding effective insights back into subsequent development cycles. If this loop continues to shorten, the pace of AI research could fundamentally change. RSI is not a bet unique to any single company—Alphabet, OpenAI, SSI, and Anthropic, nearly every leading frontier lab, is moving in the same direction.

Gemini 3.8 Flash: Alphabet's RSI Signal

On September 2, Alphabet officially released Gemini 3.8 Flash, the third Flash model launched within six weeks. The performance metrics alone are striking: in the high-reasoning mode of the independent evaluation firm Artificial Analysis, the 3.8 Flash Intelligence Index scores 59 points, approaching top-tier flagship models like GPT-5.6 Sol and Grok 4.6, with a single-task cost of only $0.58 and an input price of $0.75 per million tokens. Yet what deserves more attention is how Yao Shunyu, a DeepMind researcher at Alphabet, characterized the model: for the model itself, this is only a modest step forward, but for RSI, it represents a tremendous leap.

This statement reveals what Alphabet is truly pursuing. An internal memo at Alphabet explicitly outlines the objective: to force recursive self-evolution, transform the coding model into a fully automated AI researcher, and completely close the entire R&D loop. To achieve this, Alphabet has assembled a code strike team led by the DeepMind CTO with direct oversight from co-founder Sergey Brin, and has recruited Barret Zoph, the former head of post-training at OpenAI, as Vice President of Research, overseeing reinforcement learning and post-training. The core upgrade in 3.8 Flash is the product embodiment of this strategy: the model is designed to execute more reasoning steps on complex tasks, repeatedly invoke tools, and evaluate and optimize its own results through long-running autonomous agent loops. In other words, the model begins to take on the full "plan–execute–check–fix" workflow.

Notably, Alphabet had several candidate 3.5 Pro models internally, but ultimately rejected them all—because they did not significantly outperform the Flash series. Meanwhile, the next-generation flagship Gemini 4 remains stalled in post-training. These circumstances have inadvertently fueled the rapid iteration of the Flash series.

OpenAI's Brake and SSI's Acceleration: Giants in Sync

Around the time Alphabet released its new model, OpenAI CEO Sam Altman made an unusually candid statement in a public interview. He acknowledged that OpenAI had delayed a frontier reinforcement learning (RL) training run—the first time in the company's history that frontline training was voluntarily paused. The reason was not a single "smoking gun" event; rather, during training, the team observed a startlingly rapid improvement in model capabilities. Altman's exact words were: "The rate of capability improvement... I can only describe it as 'awe-inspiring.' We need more time for safety, alignment, and safeguards to catch up." He further added: "A year ago, I didn't think superintelligence would arrive soon. Now I think it could happen. I'm not saying it will, but we are advancing at an extremely fast pace."

Altman also clarified that this pause was specific to the frontier RL training run, not a halt to all training, and that compute clusters were not idle. He emphasized that OpenAI's commercial momentum remains strong, with enterprise revenue now exceeding consumer revenue, and that existing models still hold significant commercial value. However, the signal from this statement goes far beyond business: even OpenAI itself is braking to let safety catch up—the most direct public indication that RSI is approaching.

In addition, Nvidia recently announced a major strategic investment in SSI (Safe Superintelligence Inc.), which simultaneously declared a 10x expansion in compute over the next 12 months. SSI, founded in June 2024 by OpenAI co-founder and former chief scientist Ilya Sutskever, has as its "only goal and only product" the pursuit of safe superintelligence. The most noteworthy line in the announcement is: "Our research is worth scaling."

According to analysis cited by Wall Street Insights' "Minority Views" column, the market currently holds a common narrative: apart from coding, AI has yet to find the next trillion-dollar application, so AI capital expenditure growth will likely peak around 2028. But the SSI event reveals that this narrative may be grasping the wrong variable. The first determinant of CapEx at frontier labs has never been a specific application track; it is whether the next-generation model remains worth expanding training scale. Over the past six months, almost all frontier labs have sent a highly consistent signal: OpenAI continues to build out training clusters, Anthropic persistently raises funds to build AI infrastructure, xAI keeps expanding Colossus, Meta continues constructing gigawatt-scale AI campuses, Alphabet expands its TPU deployment, and SSI announces a tenfold compute expansion. Not a single company's actions suggest that "training is over."

RSI Reshapes Compute Logic: Training Shifts from "One-Time" to "Never-Ending"

The key may lie in how RSI (Recursive Self-Improvement) is transforming the nature of training itself. The old training model was: collect data → train → deploy → finish. The RSI-era training model is: model A generates a new algorithm → trains model B → model B optimizes the training process → trains model C → model C discovers a better RL strategy → and the loop continues indefinitely. This means training becomes Continuous Training—it will never stop. To validate a single new algorithm, where one training run was previously sufficient, labs may now run 100, 1000, or 10,000 versions simultaneously, keeping only the best one.

As analysis cited by Wall Street Insights points out, this leads to a direct inference: compute advantage directly translates into research advantage, and research advantage directly translates into model advantage. A lab with sufficient GPUs can complete all validations in a single day, while a lab lacking GPUs can only complete 5 or 10 per day—the speed gap widens immediately. OpenAI has proposed the concept of an "Automated AI Researcher"; Anthropic heavily involves Claude in model development; Alphabet's AlphaEvolve is already using AI to discover new algorithms. These are all real instances of RSI in action.

Gavin Baker's Warning: Top Labs May Actively Sacrifice Revenue for Training

What does this mean for investors? In a recent public conversation, renowned tech investor Gavin Baker provided a concrete calculation: suppose a lab has 10GW of compute, with 8GW allocated to inference, generating annualized revenue of approximately $480 billion (at $60 billion per GW). If they achieve a major research breakthrough and decide to compress inference allocation from 8GW to 2GW while expanding training from 2GW to 8GW, annualized revenue would plummet from $480 billion to $120 billion. "I genuinely believe they could make such a decision, and this is something the public markets must get used to."

Gavin Baker also stated that, based on a firm belief in Scaling Laws, top large-model companies will not prioritize free cash flow in the short term; instead, they will reinvest all profits and capital into purchasing compute and model training. He further noted that this differs from internet companies like Meta and Alphabet—the latter have relatively stable fundamentals, with no massive trade-off between infrastructure costs and revenue. Frontier labs, by contrast, may at any moment sacrifice substantial short-term commercial revenue to gain a long-term technological edge. This represents a major risk signal for AI stock valuations.

The Frontier Researcher's Mindset: The Powerlessness of Compute Dominance

This race is also altering the psychological landscape of frontline researchers. Sarah Guo, founder of Conviction and an AI investor, shared observations in a video blog interview based on conversations with about 250 frontier AI builders: over the past 12 months, a growing number of top researchers have come to believe that once AI research models with recursive self-improvement capabilities emerge, humanity may be only 1 to 2 years away from exponential intelligence. At the same time, as training budgets move toward hundreds of billions or even trillions of dollars and teams expand to thousands of people, some top researchers are experiencing two negative sentiments: "What I do doesn't matter, because models can do it themselves soon" and "The only thing that matters is compute scale—individual contributions are being diluted." This psychological shift itself is a side confirmation that RSI logic is gaining broad acceptance—when researchers begin to think "only compute matters," it precisely indicates that the link between compute and research capability has become deeply entrenched.

RSI Approaches: Market Narratives Need Updating

A conclusion with direct implications for investors is emerging: the market's narrative of "AI capital expenditure peaking in 2028" rests on the premise that CapEx is driven solely by commercial application demand. But RSI offers a different analytical framework: the future of CapEx is determined not just by inference demand and application revenue, but also by whether model capabilities remain worth expanding training scale. As long as frontier labs continue to believe "Our research is worth scaling," training investment will not cease. Based on all public signals so far, this premise has not yet been shaken.

Whether RSI will become the core variable driving the next round of AI capital expenditure is still too early to conclude. RSI has not yet been proven to reliably deliver exponential capability jumps. Improvements on benchmarks do not necessarily equate to the ability to continuously and autonomously improve models in real-world research environments. Higher agent capabilities also bring higher token consumption, more complex systems engineering, and greater safety risks. Yet judging by the signals from Alphabet, SSI, and OpenAI, frontier labs are clearly no longer treating RSI as a distant theoretical concept. Alphabet is testing low-cost models on longer, more complex, and more verifiable agent tasks; SSI is validating whether "research is worth scaling" through a tenfold compute expansion; and OpenAI is discussing how to reallocate resources among safety, commercialization, and frontier training.

These changes collectively point to one trend: the next phase of competition in the large-model industry may no longer be just "whose model answers questions better," but more importantly, "who can enable AI to participate in AI research faster and iterate within safe boundaries." If this trend continues, the core questions for the AI industry chain will also shift. The market should not only ask: is AI application revenue sufficient to support data center construction? It should also ask: is the capability curve of frontier models still rising? Can AI truly shorten the R&D cycle? Can training clusters convert compute into research outcomes? And can safety systems keep pace with the speed of model self-iteration? This may be the most significant impact RSI has on AI capital expenditure, model competition, and technology investment.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10