Why it matters
Lambert's perspective is significant as it offers a grounded view on AI's trajectory, contrasting with more speculative 'takeoff' scenarios. His insights are particularly relevant for developers navigating the current AI landscape and anticipating future directions.

On the record

Essay · Interconnects · Read the original

“I fully agree that the models will be superhuman distributed GPU engineers in a few years.”
Nathan Lambert · in the original · checked against the source text
“This will make experimentation and tinkering with the formulation of the models far easier, but it will not make our models dramatically different in nature.”
Nathan Lambert · in the original · checked against the source text
“I expect the effective cost of model intelligence to decline near-exponentially over the coming years (potentially faster than recent trends).”
Nathan Lambert · in the original · checked against the source text

Opinions and predictions (not established facts)

  • AI models will be superhuman distributed GPU engineers in a few years.
  • The bottleneck in AI development will be less and less engineering over the coming years.
  • The effective cost of model intelligence is expected to decline near-exponentially over the coming years.
  • A prediction of pretraining research, at least in architecture and data selection to serve our current class of models, being automated in 2-3 years feels reasonable to me.
  • This will be a massive trigger of Jevons paradox for agentic models.

What changed

In a recent piece for Interconnects, AI researcher Nathan Lambert explained his view on the current and future trajectory of AI development. Lambert posits that while AI capabilities are set for rapid acceleration, this progress will primarily stem from improvements in infrastructure and engineering, rather than fundamental breakthroughs leading to general superintelligence. He notes that researchers are observing significant advancements in the engineering capabilities of AI models, predicting that these models will soon surpass human capabilities in tasks like distributed GPU engineering. This will facilitate easier experimentation and model formulation but won't drastically alter the fundamental nature of AI. Lambert highlights that AI is entering an era where good ideas are more valuable than execution, partly due to the availability of coding agents that ease research. He contrasts this with the pre-deep learning era where AI was more research-focused, noting that current top researchers are judged by their ability to implement and scale ideas. The coming years, he predicts, will see a massive acceleration driven by parallelized, AI-assisted language modeling, reducing the engineering bottleneck. This will boost the diffusion of AI technology into the economy, even if economically valuable superhuman traits beyond math and coding are not achieved. Much of this near-term progress, he clarifies, comes from scaling inference-time compute with current tools, not from dramatic step-changes in AI's research abilities.

Lambert also detailed how the AI training and inference stack is highly optimizable. Metrics like tokens per second per GPU (training speed) and tokens per prompt or FLOPs per token (inference efficiency) are improvable. He expects AI agents to optimize this process end-to-end within a few years, pushing inference capabilities close to the maximum compute possible on accelerators like GPUs. He notes that companies have already achieved significant cost savings (10-30%) in serving models, and this entire stack is expected to compound, leading to a near-exponential decline in the effective cost of AI intelligence. This phase of efficiency gains, driven by flexible GPU platforms and co-design of accelerators and models, is expected to last a few years. He predicts that pretraining research, particularly in architecture and data selection for current models, could be automated within 2-3 years. This efficiency surge, he believes, will trigger Jevons paradox for agentic models, leading to increased demand as the industry focuses on better ways to orient and deliver agents. He cites Meta's Muse agent as an early indicator, expecting more such experiences tailored to different audiences and use cases, emphasizing value creation through understanding agent functionality rather than pushing performance frontiers.

Furthermore, Lambert identified improving the general quality of RL environments as another area of industrial-scale, low-hanging fruit. He observes a proliferation of RL data companies achieving significant revenue, yet notes that the average outputs from this sector are of remarkably low quality, with many researchers agreeing that much of the purchased data is subpar, despite leading labs seeing a clear return on investment. He asserts that the crudeness of many RL environments is a fixable issue.

Why it matters for builders

Lambert's analysis is crucial for AI developers and researchers as it provides a realistic outlook on the pace and nature of AI advancement. By distinguishing between engineering acceleration and fundamental breakthroughs, he helps frame expectations and guide research priorities. His emphasis on efficiency gains and agent delivery suggests that practical applications and optimization of existing models will be key areas of focus, rather than solely pursuing AGI. This perspective is valuable for understanding the near-term economic impact and diffusion of AI technologies.

Practical impact

Lambert's argument suggests a near-term future where AI development is characterized by rapid iteration and optimization of existing architectures, driven by engineering and infrastructure improvements. This implies that builders should focus on leveraging these efficiency gains and developing better interfaces and delivery mechanisms for AI agents, rather than expecting a sudden leap to general superintelligence. The predicted automation of pretraining research and the optimization of inference stacks point towards a more accessible and cost-effective AI landscape, potentially accelerating its integration into various economic sectors. The focus on agentic models and RL environments indicates a trend towards more specialized and practical AI applications.

Caveats and source limits

Lambert's piece primarily focuses on his predictions and opinions regarding AI's future development, particularly distinguishing between engineering progress and the path to superintelligence. While he cites metrics like tokens per second per GPU and FLOPs per token as optimizable, these are presented as areas of expected improvement rather than independently verified current facts. His predictions about automation timelines (e.g., pretraining research in 2-3 years) and the impact of Jevons paradox are speculative. The source does not provide specific data on the performance of current RL environments or the exact revenue figures for RL data companies, beyond stating that many have crossed significant revenue milestones.

Sources

Written with AI assistance from the linked sources; every claim below was checked against them automatically. How we produce articles.

Claim check: 6/6 supported claims - 6 evidence links - 90% avg confidence
  • AI models will be superhuman distributed GPU engineers in a few years.supported - interconnects.ai
  • Good ideas can be much more valuable than good execution in software due to the start of an era where coding agents ease research.supported - interconnects.ai
  • The bottleneck in AI development will be less and less engineering over the coming years, leading to massive acceleration.supported - interconnects.ai
  • The effective cost of model intelligence is expected to decline near-exponentially over the coming years.supported - interconnects.ai
  • Pretraining research, at least in architecture and data selection, could be automated in 2-3 years.supported - interconnects.ai
  • The average outputs from the RL data sector are remarkably low-quality, though leading labs see clear return on investment.supported - interconnects.ai

Caveats

  • This is a prediction by the author.
  • Author's interpretation of the current AI research landscape.
  • Author's assessment based on researcher consensus and observed market behavior.
  • Single-source caution: verify critical details at the linked source.
Radar score 77/100 - how it was calculated
Reliability77
Freshness100
Novelty64
Technical56
Developer56
Ecosystem70
Confidence98
  • Reliability 77: Primary on-record statement
  • Freshness 100: Fresh source date
  • Novelty 64: New statement by a tracked AI voice
  • Technical 56: Structured technical source signals
  • Developer 56: Builder relevance source signals
  • Ecosystem 70: Tracked AI voice
  • Confidence 98: Claims have reliable evidence
Share
XLinkedInHacker News

Discussion

Loading comments...

Related articles

Voices & Interviews - Oct 9, 2026Ben Recht Argues Uncertainty Quantification is Narrowly DefinedBen Recht argues that uncertainty quantification in forecasting is too narrowly defined, primarily relying on error bars and prediction intervals. He suggests a need for more imaginative approaches to better prepare for forecast failures.Voices & Interviews - Oct 6, 2026Gary Marcus Argues for AI Regulation at NYC Council HearingGary Marcus, a scientist and author, testified before the New York City Council, advocating for a regulatory regime for AI similar to the FDA for drugs. He argued that AI developers should be required to demonstrate that their products' benefits outweigh the risks before gaining market access.Voices & Interviews - Oct 4, 2026Terence Tao Explains AIM: An Invitation to Explore Mathematics TogetherMathematician Terence Tao, in a blog post, introduces AIM, a community-led initiative for mathematical exploration, aiming to leverage AI while prioritizing human collaboration and understanding. The project seeks to foster a collaborative environment for mathematicians, especially students and early-career researchers, to engage with open problems.Research Papers - Jun 25, 2026Progress Advantage: Annotation-Free Step-Level Scoring for LLM AgentsResearchers have introduced 'progress advantage,' a method that leverages reinforcement learning (RL) post-training to provide step-level evaluation for LLM agents without requiring dedicated reward model training or human annotations. This technique offers a byproduct of standard RL pipelines, enabling fine-grained scoring for applications like test-time scaling, uncertainty quantification, and failure attribution.AI Tools - Sep 22, 2026Higgsfield AI Accelerates Video Tool Release with GPT-6 AstraHiggsfield AI has leveraged GPT-6 Astra to streamline video ad creation, enabling small businesses to bring new creative tools to market more rapidly. This integration focuses on simplifying the production process for video advertisements.