On the record
“I fully agree that the models will be superhuman distributed GPU engineers in a few years.”
“This will make experimentation and tinkering with the formulation of the models far easier, but it will not make our models dramatically different in nature.”
“I expect the effective cost of model intelligence to decline near-exponentially over the coming years (potentially faster than recent trends).”
Opinions and predictions (not established facts)
- AI models will be superhuman distributed GPU engineers in a few years.
- The bottleneck in AI development will be less and less engineering over the coming years.
- The effective cost of model intelligence is expected to decline near-exponentially over the coming years.
- A prediction of pretraining research, at least in architecture and data selection to serve our current class of models, being automated in 2-3 years feels reasonable to me.
- This will be a massive trigger of Jevons paradox for agentic models.
What changed
In a recent piece for Interconnects, AI researcher Nathan Lambert explained his view on the current and future trajectory of AI development. Lambert posits that while AI capabilities are set for rapid acceleration, this progress will primarily stem from improvements in infrastructure and engineering, rather than fundamental breakthroughs leading to general superintelligence. He notes that researchers are observing significant advancements in the engineering capabilities of AI models, predicting that these models will soon surpass human capabilities in tasks like distributed GPU engineering. This will facilitate easier experimentation and model formulation but won't drastically alter the fundamental nature of AI. Lambert highlights that AI is entering an era where good ideas are more valuable than execution, partly due to the availability of coding agents that ease research. He contrasts this with the pre-deep learning era where AI was more research-focused, noting that current top researchers are judged by their ability to implement and scale ideas. The coming years, he predicts, will see a massive acceleration driven by parallelized, AI-assisted language modeling, reducing the engineering bottleneck. This will boost the diffusion of AI technology into the economy, even if economically valuable superhuman traits beyond math and coding are not achieved. Much of this near-term progress, he clarifies, comes from scaling inference-time compute with current tools, not from dramatic step-changes in AI's research abilities.
Lambert also detailed how the AI training and inference stack is highly optimizable. Metrics like tokens per second per GPU (training speed) and tokens per prompt or FLOPs per token (inference efficiency) are improvable. He expects AI agents to optimize this process end-to-end within a few years, pushing inference capabilities close to the maximum compute possible on accelerators like GPUs. He notes that companies have already achieved significant cost savings (10-30%) in serving models, and this entire stack is expected to compound, leading to a near-exponential decline in the effective cost of AI intelligence. This phase of efficiency gains, driven by flexible GPU platforms and co-design of accelerators and models, is expected to last a few years. He predicts that pretraining research, particularly in architecture and data selection for current models, could be automated within 2-3 years. This efficiency surge, he believes, will trigger Jevons paradox for agentic models, leading to increased demand as the industry focuses on better ways to orient and deliver agents. He cites Meta's Muse agent as an early indicator, expecting more such experiences tailored to different audiences and use cases, emphasizing value creation through understanding agent functionality rather than pushing performance frontiers.
Furthermore, Lambert identified improving the general quality of RL environments as another area of industrial-scale, low-hanging fruit. He observes a proliferation of RL data companies achieving significant revenue, yet notes that the average outputs from this sector are of remarkably low quality, with many researchers agreeing that much of the purchased data is subpar, despite leading labs seeing a clear return on investment. He asserts that the crudeness of many RL environments is a fixable issue.
Why it matters for builders
Lambert's analysis is crucial for AI developers and researchers as it provides a realistic outlook on the pace and nature of AI advancement. By distinguishing between engineering acceleration and fundamental breakthroughs, he helps frame expectations and guide research priorities. His emphasis on efficiency gains and agent delivery suggests that practical applications and optimization of existing models will be key areas of focus, rather than solely pursuing AGI. This perspective is valuable for understanding the near-term economic impact and diffusion of AI technologies.
Practical impact
Lambert's argument suggests a near-term future where AI development is characterized by rapid iteration and optimization of existing architectures, driven by engineering and infrastructure improvements. This implies that builders should focus on leveraging these efficiency gains and developing better interfaces and delivery mechanisms for AI agents, rather than expecting a sudden leap to general superintelligence. The predicted automation of pretraining research and the optimization of inference stacks point towards a more accessible and cost-effective AI landscape, potentially accelerating its integration into various economic sectors. The focus on agentic models and RL environments indicates a trend towards more specialized and practical AI applications.
Caveats and source limits
Lambert's piece primarily focuses on his predictions and opinions regarding AI's future development, particularly distinguishing between engineering progress and the path to superintelligence. While he cites metrics like tokens per second per GPU and FLOPs per token as optimizable, these are presented as areas of expected improvement rather than independently verified current facts. His predictions about automation timelines (e.g., pretraining research in 2-3 years) and the impact of Jevons paradox are speculative. The source does not provide specific data on the performance of current RL environments or the exact revenue figures for RL data companies, beyond stating that many have crossed significant revenue milestones.
Sources
Claim check: 6/6 supported claims - 6 evidence links - 90% avg confidence
- AI models will be superhuman distributed GPU engineers in a few years.supported - interconnects.ai
- Good ideas can be much more valuable than good execution in software due to the start of an era where coding agents ease research.supported - interconnects.ai
- The bottleneck in AI development will be less and less engineering over the coming years, leading to massive acceleration.supported - interconnects.ai
- The effective cost of model intelligence is expected to decline near-exponentially over the coming years.supported - interconnects.ai
- Pretraining research, at least in architecture and data selection, could be automated in 2-3 years.supported - interconnects.ai
- The average outputs from the RL data sector are remarkably low-quality, though leading labs see clear return on investment.supported - interconnects.ai
Caveats
- This is a prediction by the author.
- Author's interpretation of the current AI research landscape.
- Author's assessment based on researcher consensus and observed market behavior.
- Single-source caution: verify critical details at the linked source.
Radar score 77/100 - how it was calculated
- Reliability 77: Primary on-record statement
- Freshness 100: Fresh source date
- Novelty 64: New statement by a tracked AI voice
- Technical 56: Structured technical source signals
- Developer 56: Builder relevance source signals
- Ecosystem 70: Tracked AI voice
- Confidence 98: Claims have reliable evidence
Discussion
Loading comments...