AI will automate GPU engineering and inference optimization within a few years. That is the actual near-term story, not general superintelligence. The author, writing at Interconnects, argues that top researchers conflate two separate things: the coming acceleration in infrastructure and tooling, which is real and measurable, and a fundamental change in what models can do, which is not arriving on the same timeline. Inference efficiency gains are already compounding at 10-30% cost reductions per model release cycle. The author expects effective model intelligence costs to decline near-exponentially, potentially faster than recent trends.

The bottleneck is shifting from engineering back to ideas, a transition the field has seen before. Pre-deep learning AI was research-heavy. The deep learning era made implementation and scaling the dominant skill. Coding agents are now collapsing that execution cost again. The author predicts pretraining research, specifically architecture search and data selection, will be largely automated in 2 to 3 years. This is not RSI or takeoff. The author calls it parallelized, AI-assisted language modeling. Separately, RL training environments, a sector with multiple companies already crossing $100M to $1B in revenue, produce outputs researchers widely describe as low quality. That gap is a fixable inefficiency, and fixing it is a concrete near-term lever.

Read this piece for the argument it does not make. The author explicitly decouples infrastructure progress from claims about economically valuable superhuman reasoning outside math and coding. The section on Jevons paradox applied to agentic demand, and the framing of Meta's Muse as an early signal of value created by delivery rather than frontier performance, are worth the full read. So is the scientific literature argument: that AI's real near-term contribution to biology and chemistry may be cross-subdiscipline pattern matching across sparse research networks, not cure-level breakthroughs.

[READ ORIGINAL →]