OpenAI's GPT-OSS 120b, released August 2025, still moves 71.3 billion tokens per day on OpenRouter, roughly 36% of the volume commanded by Anthropic's newest frontier model, Claude Opus 4.8. GLM 5.2 leads the entire field at 495 billion daily tokens. The token market has split into tiers, and the splits are widening.
Competition is driving the segmentation fast. Anthropic shipped Opus 5 on July 24, 2026, priced below its predecessor Fable, targeting the same mid-market ground as Moonshot's Kimi 3. Poolside launched Laguna S 2.1 the same week: a 118-billion-parameter mixture-of-experts model that activates only 8 billion parameters per token. On an M5 Max, it runs at the same decode speed as a 26-billion-parameter dense model. The author's own production data, five months of MCP tool-call logs from a local coding and email agent, shows why that matters: tool-call failure rates drop from 29.4% on Ornith-1.0 35B to 20.1% on Laguna S 2.1, a 7-point reduction just from swapping the local model.
Read the original for the failure-rate chart and the architecture explanation that makes the speed result non-obvious. The core argument is structural: frontier models no longer have to serve every use case, local models are now good enough to absorb meaningful workloads, and price competition at the top is accelerating the ceiling for the tier running on your laptop.
[READ ORIGINAL →]