Z.ai revealed on August 26 that the mystery model 'Ox Alpha' on OpenRouter was GLM-5.3-Flash, a model the company had deliberately run on public traffic for a week while the AI community ran tokenizer traces and networking forensics trying to identify it. The model processed several trillion tokens daily during that period. It runs entirely on Chinese chips and infrastructure. Weights are MIT-licensed. Pricing is 7.5 cents per million input tokens and 25 cents per million output tokens through September 9 on OpenRouter, then 15 cents and 50 cents after. Artificial Analysis scores it at 57 on their intelligence index at roughly nine cents per task. GPT-5.6 Sol scores 59 at 67 cents. Grok 4.6 scores 61 at 94 cents. The forensics week that preceded the reveal is worth reading in detail.

The cost pressure is already restructuring enterprise AI spending. Uber's CTO burned $1,200 in a single two-hour demo and exhausted the company's full-year 2026 coding budget by April. Uber capped per-person per-tool spend at $1,500 by June. McKinsey's 2026 State of AI survey found 80 percent of users report faster output but only 37 percent of companies see measurable EBIT impact. On OpenRouter, Chinese models passed US models in token share in early June. The direction is not a prediction. It is already the current state.

The article proposes a three-tier model allocation: top-tier models like Fable and Opus for roughly 5 percent of tasks, specifically irreversible strategic decisions. Mid-tier models including Kimi K3, Gemini 3.7 Flash, GPT-5.6 Sol, and Grok 4.6 for 50 percent. GLM-5.3-Flash as the volume workhorse for 45 percent. The practical framework and the token-economics math behind each split, not just the conclusion, are what make the full piece worth the read. September brings expected releases from Google, xAI, Anthropic, OpenAI, and DeepSeek. The Pareto frontier may shift. The cost direction will not.

[READ ORIGINAL →]