Ant Group's Ling 3.0 Flash is live on Vercel's AI Gateway, free through August 3rd. The model is a 124B parameter Mixture-of-Experts architecture with roughly 5.1B parameters active per token, a 256K token context window, and both thinking and non-thinking inference modes.

The design target is agentic workloads at production scale: coding agents, long-context multi-turn sessions, and document processing where token budgets and latency constraints are tight. Access it via the AI SDK using the model ID inclusionai/ling-3.0-flash-free. AI Gateway passes through provider pricing with no markup and no platform fee, including on Bring Your Own Key requests.

The full changelog is worth reading for the specifics on Gateway's routing rules, Zero Data Retention support, per-key budgets, and failover configuration. These are the operational details that determine whether a model is actually usable in production, not just in a playground.

[READ ORIGINAL →]