DeepSeek V4 Flash, currently the most-used model on OpenRouter by weekly token volume, completed only 53.8% of complex agent tasks in structured testing. Composio ran the model through eight harnesses, including Claude Code, Codex, and OpenCode, across 30 multi-step workflows touching Gmail, GitHub, Slack, and Google Sheets. Of 240 total runs, 129 passed. Only 6 of the 30 workflows succeeded across every harness tested.
Simultaneously, DeepSeek is raising API prices by up to 1,100%. Flash moves from its current rate to 44 cents per million input tokens and $1.32 per million output tokens at peak. Pro hits $1.32 input and $3.96 output at peak. Cache hit pricing jumps between 52% and 1,100%. Seventeen of every 24 hours stay at half price, a structure DeepSeek frames as 'workload scheduling' incentive. The company still undercuts OpenAI, Anthropic, and Google on price, but the math now requires tighter justification.
The Composio data is the piece worth reading in full. The same model produced substantially different pass rates depending on harness configuration, caching behavior, retries, and provider stack. That finding reframes the real enterprise question: not whether V4 Flash is capable, but whether the orchestration layer around it is. ML researcher Nathan Lambert called Flash adoption 'insane' and the model itself a 'total monster.' The testing results suggest that verdict depends heavily on what is doing the orchestrating.
[READ ORIGINAL →]