Google released two variants of Gemini 3.8 Flash on Wednesday: a standard agentic model and Flash Cyber, a cybersecurity-specialized model already deployed inside Google's own infrastructure. Flash Cyber produced 2.6 times more correct patches in Chrome vulnerabilities than larger commercial models and found a critical foundational vulnerability in under 2 hours, a task Google says typically takes months. It scored 86.2% on the CyberGym benchmark and 47.2% on CWE-Bench, and Wiz reported 7.5 to 9.7 percent higher recall of real-world vulnerabilities at 2.3 to 5.2 times lower cost than leading frontier models.
The standard 3.8 Flash ranks No. 14 in Arena.ai's Agent Arena, above DeepSeek-V4-Pro, and No. 7 in Text Arena, ahead of Claude Opus 5. Its predecessor, 3.7 Flash, sits at No. 32 in Agent Arena. Priced at $0.75 per million input tokens and $3.75 per million output tokens, it carries a 1M-token input window and a 64K-token output limit. This is Google's third Flash release in six weeks. The model outperformed frontier competitors on Vals Finance Agent V2, Harvey's Legal Agent Benchmark, and DeepSWE, and scored 54.9% on Humanity's Last Exam-Verified.
Flash Cyber is restricted to vetted partners through Google's Fairwind Program, covering government authorities and critical infrastructure operators, because it ships with a more permissive set of cybersecurity safeguards. Chrome's engineering director described a 'vulnerability apocalypse' driven by generative AI, with vulnerability reports spiking overnight. One bug Flash Cyber caught had sat undetected in Chromium for 13 years despite review by potentially hundreds of engineers. The full piece details the benchmark methodology, the model's prompt injection robustness claims, and why Google explicitly deprioritized offensive exploitation capabilities in training.
[READ ORIGINAL →]