Google Unleashes New Gemini AI Models, Eyes Enterprise Savings
Google DeepMind has launched Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, new AI models designed to optimize enterprise AI agents by cutting latency and token costs. These "workhorse" models offer specialized capabilities for coding, high-volume tasks, and cybersecurity vulnerability remediation, aiming for efficiency and reliability in AI deployments.
Google DeepMind has unveiled new additions to its Gemini model family: Gemini 3.6 Flash, 3.5 Flash-Lite, and a restricted 3.5 Flash Cyber variant. These models are designed as "workhorses" to enhance the economics of running autonomous software agents in production environments, primarily by reducing latency and token costs. The focus of these releases is to deliver efficiency, latency, and reliability for customers building AI agents at scale, prioritizing throughput over parameter count, especially for background agents rather than chat interfaces.
Gemini 3.6 Flash, positioned for coding, knowledge work, and multimodal reasoning, is notable for its efficiency. Google's developer documentation highlights a 17 percent reduction in output tokens compared to its predecessor, 3.5 Flash, based on the Artificial Analysis Index, with specific synthetic tests like the Datacurve DeepSWE benchmark showing drops in token usage of up to 65 percent. Priced at $1.50 per 1 million input tokens and $7.50 per 1 million output tokens, it's optimized for continuous reasoning loops. Performance metrics show significant improvements: a 49 percent success rate on DeepSWE (up from 37 percent), a score increase from 49.7 percent to 63.9 percent on MLE Bench, and a GDPval-AA v2 score of 1421 (compared to 1349 for the older model), indicating better real-world knowledge work capabilities. Industry leaders like Figma have integrated 3.6 Flash into their prototyping infrastructure for faster design iterations, while legal technology platform Harvey and research tool Hebbia leverage it for multimodal document processing, including ingesting raw financial filings, parsing document structure, reading embedded charts, and generating draft reports for review.
For high-volume, low-latency tasks like document processing and agentic search, Google introduced Gemini 3.5 Flash-Lite. This model is the most cost-effective in its class, offering impressive speed at 350 output tokens per second, making it the fastest in the 3.5 series according to Google. Its pricing is set at $0.3 per 1 million input tokens and $2.5 per 1 million output tokens, allowing engineering teams to route simple, high-volume subagent requests to a minimal thinking level while reserving higher thinking levels for multi-step work. Gemini 3.5 Flash-Lite also shows enhanced performance, achieving a 72.2 percent success rate on Google’s GDM-MRCR v2 long-context test (up from 60.1 percent) and nearly doubling its GDPval-AA v2 score from 642 to 1140. This model is rolling out in the Gemini app and Google Search.
The specialized Gemini 3.5 Flash Cyber is designed to address the growing challenge of patching code vulnerabilities, as automated scanners often surface flaws faster than security teams can remediate them. Built to validate and remediate security flaws, its performance on the CyberGym benchmark is competitive with frontier models, though specific figures haven't been made public in detail. However, its distribution is restricted to governments and vetted partners through a pilot program, a limitation Google frames as a safeguard against the model generating exploit code for offensive use. Multiple instances of 3.5 Flash Cyber operate in parallel within Google’s CodeMender security agent, cross-checking findings before generating a single remediation report for human review.
Both Gemini 3.6 Flash and 3.5 Flash-Lite carry the same native client-side computer-use tool, which has been directly folded into the Gemini API and Gemini Enterprise platforms, removing the need for custom intermediary software. The company reports an OSWorld-Verified score of 83.0 percent, up from 78.4 percent, and says updated safeguards against chemical, biological, radiological, and nuclear misuse improve resistance to jailbreaking without raising refusal rates for benign requests.
While these releases focus on efficiency and specialized tasks, the launch is notable for what it didn’t include: the long-anticipated update to Google's flagship model, Gemini Pro, which was last updated in February. Despite teasing a release in May and reports of internal delays in meeting performance goals, Google DeepMind product lead Logan Kilpatrick confirmed that Gemini 3.5 Pro is currently undergoing partner testing with an anticipated release "soon." Meanwhile, pre-training for the next-generation Gemini 4 architecture is already actively underway, highlighting Google's ongoing commitment to advancing its AI capabilities.