Google releases faster Gemini 3.6 Flash and cybersecurity models as 3.5 Pro stays delayed

Reviewed byNidhi Govil

33 Sources

Share

Google DeepMind unveiled three new Gemini models including the efficiency-focused Gemini 3.6 Flash, which uses 17% fewer tokens and costs less than its predecessor. The company also introduced Gemini 3.5 Flash Cyber for cybersecurity and Flash-Lite for high-volume tasks. However, the flagship Gemini 3.5 Pro remains delayed despite promises of a June release, while Google has already begun pre-training for Gemini 4.

Google Gemini 3.6 Flash Arrives with Enhanced Efficiency

Google DeepMind released three new Gemini models on Tuesday, marking a strategic shift toward cost-efficient AI models even as its flagship offering remains conspicuously absent

1

. The star of the announcement is Gemini 3.6 Flash, which Google positions as its "workhorse model" designed to deliver improved coding capabilities, knowledge work performance, and multimodal tasks while using up to 17% fewer tokens than its predecessor

2

. This token efficiency translates directly to lower costs for developers, with pricing set at $1.50 per million input tokens and $7.50 per million output tokens, down from $1.50 and $9 respectively for the now-deprecated 3.5 Flash

1

.

Source: Analytics Insight

Source: Analytics Insight

The new Gemini models were built in direct response to user feedback about the 3.5 release, particularly around code generation performance that didn't meet Google's initial promises

1

. In the DeepSWE test for coding, Gemini 3.6 Flash scores 49% compared to 37% for 3.5 Flash, while the OSWorld test for computer use shows improvement to 83% from 78.4%

1

. The model now supports computer use as a standard feature in the Gemini API, enabling more sophisticated agentic systems that can complete tasks more accurately and in fewer steps

3

.

Specialized Models Target Cybersecurity and High-Volume Workloads

Alongside the flagship update, Google introduced Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber, each targeting distinct use cases within Google's AI strategy

2

. Flash-Lite achieves an impressive 350 output tokens per second, making it Google's fastest and most cost-effective model at $0.30 per million input tokens and $2.50 per million output tokens. This positions it for high-volume workloads and smaller tasks within larger AI-agent systems, with Google already deploying it extensively in Google Search for AI Overviews

1

.

Gemini 3.5 Flash Cyber represents Google's first LLM tuned specifically for cybersecurity, designed to detect and patch software vulnerabilities

5

. Google claims this specialized model performs nearly as well at finding and fixing cybersecurity issues as the much larger and more expensive Claude Mythos from Anthropic, while maintaining the efficiency of a Flash model

1

. The model is already finding and fixing bugs in Google's internal codebases for Android, Chrome, and YouTube

3

. However, acknowledging the dual-use nature of vulnerability detection tools, Google is restricting access through a limited pilot program exclusively available to governments and trusted partners via Google DeepMind's CodeMender agent

2

.

Source: VentureBeat

Source: VentureBeat

Gemini 3.5 Pro Delay Raises Questions About Competitive Position

The most notable aspect of Tuesday's announcement is what Google didn't release. Gemini 3.5 Pro, the company's flagship reasoning model promised for June at the I/O conference in May, remains in testing with unnamed partners

1

. Google DeepMind product lead Logan Kilpatrick stated the company hopes to "land soon" but provided no concrete timeline

2

. Bloomberg reported that Google is facing internal delays as it struggles to meet performance goals, particularly in coding where the model reportedly falls short compared to offerings from OpenAI and Anthropic

4

.

Source: 9to5Google

Source: 9to5Google

This delay comes at a critical moment as competitors accelerate their release cycles. Since Google last updated its Pro model in February, OpenAI has released GPT-5.5 and begun rolling out GPT-5.6, while Anthropic has launched Claude Opus 4.8, Claude Sonnet 5, and expanded access to its frontier Fable 5 model

2

. Chinese rivals are also gaining momentum, with Moonshot AI's Kimi K3 drawing enough demand to limit new subscriptions due to capacity constraints, and Alibaba teasing Qwen 3.8 Max

5

.

Looking Ahead: Gemini 4 and Monthly Release Cadence

During Alphabet's quarterly earnings call, CEO Sundar Pichai addressed concerns about the Gemini 3.5 Pro delay by pivoting to future plans, announcing that Google has already begun its most ambitious pre-training run yet for Gemini 4

4

. Pichai also outlined plans to release subsequent AI models at an almost monthly cadence, suggesting a fundamental shift in Google's development and deployment strategy

4

.

Google's emphasis on price and efficiency may help offset slower timing in key product categories. Artificial Analysis data shows Gemini Flash already undercuts comparable models from Anthropic, OpenAI, and Chinese rivals on cost

5

. The company is also reportedly developing a specialized chip designed to run Google Gemini up to 10 times more efficiently, part of a broader push to lower serving costs through its custom chips, cloud infrastructure, and integrated hardware-software design

5

. All new models are available starting today for developers through the Gemini API via Google AI Studio and Android Studio, with Gemini 3.6 Flash also available in the Gemini app

3

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved