33 Sources
[1]
Google reveals faster and cheaper Gemini 3.6 Flash, says 3.5 Pro is still in testing
Google announced a significant evolution of its AI models at I/O in May with the release of Gemini 3.5 Flash, and it's not slowing down. The company has revealed three new AI models today, including its first version of Gemini geared toward cybersecurity. However, none of the new models is the delayed Gemini 3.5 Pro, which was supposed to launch in June. Gemini 3.5 Flash, which was the star of the show at I/O, has already been deprecated. In its place, developers and users will find Gemini 3.6 Flash. Google makes the usual claims about this model -- it's marginally more capable and better at coding, and it has great multimodal features. Google says the changes to 3.6 Flash were made in response to user feedback on the 3.5 release. In general, Gemini 3.5 Flash didn't appear to live up to Google's promises around code generation. Perhaps that's simply a consequence of Google's intense focus on efficiency as businesses have started to fret over the cost of AI tokens. In the DeepSWE test for coding, 3.6 Flash jumps to 49 percent versus 37 percent for 3.5 Flash. The new model now supports computer use as a standard feature in the Gemini API, too. The OSWorld test for computer use shows a modest boost to 83 percent from 3.5's 78.4 percent score. Efficiency was a big focus for Gemini 3.5 Flash, and Google says that effort has been amped up with 3.6. Even with small benchmark gains, Gemini 3.6 Flash uses about 17 percent fewer tokens. In agentic workflows (like the one below), Gemini 3.6 Flash should complete tasks more accurately, in fewer steps, and with fewer tokens. That could save developers (and Google) a lot of money. The new model has a lower API cost, at $1.50/1M input tokens and $7.50/1M output tokens. It was $1.50 and $9, respectively, for 3.5 Flash. Google is not done with the 3.5 branch yet, though. It has also released Gemini 3.5 Flash Lite and 3.5 Flash Cyber. The new Flash Lite is Google's most efficient modern AI, hitting an impressive 350 tokens per second. The company claims this model is ideal for scaling agentic systems without breaking the bank. Based on benchmark numbers, the new Flash Lite is almost on par with frontier models from about a year ago, but it's cheap. Pricing is set at $0.30/1M input tokens and $2.50/1M output tokens, though that is slightly higher than the previous 3.1 Flash Lite ($0.25 and $1.50). Gemini 3.6 Flash will begin rolling out in the API today, and it will take over from 3.5 Flash in the Gemini app. Likewise, Gemini 3.5 Flash Lite is available to developers and in the Gemini app. Google also notes that you'll see a lot of 3.5 Flash Lite in Google search, where its higher speed probably makes it ideal for AI Overviews. So those may get a smidge better. Google's AI roadmap Google also has some news on upcoming AI models. First up will be a limited release of Gemini 3.5 Flash Cyber. This is Google's first LLM tuned specifically for cybersecurity. The company says this model is almost as good at finding and fixing cybersecurity issues as the much larger and more expensive Claude Mythos. At the same time, it has the efficiency of a Flash model. Of course, Google acknowledges the "dual-use" nature of such models, which can just as easily be used to identify vulnerabilities for malicious purposes. The company borrows a page from Anthropic here, painting Gemini 3.5 Flash Cyber as too dangerous to release publicly. Instead, the model will launch soon as a limited pilot in Google DeepMind's CodeMender agent, which is available exclusively to trusted partners and governments. Then there's the curious case of Gemini 3.5 Pro, which Google claimed was slated for a June release back at I/O. That never happened, and Google hasn't had anything to say about it until now. There's not much of an update, though. The company claims its new flagship model, which is supposed to rival GPT 5.6 and Claude Fable/Sonnet 5, is currently in testing with unnamed partners. The model will be released "as soon as it's ready." Earlier reports claimed that Google delayed 3.5 Pro because it couldn't match competing models in coding. Gemini 3.5 Flash may not be top-of-the-line for long when it arrives. Google also notes that it has started pre-training for Gemini 4, a process that is apparently more ambitious than its previous AI efforts. There's no timeline for when we'll see Gemini 4, and we don't know if there will be more 3.x releases before that.
[2]
Google releases three new Gemini models -- but no 3.5 Pro
On Tuesday, Google DeepMind released Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. Gemini 3.6 Flash is Google's "workhorse model" that promises improved capabilities in coding, knowledge work, and multimodal performance while reducing token usage by up to 17%, making it cheaper than its predecessor 3.5 Flash. Gemini 3.5 Flash-Lite is the most cost-effective model in the class, and 3.5 Flash Cyber is a specialized model that was fine-tuned for finding and fixing cybersecurity vulnerabilities at a decent price point. This model will be exclusively available to governments and trusted partners as part of a limited access pilot program, according to Google. Google says the focus on these releases is to deliver efficiency, latency, and reliability to customers that are building AI agents at scale. The launch is notable not just for what Google shipped -- cheaper, faster models optimized for coding, efficiency, and cybersecurity -- but for what it didn't. The update doesn't include the long-anticipated update to Google's flagship model, Gemini Pro, which was last updated in February. In the time since that launch, OpenAI has released GPT-5.5 and begun rolling out GPT-5.6, while Anthropic has launched Claude Opus 4.8, Claude Sonnet 5, and expanded access to its frontier Fable 5 model, highlighting the intense release pace of the rival labs. Google teased the release of Pro as part of the 3.5 Flash release in May, saying the Pro version was "already being used internally, and we look forward to rolling it out next month." Last week, Bloomberg reported that Google was facing internal delays in launching the 3.5 Pro as it struggled to meet internal performance goals. Gemini Pro models are generally Google's highest-capability offerings for complex reasoning and coding tasks, while Flash models prioritize lower cost and faster response times for production applications. Google DeepMind product lead Logan Kilpatrick said Tuesday that the company is currently testing Gemini 3.5 Pro with partners and hopes to "land soon." He also noted that the team has started its most ambitious pre-training run yet for Gemini 4.
[3]
Google Releases 3 New Gemini Models, 3.5 Pro Still Not Available - CNET
Blake has over a decade of experience writing for the web, with a focus on mobile phones, where he covered the smartphone boom of the 2010s and the broader tech scene. When he's not in front of a keyboard, you'll most likely find him playing video games or watching horror movies. Google released three new AI models on Tuesday, all built on Gemini 3.5 Flash. The new models are more token-efficient, faster and more reliable across the board. The tech giant also provided an update on the much-anticipated Gemini 3.5 Pro and what's to come after. Here's what's new in the latest Gemini models from today's announcement. Gemini 3.6 Flash Google called 3.6 Flash its "workhorse" model that's now better at coding, knowledge work and multimodal performance. It also promises reduced token usage by up to 17%, and at a lower cost per token versus its predecessor, 3.5. Google says it built the model based on both developer and customer feedback. A series of benchmarks shows 3.6 Flash's gains in performance and average tokens per task compared to its predecessor. Gemini 3.5 Flash-Lite Google's fastest and most cost-effective model can deliver 350 output tokens per second and "significantly" outperforms previous generations when it comes to agentic workflows, according to the blog post. Like 3.6 Flash, this model now supports computer use as a built-in tool to take on more agentic tasks. Gemini 3.5 Flash Cyber 3.5 Flash Cyber is a specialty model that prioritizes cybersecurity workflows in order to find and fix vulnerabilities. It works alongside an infrastructure agent called CodeMender to help cybersecurity teams quickly identify and patch issues. According to a separate article from Google DeepMind, the new model is already finding and fixing bugs in Google's internal codebases in Android, Chrome and YouTube. This model will initially be limited to governments and trusted partners, but access will expand in the future. Gemini 3.5 Pro is still on the way While the three latest models are the primary focus for Tuesday's announcements, Google gave a brief update to its upcoming flagship AI model, Gemini 3.5 Pro. The model is said to currently be in testing with partners, and it plans to make it available as soon as it's ready. How long that will take is anyone's guess, but Google's also already looking ahead to the next generation of AI, too. Google says it has already begun pretraining for Gemini 4, which will be released at an undetermined date. Both Gemini 3.6 Flash and 3.5 Flash-Lite are available starting today for developers in the Gemini API via Google AI Studio and Android Studio and the Gemini app. 3.6 Flash is also available in Google Antigravity, and 3.5 Flash-Lite is rolling out to Google Search.
[4]
Gemini 3.5 Pro is late but Gemini 4 will be great, says Google CEO
A more frequent model release cycle may bring faster improvements to capabilities and pricing, but CIOs will pay the price in testing, governance and validation, analysts say. Google CEO Sundar Pichai has sought to allay concerns over the delayed release of the Gemini 3.5 Pro large language model. He dodged questions about it in Google's quarterly earnings call on Wednesday by focusing on the company's next frontier AI model, Gemini 4, and plans to release subsequent LLMs at an almost monthly cadence. His comments came a day after Google unveiled Gemini 3.6 Flash and 3.5 Flash Cyber but offered no update on the release of Gemini 3.5 Pro, the company's delayed flagship reasoning model that many developers had expected to arrive weeks earlier. Google introduced the Gemini 3.5 family at its annual I/O conference, promising to release the Pro model in June. That timeline has since slipped, with Bloomberg suggesting Gemini 3.5 Pro is months late because the model's coding performance is falling short of internal expectations, especially when compared to better performance by similar models from OpenAI and Anthropic.
[5]
Google expands Gemini lineup with cheaper models and new Mythos rival
The launch comes one day before Alphabet earnings as Chinese rivals gain ground and Google works to improve the timing, scale and efficiency of its AI releases. Alphabet is releasing three new Gemini models on Tuesday, including its clearest answer yet to Anthropic's lead in cybersecurity, as the company looks to show progress across a product pipeline that has faced delays and mounting competition. Gemini 3.5 Flash Cyber is designed to detect and patch software vulnerabilities and will initially be available only to governments and trusted partners through a limited-access pilot. Google said the specialized model runs at a lower price per token than larger models. That could help Google narrow its cybersecurity gap with Anthropic, which has built an early lead in automated code defense. Google is also launching Gemini 3.6 Flash, which improves coding, multimodal and knowledge-work performance while using up to 17% fewer tokens and costing less per token than the previous model -- a meaningful reduction in the cost of running high-volume workloads. Gemini 3.5 Flash-Lite, meanwhile, is Google's fastest and least expensive model in the 3.5 family, built for high-volume workloads and smaller tasks within larger AI-agent systems. The broader lineup reflects Google's bet that price and efficiency can help offset its slower timing in several key product categories. Artificial Analysis data shows Gemini Flash already undercuts comparable models from Anthropic, OpenAI and Chinese rivals on cost. According to the company, Gemini 3.6 Flash -- the stronger of the two new models -- is cheaper per task than GPT-5.6 Terra Max, Kimi K3 and Qwen 3.7 Max, while 3.5 Flash-Lite costs just a fraction of that. The rollout comes on the eve of Alphabet earnings and as Chinese rivals gain momentum. Moonshot AI's Kimi K3 drew enough demand that the company limited new subscriptions and API access because of capacity constraints, while Alibaba is teasing Qwen 3.8 Max, which it said trails only Anthropic's Fable 5 in overall performance. That demand highlights the other side of the AI race: Building a competitive model is only part of the challenge. Companies also need enough computing capacity to serve it at scale. Google has a potential advantage through its custom chips, cloud infrastructure and ability to design models and hardware together, although the company has faced capacity constraints of its own. Tuesday's model launches come as Google is reportedly developing a specialized chip designed to run Gemini up to 10 times more efficiently, part of a broader push to lower the cost of serving AI. A Google Cloud spokesperson told CNBC in a statement that its teams are "constantly researching and experimenting with new innovations to deliver maximum performance and efficiency for our users and customers" and that "while not every project moves into production, this rigorous exploration is central to our full stack approach." "By co-designing our hardware and software from the ground up, we ensure our systems are integrated and highly optimized for real-world workloads," continued the statement. Google is also offering more visibility into its roadmap after questions about delays. Gemini 3.5 Pro is now being tested with partners ahead of broader availability, while the company has begun its largest-ever pre-training run for Gemini 4. Choose CNBC as your preferred source on Google and never miss a moment from the most trusted name in business news.
[6]
Google Releases Three New Gemini A.I. Models
The models include one that is the company's most powerful and another that is fine-tuned for cybersecurity, as Google competes with rivals like OpenAI and Anthropic. Google on Tuesday released three new artificial intelligence models, including Gemini 3.6 Flash, its most powerful, and Gemini 3.5 Flash Cyber, which is fine-tuned for cybersecurity. Google is aiming to compete with rivals such as Anthropic and OpenAI on cutting-edge models and on cybersecurity, an area in which the two start-ups have forged ahead this year with powerful systems that can identify security vulnerabilities in software. Gemini Flash 3.6 improves on the capabilities of Google's previous version of Flash, which performed strongly on benchmarks like coding and finance tasks, according to A.I. leaderboards. The new model is also cheaper for users. Gemini 3.5 Flash Cyber can find and patch security vulnerabilities at a lower cost than other larger models, the company said. The third model, Gemini 3.5 Flash-Lite, is designed for tasks like managing A.I. "agents," which are bots that can act autonomously to function like digital personal assistants. The new Flash models were built "to meet the sweet spot of efficiency and quality," Tulsee Doshi, a senior director of product management on Google's Gemini team, said in a statement. Google has also been working on a flagship A.I. model called Gemini 3.5 Pro, which is expected to be its highest performing system. The Silicon Valley company had said in May that Pro would be rolled out in June, but it remains unavailable. Google said Pro was still in testing and would be made "broadly available as soon as it's ready." Since OpenAI's ChatGPT chatbot kicked off the A.I. boom in 2022, Google has poured money and effort into leading the race. Its Gemini A.I. models soon ranked among the top A.I. tools. But in recent months, Google has been upstaged by advanced new models from Anthropic and OpenAI, as well as new models from Chinese start-ups such as Moonshot AI. This month, Meta, the owner of Facebook and Instagram, also released a model that crept ahead of Gemini's capabilities, according to several leaderboards. Cybersecurity has increasingly been a focus for A.I. models. The systems can find flaws in software and fix them more rapidly than humans, but they can also be used to exploit vulnerabilities and carry out cyberattacks. In April, Anthropic released a cybersecurity-focused model called Mythos, which it said could create a cybersecurity "reckoning." To prevent the powerful model from falling into the wrong hands, the company made it only available to a small group of organizations so they could defend against cyberattacks. Anthropic later released Fable, a model with similar capabilities and guardrails meant to prevent it from being used for hacking. OpenAI soon introduced its own cybersecurity model and made it available to only a limited group of organizations to prepare their defenses, before rolling it out more broadly. (The New York Times has sued OpenAI and Microsoft, claiming copyright infringement of news content related to A.I. systems. The two companies have denied those claims.) Google said it was taking a similar approach with Gemini 3.5 Flash Cyber. In testing, the model found 55 problems in a complex piece of open-source code known as V8 JavaScript Engine, including 10 vulnerabilities that no other model had uncovered before, the company said. "The model will be exclusively available to governments and trusted partners," Ms. Doshi said. "This will give frontline defenders a head start in finding and fixing critical vulnerabilities before they can be exploited, while mitigating against broader misuse." Google said the new Flash models would be available for a lower cost than its previous offerings. The systems use 17 percent fewer tokens, which are the units of measurement for A.I. use, than earlier models. That means developers will spend less money to complete the same tasks.
[7]
Google launches Gemini 3.6 Flash and a Mythos rival
Google just flooded the market with three cheap, fast Gemini models, including a security model aimed straight at Anthropic. But it held back the one everyone actually wanted, its delayed flagship, and teased a Gemini 4 that so far exists only as a training run. Google's answer to a summer of being outrun is not a bigger model. It is cheaper ones. On Tuesday the company launched three new Gemini models, all at its fast, low-cost "Flash" tier, a day before Alphabet reports earnings. The message is efficiency over raw power. The catch is what is still missing. Cheap, fast, and everywhere but the top The workhorse is Gemini 3.6 Flash. Google says it does better coding and knowledge work than its predecessor while using about 17% fewer output tokens, and it costs less. That is $7.50 per million output tokens, down from $9. Its knowledge now runs to March 2026. Alongside it sits 3.5 Flash-Lite, the fastest of the family at 350 tokens a second, and cheaper still. Both lean into one bet: that most AI work does not need a frontier brain, just a quick, affordable one. That bet has a champion. "Companies are already blowing through their annual token budgets, and it's only May," chief executive Sundar Pichai said earlier this year. A mix of Flash models, he told Business Insider, could save firms over $1 billion a year. A cheaper shot at Mythos The most pointed release is Gemini 3.5 Flash Cyber. It is tuned to find and patch software vulnerabilities, running inside Google's CodeMender agent. Google calls it a "cost-efficient" alternative to large security models. The unnamed target is Anthropic's Mythos, which costs $10 per million input tokens and $50 per million output, as The Verge notes. Google claims Flash Cyber matches frontier performance on a key benchmark at a fraction of the price, a direct swipe at Anthropic's lead in AI-driven security. Because a bug-finder also helps attackers, Google is keeping it on a leash. Flash Cyber goes only to governments and trusted partners, in a limited pilot. The model they did not ship Then there is the absence. Gemini 3.5 Pro, the flagship Google promised for June, is still in testing, reportedly held back after falling short on coding. Google has no model in the public top ten. The timing stings. In roughly a week, xAI's Grok 4.5, three versions of OpenAI's GPT-5.6, and Moonshot's Kimi K3 all shipped. Anthropic's Fable 5, meanwhile, sits atop the leaderboards, as Reuters reported. Gemini 4, on paper Google's reply is to point further ahead. It says it has begun its "most ambitious pre-training run yet" for Gemini 4. That is a statement of intent, not a shipped capability. The efficiency drive runs deeper than software. Google is also building a custom chip to serve Gemini far more cheaply. One caveat is worth keeping, though: every benchmark here is Google's own, and no outsider has checked them yet. The bet The plan is coherent. In a year when companies are counting tokens, cheap and fast can win the middle of the market while the flagship catches up. Whether it holds rests on two things: 3.5 Pro finally shipping, and Gemini 4 turning out to be more than a training run.
[8]
Google expands Gemini 3.5 line with trio of new models -- and shares an update on Gemini 3.5 Pro
Google says it plans to make Gemini 3.5 Pro available broadly soon, and the team has started pre-training Gemini 4. Back in May, Google released Gemini 3.5 Flash, its smartest speed model at the time. It's been only about two months since then, but a new shiny AI model is ready to take its place. Google has announced the launch of Gemini 3.6 Flash. Along with 3.6 Flash, the company is also launching 3.5 Flash-Lite and 3.5 Flash Cyber in CodeMender. On top of that, the model will be available at a lower cost than 3.5 Flash. The company has set the pricing at $1.50/1M input tokens and $7.50/1M output tokens. But it's not just about efficiency and cost; performance also improves with this new model. According to Google, 3.6 Flash is more precise, delivers fewer unwanted code edits, and reduces execution loops. Computer use has improved from 78.4% to 83% and there's now a built-in client-side tool via the Gemini API and Gemini Enterprise. As the graphs above show, knowledge work benchmarks also favor the new 3.6 Flash. Google says that 3.5 Flash-Lite offers "significantly better quality than 3.1 Flash-Lite." It's also said to significantly outperform the previous model. And 3.5 Flash-Lite doesn't just outclass its predecessor, Google claims it even outperforms Gemini 3 Flash. In an update, Google adds that Gemini 3.5 Pro is making progress, as it continues to test the model with partners. The company says it plans to roll out 3.5 Pro broadly soon. Additionally, Google says that it has started pre-training the next major model, Gemini 4.
[9]
Google launches Gemini 3.6 Flash and 3.5 Flash-Lite, teases Gemini 4
As we wait for 3.5 Pro, Google today announced Gemini 3.6 Flash and 3.5 Flash-Lite, while providing updates on what comes next. Gemini 3.6 Flash Gemini 3.6 Flash follows the last release at I/O 2026. This update takes into account developer and customer feedback since May by being "more token efficient across tasks." Compared to 3.5 Flash, it consumes 17% fewer output tokens (per the Artificial Analysis Index), while taking "fewer reasoning steps and tool calls to accomplish multi-step workflows." At the same time, it's priced lower at $1.50/1M input tokens and $7.50/1M output tokens (versus $9/1M output). In terms of coding performance, Gemini 3.6 Flash "delivers higher precision with fewer unwanted code edits and reduced execution loops" than its predecessor. Specifically: It generates higher quality and more reliable, production-ready code as seen in DeepSWE (49% vs. 37%), and shows significant improvement in ML Research, as seen in MLE Bench (63.9% vs. 49.7%). For knowledge work, the model scores 1421 on GDPval-AA (versus 1349). Computer use capabilities go from 78.4% on OSWorld-Verified to 83%. The knowledge cutoff date finally advances from January 2025 to March 2026. Gemini 3.5 Flash-Lite The company also announced Gemini 3.5 Flash-Lite for high-throughput and low-latency tasks, like agentic search and document processing. Google says it offers "significantly better quality than 3.1 Flash-Lite" from March, while pricing is $0.30/1M input tokens and $2.50/1M output tokens. It's a significant step up in coding and agentic tasks as seen in Terminal-Bench 2.1 (54% vs 31%), long context as seen in GDM-MRCR v2 (72.2% vs. 60.1%), and real-world task execution as seen in GDPval-AA v2 (1140 vs. 642). Google notes that 3.5 Flash-Lite outperforms 3 Flash: SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%). Gemini 3.5 Flash Cyber The final announcement today is Gemini 3.5 Flash Cyber for finding and fixing security vulnerabilities. Flash serves as the foundation given its performance and efficiency, with the goal of detecting, validating, and patching "code security issues at scale" and at a "lower price per token than larger models." Google's CodeMender tool uses multiple 3.5 Flash Cyber agents. Access is initially for governments and trusted partners "as part of a limited-access pilot program." This will give frontline defenders a head start in finding and fixing critical vulnerabilities before they can be exploited, while mitigating against broader misuse. Availability Gemini 3.6 Flash and 3.5 Flash-Lite are available today in the Gemini app, with the latter model also coming to Search. Developers can access them through Google Antigravity, AI Studio, and Android Studio. Looking ahead, Google reiterates that "Gemini 3.5 Pro is currently testing with partners," with broad availability "as soon as it's ready." Meanwhile, the DeepMind team is "already focusing on building the next generation of models." We have already started our most ambitious pre-training run yet, for Gemini 4, and can't wait to share more.
[10]
Google has new Gemini models, but the one everyone wants isn't here
Google promises Gemini 3.5 Pro is coming "soon" after previous delays, leaving users with Flash models suitable for everyday tasks but not advanced work. Google just unleashed three new versions of its speedy and efficient Gemini Flash models, including an all-round workhorse and a model geared toward cybersecurity. But the arrival of the new Flash models also highlights what's missing from the lineup: a new "Pro" model. Both Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are available now in the Gemini app, while Gemini 3.5 Flash Cyber will initially be released only for government use and to "trusted partners." The limited release follows the precedent set by Anthropic and OpenAI, which offered Mythos 5 and GPT-5.6 Sol, respectively, in limited release due to their advanced cybersecurity abilities. (Mythos 5 is still restricted to "trusted partners.") But while we now have new Gemini Flash models, we're still stuck with Gemini 3.1 Pro, which was released way back in February -- ancient history as far as the AI market is concerned. In contrast, OpenAI's latest frontier model, GPT-5.6 Sol, came out at the end of last month. Anthropic's Claude Fable 5 and Mythos 5 also landed in June, with Fable 5 briefly getting yanked after the sudden imposition of US government export controls. Prior to that, we also saw the release of GPT-5.5 (April) and Claude Opus 4.8 (in May). In a press release, Google said it will be "releasing 3.5 Pro soon." Many (myself included) expected Gemini 3.5 Pro to arrive in May during Google's I/O conference, but that didn't happen. Speculation about 3.5 Pro's status grew following a recent Bloomberg article, which claimed the model remains in development after "disappointing" testing results. For its part, Google says Gemini 3.6 Flash is good for coding and "knowledge work," as well as multimodal tasks like analyzing charts and parsing documents. Gemini 3.5 Flash-Lite, meanwhile, is optimized for low-latency tasks such as agentic search. While speedy and efficient models like Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are great for summarizing documents, planning meals for the week, or performing everyday AI duties, "pro"-tier models like Claude Opus and Fable or GPT-5.6 Sol are best suited for more complex tasks, like full-blown coding, working in large Excel documents, or composing lengthy prompts. Personally, I've largely abandoned Gemini Pro in favor of newer pro AI models like Claude Opus 4.8 and GPT-5.6, although that could all change once Gemini 3.5 Pro finally drops. The big question, though, is when?
[11]
Google's next Gemini Pro is months behind schedule as coding capabilities fall short of internal goals
Google is months behind on its flagship Gemini Pro upgrade as coding falls short, frustrating engineers who are leaving for Anthropic. Google is months behind schedule on delivering the next version of its flagship AI model, Gemini Pro, because the technology has fallen short of internal goals in coding, Bloomberg reported on Thursday citing 10 current and former employees. The company was widely expected to release the upgrade at its May developer conference but has been unable to close the gap with Anthropic and OpenAI, which have both released models that outperform Google's current offerings in writing code. Alphabet shares slipped more than three percent on the news. Late last month, Google updated the data used to train Gemini in an attempt to improve its coding abilities, but the results were disappointing, according to one of the people Bloomberg spoke with. Both OpenAI and Meta recently released new models that further outpace Google's current AI for writing code, intensifying pressure on a team already struggling to ship. A Google spokesperson said the company is "shipping quickly across a wide range of models" and is testing the upgraded Pro, a new Flash model, and other models with partners. Part of the problem is structural. Google Cloud, DeepMind, and the Android team are all building AI coding tools for developers, with involvement from consumer product teams as well, creating internal competition that has slowed progress. Co-founder Sergey Brin has been pushing for the company to move faster on AI coding, but his efforts have been hampered by competing factions and by engineers who believe important code should still be written by humans to meet Google's standards, according to former employees. Google has taken steps to consolidate its fragmented coding efforts. Chief AI Architect Koray Kavukcuoglu is working to unite the company's internal AI coding tools, and a new team within DeepMind led by research engineer Sebastian Borgeaud has been formed specifically to tackle the problem. The company said at its most recent Cloud conference that 75 percent of code at Google is now AI-generated and that it has consolidated most of its developer tooling under Antigravity, the internal platform that manages data, memory, and safety protocols for AI applications. The delays have contributed to a wave of senior departures to Anthropic and other labs, with former employees saying frustration with Google's competitive position is a driving factor. Engineers who try to use AI for their own work often hit capacity constraints due to internal competition for computing power, a problem that extends to external customers as well. Only some teams inside Google are even allowed to use Anthropic's Claude, with access restricted to groups doing cutting-edge research. Customers waiting for the Pro upgrade have had mixed experiences with the current Flash model. Rodrigo Davies, a product manager at Figma, said the model hit "a sweet spot of speed and quality" for the design platform's AI assistant. But Freddy Vega, CEO of Latin American education platform Platzi, said the Flash model is more expensive and slower than its predecessor while remaining far less capable than competitors, and his team has shifted to Anthropic instead.
[12]
Google's next flagship Gemini model reportedly stuck months behind schedule
Coding performance appears to be a major problem, despite Google recently updating the model's training data. Google will certainly feel that it belongs at the front of the AI race with the other big hitters, but its next big Gemini model is said to be struggling to get over the line. Gemini 3.5 Pro is reportedly months behind schedule as Google works to bring it up to its internal standards. According to Bloomberg, the delay is based on information from people familiar with the matter, along with ten current and former Google employees. The setback has reportedly frustrated engineers, researchers, and managers inside the company, with some worried that Anthropic and OpenAI are beginning to pull further ahead. Google had reportedly been widely expected to unveil Gemini 3.5 Pro at its developer conference in May. Coding appears to be one of the main sticking points, with Bloomberg saying Google updated the model's training data late last month in an effort to improve those abilities, only for the results to disappoint. The report suggests Google's sheer size may be working against it. Multiple teams across Google Cloud, DeepMind, Android, and other parts of the company are building AI coding tools, while several layers of stakeholders are involved in preparing models for release. Employees are also said to face competition for computing power when trying to use AI internally. Google pushed back on the idea that it is moving too slowly, with a spokesperson saying it is "shipping quickly across a wide range of models" while keeping them cost-effective. The company also confirmed that it is testing Gemini 3.5 Pro, an upgraded Flash model, and other models with partners, while discussing model testing and safety standards with the US government.
[13]
Google releases series of new cheaper Gemini models
Why it matters: The AI deployment race has shifted from benchmark bragging rights to who can provide the best model at the lowest price. * Gemini 3.6 Flash: a faster, cheaper successor to 3.5 Flash that Google says uses up to 17% fewer output tokens while improving coding, reasoning and multimodal performance. * Gemini 3.5 Flash-Lite, an even lower-cost model designed for high-volume AI agents and document processing. * Gemini 3.5 Flash Cyber, a security-focused model built to identify and patch software vulnerabilities, initially available only to governments and selected partners through Google's CodeMender platform. Yes, but: The release does not include Gemini 3.5 Pro, a more powerful model that Bloomberg last week reported was months behind schedule. * Google Tuesday said Gemini 3.5 Pro is in testing and will be broadly available once it's ready, while Gemini 4 is in pre-training. * Shares of Google parent Alphabet were down 0.8% midday Tuesday and nearly 6% since Thursday. Zoom in: Each of the models introduced Tuesday appears to address a key focus for enterprise customers. * 3.6 Flash costs less per token than its predecessor while improving coding accuracy and reducing unnecessary reasoning steps, which has become an increasingly important metric as companies try to rein in AI costs. * Flash-Lite targets the growing market for AI agents that need to execute millions of relatively simple tasks cheaply rather than maximize benchmark performance, a sign of the increase in enterprise agent usage. * The cybersecurity model reflects Google's broader push to commercialize DeepMind research through enterprise security products. Zoom out: Rather than betting on a single flagship model, Google is building a portfolio optimized for different workloads, prices and industries. State of play: The releases arrive during one of the fiercest recruiting wars in Silicon Valley. * Google has spent months trying to retain key DeepMind researchers as top talent have left the company, citing its ties with the U.S. military. * Frontier competitors, meanwhile, including OpenAI, Meta and Thinking Machines, have been aggressively recruiting top talent with lucrative compensation packages. * Meta alone has hired a number of prominent Google researchers in recent months, highlighting that the battle for AI leadership increasingly hinges on people as much as models. The bottom line: The AI race is increasingly becoming less about producing the single smartest model and more about giving developers and businesses choice about which model is best for the job.
[14]
Google releases two new Gemini models, but no Gemini 3.5 Pro
Google has introduced two additions to its Gemini "Flash" lineup, along with a specialized cybersecurity model, 3.5 Flash Cyber, while holding off on releasing a long-awaited Gemini 3.5 Pro. In a July 21 blog post, Google announced Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, describing them as built to give developers the "efficiency, low latency, and reliability" needed to run AI agents at scale. The post also detailed Gemini 3.5 Flash Cyber, a specialized model paired with Google's CodeMender security agent. Our big Guessing Game is back! Enter now for a chance to win an Apple Watch. The announcement also included a brief update on Gemini 3.5 Pro, which Google says "is currently testing with partners," with general availability "as soon as it's ready." The post added that Google has already begun training runs for a future Gemini 4 model. At the Google I/O event in May, Alphabet CEO Sundar Pichai announced that Gemini 3.5 Pro would launch in June. Since then, Anthropic and OpenAI have launched new frontier models, Fable 5 and GPT-5.6 Sol, while the Chinese AI lab Moonshot recently launched Kimi K3, a more affordable but still competitive open-source AI model. It's against this backdrop that Google announced its new trio of Gemini models. According to the post, Gemini 3.6 Flash is positioned as the company's general-purpose "workhorse" model, offering improvements in coding, knowledge work, and multimodal tasks compared to 3.5 Flash. Google also priced the model lower than 3.5 Flash, at $1.50 per million input tokens and $7.50 per million output tokens. Google released benchmark evaluations of Gemini 3.6 Flash, showing improvements in a variety of domains. Gemini 3.5 Flash-Lite, meanwhile, is billed as the fastest model in the 3.5 series, reportedly capable of 350 output tokens per second, per Artificial Analysis figures cited in the post, at a lower price point of $0.3 per million input tokens and $2.5 per million output tokens. Google writes that the model is intended for high-throughput, latency-sensitive workloads like agentic search and document processing, and that it outperforms the prior Flash-Lite generation across agentic benchmarks. The most narrowly targeted release, Gemini 3.5 Flash Cyber, is designed specifically to detect and patch software vulnerabilities and is being deployed inside CodeMender, Google's code security agent. Per the post, the model will not be broadly available; Google says it plans to offer it exclusively to governments and select partners through a limited-access pilot, citing the dual-use risks associated with a model capable of finding security flaws. Google also said the newly released 3.6 Flash model includes expanded safety measures to resist jailbreak attempts related to chemical, biological, radiological, and nuclear misuse, as well as cyber offenses, while aiming to avoid unnecessary refusals for legitimate use cases. Disclosure: Ziff Davis, Mashable's parent company, in April 2025 filed a lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.
[15]
Google's Gemini 3.6 Flash model cuts AI agent token costs by up to 65% on long horizon engineering tasks -- and 3.5 Pro is on the way
Google DeepMind today released three new proprietary AI models it says are among its most token-efficient yet: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. The models aim to make AI agents faster, smarter, and cheaper at scale. Google is pricing Gemini 3.6 Flash at $1.50 per one million input tokens and $7.50 per one million output tokens through its application programming interface (API), while Gemini 3.5 Flash-Lite costs a staggeringly cheap $0.30/$2.50 per million tokens in/out. Compare that to the $1.50/$9.00 per 1M tokens for Gemini 3.5 Flash, and the $2/$12 for Gemini 3.1 Pro Preview, and the savings are considerable. However, Google's prior generation Gemini 3.1 Flash-Lite still remains the search giant's "most cost-efficient" model at $0.25/$1.50 per 1M tokens. Yet, it remains 2X slower than the new, more expensive Gemini 3.5 Flash-Lite, giving those enterprises who value speed more "bang" for their buck. VB Frontier AI Model API Pricing Comparison Chart (Late July 2026 Shortlist) No price was provided yet for the specialty Gemini 3.5 Flash Cyber model, which, as its name would imply, is designed for cybersecurity researchers and red teamers to patch bugs. While the prices are among the middle-low end of all major AI models globally, the fact that Google designed them to use less tokens overall also should drive down costs for enterprises beyond what the sticker price shows (since you'll be paying for fewer total tokens at any rate). Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are available immediately through the Gemini API in Google AI Studio and Android Studio, as well as within the consumer Gemini application and Google Search. According to a separate Google blog post, Gemini 3.5 Flash Cyber will be available "exclusively available to governments and trusted partners via CodeMender soon" -- CodeMender being Google's proprietary AI code bug-fixing agent released last year. As with previous Gemini models, these are all proprietary and "closed source," thus, they can only be obtained through Google's official API and that of its partners, as opposed to an open-source license like MIT or Apache 2.0. One conspicuous omission noted by developers on X and social media: where is the larger, more powerful, flagship Gemini 3.5 Pro model Google previously alluded would be released this summer? After all, Gemini 3.1 Pro, the prior flagship, debuted back in February 2026, and rivals OpenAI and Anthropic have since released several more generations of flagship updates far more powerful than Google's. Google technical staffer Logan Kilpatrick responded to one such inquiry on X, writing: "Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it's ready." Google's release signals that the immediate future of AI lies in agentic capabilities -- systems that operate autonomously over extended periods. If early large language models are akin to massive, fuel-hungry freight trains capable of hauling incredible loads at immense cost, the new Flash series represents a fleet of nimble, hyper-efficient hybrid delivery vans. Efficiency gains ranging from 17% to 65% reduced tokens for strong results on third-party benchmarks Under the hood, Gemini 3.6 Flash achieves significant efficiency gains. The model reduces output token usage by 17% compared to its predecessor, Gemini 3.5 Flash, according to the Artificial Analysis Index maintained by the independent third-party AI benchmarking group of the same name. In specific long-horizon software engineering benchmarks like DeepSWE, which measures how well agents complete multi-step engineering tasks from scratch, the token savings reach up to 65%. This reduction means the model requires fewer reasoning steps and tool calls to complete the exact same multi-step workflow. Think of token efficiency like fuel economy in a vehicle. When an AI model takes a convoluted path to solve a problem, it burns through more computational fuel, driving up the final cost for the developer. By streamlining its internal logic, Gemini 3.6 Flash arrives at the correct answer faster and cheaper. While Google's materials did not specify the exact architectural or algorithmic changes used to achieve this token efficiency, they noted that the model "takes fewer reasoning steps and tool calls to accomplish multi-step workflows" and exhibits reduced "verbosity." The official model cards released by Google reveal that both Gemini 3.6 Flash and Gemini 3.5 Flash-Lite feature a 1-million-token input context window alongside a max output limit of 64,000 tokens, with both models sharing a knowledge cutoff date of March 2026. Respectable benchmark performance at low cost The technological improvements extend to concrete capabilities. Gemini 3.6 Flash scores 49% on the DeepSWE benchmark, a notable increase from the 37% achieved by version 3.5. It also pushes machine learning engineering performance higher, scoring 63.9% on MLE-Bench compared to 49.7% previously. Furthermore, Google integrates computer use as a built-in client-side tool via the Gemini API and Gemini Enterprise, reflecting an OSWorld-Verified score of 83.0%, up from 78.4%. The model also tackles knowledge work with greater proficiency, outperforming its predecessor on benchmarks like GDPval-AA v2 by moving from a score of 1349 to 1421. To ensure safety amidst these capability upgrades, Google deploys enhanced Frontier Safety safeguards. These protections harden the model against jailbreaks and mitigate risks in Chemical, Biological, Radiological, and Nuclear domains, as well as cyber offense misuses. The engineering team trains the model to minimize refusals for beneficial uses, striking a necessary balance between strict security and practical utility. Models for low-cost coding, agentic, and cybersecurity use cases -- respectively Google divided its new offerings into three distinct products tailored for different operational needs. Gemini 3.6 Flash serves as the heavy-duty workhorse of the trio. It handles complex coding, intricate knowledge work, and multimodal processing with improved precision. Enterprise customers utilize it for demanding tasks such as complex document parsing, intricate chart and data analysis, and long-form report drafting. The model executes complex code migrations using multi-agent orchestration frameworks with lower latency and higher quality than earlier iterations. Furthermore, 3.6 Flash aids in developing photographic texture extractors for 3D workflows using canvas interfaces. Gemini 3.5 Flash-Lite targets environments where high throughput and absolute minimal latency are non-negotiable. Google designates it as the fastest model in the 3.5 series. As measured by Artificial Analysis, the model processes 350 output tokens per second, making it highly effective for agentic search and massive document processing workloads. Artificial Analysis notes this is about twice as fast as prior generation model Gemini 3.1 Flash-Lite. Developers can configure 3.5 Flash-Lite to prioritize low-latency execution for high-volume tasks using minimal thinking levels, or engage higher thinking levels to process complex multi-step subagent workloads. Despite its lite designation, it outperforms the standard Gemini 3 Flash on several key agentic and coding evaluations, including SWE-Bench Pro, where it scores 54.2% compared to 49.6%, and OSWorld-Verified, scoring 74.0% versus 65.1%. The model extracts product features from massive datasets, generates interactive web design concepts, and scales receipt translation seamlessly. The third product, Gemini 3.5 Flash Cyber, represents a highly specialized deployment. Google fine-tuned this model specifically to find and fix cybersecurity vulnerabilities. It integrates directly with Google's CodeMender agent. In practice, multiple 3.5 Flash Cyber agents work concurrently to produce a single, comprehensive vulnerability report, achieving competitive performance at the frontier on the CyberGym benchmark. Google did not specify an exact numerical cost for 3.5 Flash Cyber, stating only that it is fine-tuned "at a lower price per token than larger models. Commercial licensing only The licensing framework for the new Gemini models carries profound implications for developers and enterprise users. Google deploys Gemini 3.6 Flash and 3.5 Flash-Lite under a commercial, proprietary API model. Unlike open-source software governed by licenses such as the MIT License or the GNU General Public License, developers do not gain access to the underlying model weights, training data, or source code. An MIT or GPL license grants users the freedom to download the codebase, modify the internal architecture, self-host the deployment, and distribute the software infrastructure independently. In contrast, Google's API approach means developers essentially rent access to the intelligence on a strict metered basis. Every prompt and generated response travels through Google's managed servers, incurring a cost based on the strict pricing structure of $1.50 per million input tokens for 3.6 Flash. This commercial tethering restricts deployment flexibility. Enterprises cannot air-gap the models entirely on their own local secure hardware without establishing specialized, high-tier enterprise agreements with Google Cloud. Developers remain bound by Google's acceptable use policies, arbitrary rate limits, and network requirements, creating a permanent dependency on Google's infrastructure uptime and terms of service. The licensing for Gemini 3.5 Flash Cyber proves even more restrictive. Acknowledging the dual-use nature of cybersecurity AI -- which attackers can weaponize just as easily as defenders can use it to patch systems -- Google is for now making the model only available behind a limited-access pilot program, similar to the trend kicked off by Anthropic's Mythos model with its Project Glasswing program, and continued by OpenAI with its staggered rollout for GPT-5.6. In this case, Google is making 3.5 Flash Cyber exclusively available to governments and trusted partners. This strict gatekeeping prevents open access, prioritizing systemic security over widespread developer innovation. Looking ahead Google DeepMind continues to iterate rapidly, but the gap in its product line remains apparent. While the Flash series excels in speed and economy, the industry eagerly awaits the deployment of Gemini 3.5 Pro to gauge Google's absolute frontier capabilities. Simultaneously, the company confirms that pre-training for Gemini 4 has already commenced. Until the next major flagship release materializes, developers must optimize their systems using the highly efficient, yet purposefully constrained, Flash architecture.
[16]
Google launches 3 new Gemini AI models including cybersecurity tool
Gemini 3.5 Flash Cyber, built to find and patch software vulnerabilities, will initially be available only to governments and trusted partners Google $GOOGL announced three new Gemini AI models on Tuesday, including a cybersecurity-focused model designed to detect and patch software vulnerabilities, which the company said will be restricted to governments and trusted partners at launch. The cybersecurity model, Gemini 3.5 Flash Cyber, uses Gemini 3.5 Flash as its foundation and has been specialized to detect, confirm, and remediate code vulnerabilities. It operates within Google's CodeMender agent, an AI-powered tool for vulnerability discovery and patching that the company unveiled in October 2025, according to The Hacker News. By invoking 3.5 Flash Cyber repeatedly, CodeMender is able to broaden its coverage across code paths and consolidate findings into one report. Google said the model is available at a lower price per token than larger cybersecurity models and achieves competitive performance on the CyberGym benchmark, which tests AI agents against real-world software vulnerabilities. The company added that when tested on the V8 JavaScript Engine, Gemini 3.5 Flash Cyber found 55 unique confirmed vulnerabilities, compared with 47 for Gemini 3.5 Flash and 36 for Anthropic Claude Opus 4.6, including 10 issues no other model identified. "Given the dual-use nature of this technology, we have taken an intentional approach to how we deploy 3.5 Flash Cyber," DeepMind's Gemini Security Lead Raluca Ada Popa and Vice President of Security and Privacy Four Flynn said in a statement. "This will give frontline defenders a head start in finding and fixing critical vulnerabilities before they can be exploited, while mitigating against broader misuse." According to CNBC, the release positions Google to close ground on Anthropic, which has established itself as an early leader in automated code defense through its Mythos model. The other two models launched Tuesday are available to the public. Gemini 3.6 Flash offers stronger coding and multimodal capabilities than the model it replaces and requires 17% fewer output tokens to complete tasks, the company said. It is priced at $1.50 per million input tokens and $7.50 per million output tokens, down from $9 per million output tokens for Gemini 3.5 Flash. Gemini 3.5 Flash-Lite sits at the bottom of the 3.5 series on both speed and cost, generating output at 350 tokens per second with pricing set at $0.30 per million input tokens and $2.50 per million output tokens. Both models are available through the Gemini app, Google AI Studio, and Android Studio. Google added that Gemini 3.5 Pro remains in a partner testing phase before a wider rollout, and that work on Gemini 4 is underway with the most ambitious pretraining run the company has undertaken.
[17]
Gemini 3.5 Pro delays due to coding performance, upgraded Flash model in testing
In mid-May, Google announced Gemini 3.5 Flash at I/O 2026 and said the Pro version would arrive in June. On stage, Google said it was "showing great improvements." That deadline has passed with no update on when to expect it. According to Bloomberg, Google is "taking time to try to improve [Gemini 3.5 Pro's] capabilities, particularly in coding." In late June, "Google updated the data being used to train Gemini in an attempt to improve [coding] skills, but the results were disappointing." That timeline suggests that development saw a reset between I/O and the missed launch. It's unclear how the latest models are performing in other domains. Gemini 3.1 Pro dates back to February. In a statement, the company said it is "currently testing 3.5 Pro, an upgraded Flash model, and other models with partners." Google also added: "We're shipping quickly across a wide range of models while keeping them highly cost-effective for customers." Since the developer conference, updates to the Gemini app have focused on improving the user experience and rolling out the Spark agent. Today's article also has insight on the use of coding tools internally. As of April, "75% of all new code at Google is now AI-generated and approved by engineers, up from 50% last fall." Efforts to win at coding have also been up against some engineers at Google with a more purist stance, who believe that all important code should be human-written to adhere to Google standards, ex-employees said. Additionally, according to reports, engineers internally are facing AI capacity restraints with tools. An effort to "unite the company's internal artificial intelligence coding tools" is underway. In terms of developing AI coding tools for the public, Google DeepMind (AI Studio), Cloud (Vertex), and the Android team (Android Studio) all have their own efforts.
[18]
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Developers and customers building production AI agents need higher token efficiency, lower latency, and more reliable performance. Our Flash series of models is built to meet the sweet spot of efficiency and quality to enable scaling agentic workflows. Building on Gemini 3.5 Flash, we're introducing new Gemini models: * 3.6 Flash: Our workhorse model that delivers better coding, knowledge work, and multimodal performance. According to the Artificial Analysis Index, it reduces output token usage by 17% compared to 3.5 Flash, and in some benchmarks like DeepSWE by Datacurve, we observe up to 65%, all at a lower cost per output token. * 3.5 Flash-Lite: Our fastest, most cost-effective 3.5-class model, delivering 350 output tokens per second according to the Artificial Analysis Index, also significantly outperforming prior Flash-Lite generations in agentic workflows. * 3.5 Flash Cyber in CodeMender: Successful cybersecurity applications require careful orchestration of a model alongside an agent infrastructure. We're introducing a combination of a new, highly efficient, specialized cyber-focused model paired with our CodeMender code security agent that delivers competitive performance at the frontier. Beyond today's releases, Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it's ready. In parallel, our team is already focusing on building the next generation of models. We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress. 3.6 Flash: More efficient and better quality than 3.5 Flash Gemini 3.6 Flash builds directly on developer and customer feedback from 3.5 Flash. 3.6 Flash not only delivers a step up in coding and knowledge work, but it does this while meaningfully improving token efficiency. For example, on the Artificial Analysis Index, we see 3.6 Flash consuming 17% fewer output tokens than 3.5 Flash. It also takes fewer reasoning steps and tool calls to accomplish multi-step workflows. This enhanced efficiency is also combined with a lower price than 3.5 Flash. At $1.50/1M input tokens and $7.50/1M output tokens, 3.6 Flash reduces the overall cost per agentic task, making agents more cost-effective to build and run.
[19]
Google Ships New Gemini Flash Models, But Pro Is Still Missing
Google confirmed it has begun pre-training for Gemini 4, which it calls "our most ambitious pre-training run yet." Google launched three new AI models today: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. That wasn't what most people expected. After unveiling Gemini 3.5 Flash at Google I/O 2026 in May and promising a Pro version within a month, Google quietly missed its own deadline. Gemini 3.5 Pro was held back because it fell short of internal targets, per Bloomberg, particularly on coding tasks. A late-June attempt to fix it by updating the training data -- the massive datasets a model learns from -- produced disappointing results. Alphabet stock fell roughly 4.4% on the report, erasing an estimated $200 billion in market cap in a single session. The last Pro-tier model Google shipped was Gemini 3's successor, Gemini 3.1 Pro, back in February. The Flash series is Google's line of speed-optimized models -- fast, cost-effective, and built for AI agents, which are programs that operate semi-autonomously to handle tasks like managing documents, processing data pipelines, or browsing the web without a human clicking through each step. Pro models are the heavy lifters: slower, pricier, and built for complex reasoning where raw power matters more than speed. What each AI model does -- and who it's for Gemini 3.6 Flash is the main release. It uses 17% fewer output tokens -- tokens being the basic unit AI processes, roughly three-quarters of a word -- than 3.5 Flash, per the Artificial Analysis Index. It's also cheaper: $1.50 per million input tokens and $7.50 per million output tokens, down from $9 on the output side for 3.5 Flash. For businesses running agents at scale, that difference compounds fast. On benchmarks -- standardized tests that score AI by percentage of tasks completed correctly -- 3.6 Flash hit 49% on DeepSWE v1.1, which tests long-horizon software engineering like building and debugging full codebases, versus 37% for 3.5 Flash. On MLE-Bench, a machine learning engineering test, it scored 63.9% versus 49.7%. It topped the table on OSWorld-Verified -- a test where the AI takes control of a computer screen to complete real tasks -- at 83.0%, ahead of Claude Sonnet 5 (81.2%) and GPT-5.6 Luna (72.6%). Rivals in the same category still lead elsewhere: GPT-5.6 Luna scores 67% on DeepSWE and 84.7% on Terminal-Bench 2.1, which tests agentic terminal coding. Claude Sonnet 5 tops knowledge work on GDPval-AA v2 -- a benchmark scored on an Elo rating scale like chess, where higher numbers mean better real-world task performance -- at 1607 versus 3.6 Flash's 1421. We tried the model for coding and the results were... underwhelming to say the least. Our simple coding test ended up with an unusable file. The HTML was not properly formatted, and elements were not rendered correctly. Subsequent attempts to vibe code a way to solve the issues were not successful. We asked Deepseek to turn the first model into something playable by simply fixing the bugs. It identified 11 bugs and implemented 8 key fixes, which resulted in a decent game. Deepseek's small tweaks fixed the game, which means, Gemini's core thinking was correct, but the details and inaccuracies made the result unuseful. Prepare for long vibe coding sessions with a cheap yet poor performing model if you pretend to use the model for that. The second model released by Google, Gemini 3.5 Flash-Lite, is built purely for volume: 350 output tokens per second at $0.30/million input and $2.50/million output. It's aimed at high-throughput pipelines -- think document processing at massive scale or agentic search systems -- and outperforms the older 3 Flash on key coding tasks, including Terminal-Bench 2.1 (54% vs. 31%), despite being significantly cheaper. It could also be a great session compactor (analyzing long sessions and extracting the key elements so your agent doesn't collapse with noise) for those relying on Hermes and Openclaw. The third model, Gemini 3.5 Flash Cyber, won't be publicly available. Google is restricting it to governments and vetted partners who need to find and fix software vulnerabilities -- a dual-use capability the company is not comfortable releasing broadly. Meanwhile, Google's DeepMind team is already moving on. Google confirmed in the official announcement that it has started "our most ambitious pre-training run yet, for Gemini 4," and the team is already hyping it up. Pre-training is the foundational phase where a model learns from massive datasets before task-specific fine-tuning begins -- meaning Gemini 4 is being built, not planned. Both 3.6 Flash and 3.5 Flash-Lite are live today in the Gemini app, Google AI Studio, and via the API. Gemini 3.5 Pro will ship, per Google, "as soon as it's ready," whenever that is.
[20]
Where is Gemini 3.5 Pro? The case of the missing AI model.
Gemini users who were hoping to see the launch of Gemini 3.5 Pro at Google I/O 2026 left disappointed. At the May developers' conference, the company launched a lighter-weight Gemini 3.5 Flash model for everyday use. However, Google CEO Sundar Pichai assured the audience that Gemini 3.5 Pro would follow in June. "We are also excited for 3.5 Pro," Pichai said at a pre-Google I/O media briefing. "We are using it internally. It's showing great improvements. We are still testing and refining it, and it will roll out to everyone next month." As of July 17, there's still no sign of the model. Our big Guessing Game is back! Enter now for a chance to win an Apple Watch. So, where is Gemini 3.5 Pro? Yesterday, Bloomberg published a report on the delayed launch, with reporters Julia Love and Davey Alba writing that "The delay has caused frustration among Google engineers, AI researchers, and managers, who are concerned the company risks losing its edge in the market to rivals Anthropic and OpenAI." Mashable reached out to Google with questions about the Gemini 3.5 Pro launch timeline, and the company provided the same statement it shared with Bloomberg. "We're shipping quickly across a wide range of models while keeping them highly cost-effective for customers. We're currently testing 3.5 Pro, an upgraded Flash model, and other models with partners, and we're productively engaged with the U.S. government on model testing and broader frameworks." While a one-month delay isn't normally a massive problem, the AI industry has been moving at lightning speed in recent months. The longer Google waits to release Gemini 3.5, the higher the performance bar it has to clear to maintain equal footing with its rivals. According to Bloomberg, even Meta has released a new model that outpaces Google Gemini. Bloomberg's report suggests that there are two reasons driving the delay of Gemini 3.5 Pro. The first is bureaucratic. Because of the size of Google's organization and the number of products integrated with Gemini, delays are inevitable compared to leaner AI startups. Second, Bloomberg found that Google leaders are worried that Gemini 3.5 Pro may not be competitive with their rivals' recent releases. Since Google I/O 2026, Anthropic announced it was launching its most advanced model ever, Claude Mythos Preview. The AI company said the model had such advanced cybersecurity capabilities that it would only be shared with trusted partners. Anthropic eventually did release a version of Claude Mythos called Fable 5 on June 9. On July 9, OpenAI announced its own next-generation model with advanced cybersecurity coding abilities, GPT‑5.6 Sol. This week, Chinese AI lab Moonshot released Kimi K3, a massive open-source model with 2.8 trillion parameters. Early testers say it has similar capabilities as Fable 5 and GPT-5.6 Sol, only with a much lower cost. Without a new frontier model of its own, Google has taken a tumble down AI leaderboard rankings, despite its massive advantages in the AI arms race. Google not only has unprecedented access to the world's data, but it can also put its AI tools directly into the hands of billions of Android users worldwide. Gemini 3.5 Pro may be launching soon, but while Google readies the model for release, the competition is racing ahead.
[21]
Google's Gemini Flash 3.6 model cuts AI agent token costs by up to 65% on long horizon engineering tasks -- and 3.5 Pro is on the way
Google DeepMind today released three new proprietary AI models it says are among its most token-efficient yet: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. The models aim to make AI agents faster, smarter, and cheaper at scale. Google is pricing Gemini 3.6 Flash at $1.50 per one million input tokens and $7.50 per one million output tokens through its application programming interface (API), while Gemini 3.5 Flash-Lite costs a staggeringly cheap $0.30/$2.50 per million tokens in/out. Compare that to the $1.50/$9.00 per 1M tokens for Gemini 3.5 Flash, and the $2/$12 for Gemini 3.1 Pro Preview, and the savings are considerable. However, Google's prior generation Gemini 3.1 Flash-Lite still remains the search giant's "most cost-efficient" model at $0.25/$1.50 per 1M tokens. Yet, it remains 2X slower than the new, more expensive Gemini 3.5 Flash-Lite, giving those enterprises who value speed more "bang" for their buck. VB Frontier AI Model API Pricing Comparison Chart (Late July 2026 Shortlist) No price was provided yet for the specialty Gemini 3.5 Flash Cyber model, which, as its name would imply, is designed for cybersecurity researchers and red teamers to patch bugs. While the prices are among the middle-low end of all major AI models globally, the fact that Google designed them to use less tokens overall also should drive down costs for enterprises beyond what the sticker price shows (since you'll be paying for fewer total tokens at any rate). Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are available immediately through the Gemini API in Google AI Studio and Android Studio, as well as within the consumer Gemini application and Google Search. According to a separate Google blog post, Gemini 3.5 Flash Cyber will be available "exclusively available to governments and trusted partners via CodeMender soon" -- CodeMender being Google's proprietary AI code bug-fixing agent released last year. As with previous Gemini models, these are all proprietary and "closed source," thus, they can only be obtained through Google's official API and that of its partners, as opposed to an open-source license like MIT or Apache 2.0. One conspicuous omission noted by developers on X and social media: where is the larger, more powerful, flagship Gemini 3.5 Pro model Google previously alluded would be released this summer? After all, Gemini 3.1 Pro, the prior flagship, debuted back in February 2026, and rivals OpenAI and Anthropic have since released several more generations of flagship updates far more powerful than Google's. Google technical staffer Logan Kilpatrick responded to one such inquiry on X, writing: "Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it's ready." Google's release signals that the immediate future of AI lies in agentic capabilities -- systems that operate autonomously over extended periods. If early large language models are akin to massive, fuel-hungry freight trains capable of hauling incredible loads at immense cost, the new Flash series represents a fleet of nimble, hyper-efficient hybrid delivery vans. Efficiency gains ranging from 17% to 65% reduced tokens for strong results on third-party benchmarks Under the hood, Gemini 3.6 Flash achieves significant efficiency gains. The model reduces output token usage by 17% compared to its predecessor, Gemini 3.5 Flash, according to the Artificial Analysis Index maintained by the independent third-party AI benchmarking group of the same name. In specific long-horizon software engineering benchmarks like DeepSWE, which measures how well agents complete multi-step engineering tasks from scratch, the token savings reach up to 65%. This reduction means the model requires fewer reasoning steps and tool calls to complete the exact same multi-step workflow. Think of token efficiency like fuel economy in a vehicle. When an AI model takes a convoluted path to solve a problem, it burns through more computational fuel, driving up the final cost for the developer. By streamlining its internal logic, Gemini 3.6 Flash arrives at the correct answer faster and cheaper. While Google's materials did not specify the exact architectural or algorithmic changes used to achieve this token efficiency, they noted that the model "takes fewer reasoning steps and tool calls to accomplish multi-step workflows" and exhibits reduced "verbosity." The official model cards released by Google reveal that both Gemini 3.6 Flash and Gemini 3.5 Flash-Lite feature a 1-million-token input context window alongside a max output limit of 64,000 tokens, with both models sharing a knowledge cutoff date of March 2026. Respectable benchmark performance at low cost The technological improvements extend to concrete capabilities. Gemini 3.6 Flash scores 49% on the DeepSWE benchmark, a notable increase from the 37% achieved by version 3.5. It also pushes machine learning engineering performance higher, scoring 63.9% on MLE-Bench compared to 49.7% previously. Furthermore, Google integrates computer use as a built-in client-side tool via the Gemini API and Gemini Enterprise, reflecting an OSWorld-Verified score of 83.0%, up from 78.4%. The model also tackles knowledge work with greater proficiency, outperforming its predecessor on benchmarks like GDPval-AA v2 by moving from a score of 1349 to 1421. To ensure safety amidst these capability upgrades, Google deploys enhanced Frontier Safety safeguards. These protections harden the model against jailbreaks and mitigate risks in Chemical, Biological, Radiological, and Nuclear domains, as well as cyber offense misuses. The engineering team trains the model to minimize refusals for beneficial uses, striking a necessary balance between strict security and practical utility. Models for low-cost coding, agentic, and cybersecurity use cases -- respectively Google divided its new offerings into three distinct products tailored for different operational needs. Gemini 3.6 Flash serves as the heavy-duty workhorse of the trio. It handles complex coding, intricate knowledge work, and multimodal processing with improved precision. Enterprise customers utilize it for demanding tasks such as complex document parsing, intricate chart and data analysis, and long-form report drafting. The model executes complex code migrations using multi-agent orchestration frameworks with lower latency and higher quality than earlier iterations. Furthermore, 3.6 Flash aids in developing photographic texture extractors for 3D workflows using canvas interfaces. Gemini 3.5 Flash-Lite targets environments where high throughput and absolute minimal latency are non-negotiable. Google designates it as the fastest model in the 3.5 series. As measured by Artificial Analysis, the model processes 350 output tokens per second, making it highly effective for agentic search and massive document processing workloads. Artificial Analysis notes this is about twice as fast as prior generation model Gemini 3.1 Flash-Lite. Developers can configure 3.5 Flash-Lite to prioritize low-latency execution for high-volume tasks using minimal thinking levels, or engage higher thinking levels to process complex multi-step subagent workloads. Despite its lite designation, it outperforms the standard Gemini 3 Flash on several key agentic and coding evaluations, including SWE-Bench Pro, where it scores 54.2% compared to 49.6%, and OSWorld-Verified, scoring 74.0% versus 65.1%. The model extracts product features from massive datasets, generates interactive web design concepts, and scales receipt translation seamlessly. The third product, Gemini 3.5 Flash Cyber, represents a highly specialized deployment. Google fine-tuned this model specifically to find and fix cybersecurity vulnerabilities. It integrates directly with Google's CodeMender agent. In practice, multiple 3.5 Flash Cyber agents work concurrently to produce a single, comprehensive vulnerability report, achieving competitive performance at the frontier on the CyberGym benchmark. Google did not specify an exact numerical cost for 3.5 Flash Cyber, stating only that it is fine-tuned "at a lower price per token than larger models. Commercial licensing only The licensing framework for the new Gemini models carries profound implications for developers and enterprise users. Google deploys Gemini 3.6 Flash and 3.5 Flash-Lite under a commercial, proprietary API model. Unlike open-source software governed by licenses such as the MIT License or the GNU General Public License, developers do not gain access to the underlying model weights, training data, or source code. An MIT or GPL license grants users the freedom to download the codebase, modify the internal architecture, self-host the deployment, and distribute the software infrastructure independently. In contrast, Google's API approach means developers essentially rent access to the intelligence on a strict metered basis. Every prompt and generated response travels through Google's managed servers, incurring a cost based on the strict pricing structure of $1.50 per million input tokens for 3.6 Flash. This commercial tethering restricts deployment flexibility. Enterprises cannot air-gap the models entirely on their own local secure hardware without establishing specialized, high-tier enterprise agreements with Google Cloud. Developers remain bound by Google's acceptable use policies, arbitrary rate limits, and network requirements, creating a permanent dependency on Google's infrastructure uptime and terms of service. The licensing for Gemini 3.5 Flash Cyber proves even more restrictive. Acknowledging the dual-use nature of cybersecurity AI -- which attackers can weaponize just as easily as defenders can use it to patch systems -- Google is for now making the model only available behind a limited-access pilot program, similar to the trend kicked off by Anthropic's Mythos model with its Project Glasswing program, and continued by OpenAI with its staggered rollout for GPT-5.6. In this case, Google is making 3.5 Flash Cyber exclusively available to governments and trusted partners. This strict gatekeeping prevents open access, prioritizing systemic security over widespread developer innovation. Looking ahead Google DeepMind continues to iterate rapidly, but the gap in its product line remains apparent. While the Flash series excels in speed and economy, the industry eagerly awaits the deployment of Gemini 3.5 Pro to gauge Google's absolute frontier capabilities. Simultaneously, the company confirms that pre-training for Gemini 4 has already commenced. Until the next major flagship release materializes, developers must optimize their systems using the highly efficient, yet purposefully constrained, Flash architecture.
[22]
Google expands Gemini with cheaper models and a bug-hunter it keeps on a leash
Google LLC today launched three new Gemini Flash models and moved its CodeMender code-security agent into preview, part of a push to run artificial intelligence agents more cheaply and to automate the finding and fixing of software vulnerabilities. The new models are Gemini 3.6 Flash, an updated version of the workhorse model Google positions for coding and knowledge work; Gemini 3.5 Flash-Lite, built for high-throughput tasks; and Gemini 3.5 Flash Cyber, a security-tuned model the company is restricting to governments and trusted partners. Gemini 3.6 Flash is the centerpiece. Google said it uses up to 17% fewer output tokens than the previous 3.5 Flash on the Artificial Analysis Index and takes fewer reasoning steps and tool calls to complete multistep jobs. It's priced at $1.50 per million input tokens and $7.50 per million output tokens, below the cost of 3.5 Flash. The company reported gains across coding and knowledge benchmarks. It put the model at 49% on DeepSWE against 37% for 3.5 Flash, at 63.9% on the MLE Bench measure of machine-learning research against 49.7%, and at 83% on the OSWorld-Verified test of computer use against 78.4%. Computer use is now a built-in tool in the Gemini API and Gemini Enterprise. On GDPval-AA v2, a benchmark for knowledge work, Google scored the model at 1421 against 1349. Early customers include legal AI company Harvey AI Corp. and financial research platform Hebbia Inc., which Google said used the model for document parsing, data analysis and report drafting. "Gemini 3.6 Flash excels at document drafting and review in practice areas like capital markets and corporate M&A," Niko Grupen, head of applied research at Harvey, said in a testimonial provided by Google. "Compared to its predecessor, Gemini 3.6 Flash showed strong gains in performance on our benchmarks and was notably more efficient, completing tasks 12% faster on average." Google said 3.6 Flash ships with expanded Frontier Safety safeguards covering chemical, biological, radiological and nuclear risks and cyber offense and has been trained to resist jailbreaks while cutting refusals of benign requests. Gemini 3.5 Flash-Lite is the cheapest and fastest of the three. Priced at 30 cents per million input tokens and $2.50 per million output tokens, it delivers the highest throughput of the 3.5 series, according to Artificial Analysis. Google said it significantly outperforms the earlier 3.1 Flash-Lite on tests, including Terminal-Bench 2.1. On some coding and agentic evals such as SWE-Bench Pro and OSWorld-Verified, it beats the larger 3 Flash. The company is positioning it as a migration path for workloads running on its 2.5 and 3 Flash models and is rolling it out in Google Search. The third model is aimed squarely at security. Gemini 3.5 Flash Cyber is fine-tuned to find, validate and patch software vulnerabilities and runs inside CodeMender. Google said that when CodeMender calls the model up to five times to produce a single report, it reaches competitive performance against much larger models on the CyberGym benchmark. In testing by Google DeepMind's Big Sleep team on complex codebases such as Chrome and Safari, the company said Flash Cyber outperformed its own 3.5 Flash and 3.6 Flash models as well as Anthropic PBC's Claude Opus 4.6. On the V8 JavaScript engine, Google said the model found 55 unique confirmed issues, compared with 47 for 3.5 Flash and 36 for Claude Opus 4.6, including 10 that no other model caught. Google said it benchmarked against Claude Opus 4.6 rather than newer competitor models because those more recent releases perform worse at finding vulnerabilities, which it attributed to their safety guardrails. In a separate exercise, Google's Cloud Vulnerability Research team used the model to uncover remote code execution flaws in public application programming interface and a memory-corruption bug in a production service within two hours, then generated a working exploit that bypassed standard memory protections. Citing the dual-use nature of vulnerability research, Google said it will make Flash Cyber available only to governments and trusted partners through CodeMender under a limited-access pilot. CodeMender, the agent that runs the model, entered preview today. Built on Google DeepMind research, CodeMender works in three steps Google calls scan, verify and remediate. It first scans a repository for flaws such as memory corruption, injection and cryptographic weaknesses. It then tries to prove each one is real, building an exploit and running it in a sandbox the customer controls to weed out false positives. If the exploit succeeds, the agent writes a patch and returns it as a code diff for the developer to approve. It supports C/C++, Go, Java, Python, Ruby, Rust and TypeScript. Google is not tying customers to one model. They can pick whichever fits the cost, speed and scanning depth a job needs, and the company said it will add support for third-party frontier models later this year. It's available through the Gemini Enterprise Agent Platform and as a component of Google's AI Threat Defense, where the company's Wiz cloud-security platform enriches CodeMender's findings in the Wiz Security Graph and triggers automated penetration testing. Google said Wiz will also be able to call CodeMender to scan code, a capability it described as coming soon. CodeMender keeps a human in the loop. Developers approve patches before they are committed, though the agent can be wired into continuous integration pipelines for autonomous operation. Google said source code data is encrypted, isolated and not retained. Early enterprise testers include Salesforce Inc., Robinhood Markets Inc. and Palo Alto Networks Inc. Iain Mulholland, chief information security officer at Salesforce, said CodeMender "brings AI into a critical part of the security lifecycle by accelerating the path from validated vulnerability to tested fix." Scott Ponte, head of security operations at Robinhood, said the agent "consistently identified critical vulnerabilities that our other AI-enabled tools completely missed." Gemini 3.6 Flash and 3.5 Flash-Lite are available starting today through the Gemini API in Google AI Studio and Android Studio, with 3.6 Flash also in Google's Antigravity coding tool and the Gemini Enterprise app. Both reach consumers through the Gemini app. Google said Gemini 3.5 Pro is now testing with partners ahead of a broader release and that its most ambitious pretraining run yet is under way for Gemini 4.
[23]
Google's Gemini Flash 5.6 model cuts AI agent token costs by up to 65% on long horizon engineering tasks -- and 3.5 Pro is on the way
Google DeepMind today released three new proprietary AI models it says are among its most token-efficient yet: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. The models aim to make AI agents faster, smarter, and cheaper at scale. Google is pricing Gemini 3.6 Flash at $1.50 per one million input tokens and $7.50 per one million output tokens through its application programming interface (API), while Gemini 3.5 Flash-Lite costs a staggeringly cheap $0.30/$2.50 per million tokens in/out. Compare that to the $1.50/$9.00 per 1M tokens for Gemini 3.5 Flash, and the $2/$12 for Gemini 3.1 Pro Preview, and the savings are considerable. However, Google's prior generation Gemini 3.1 Flash-Lite still remains the search giant's "most cost-efficient" model at $0.25/$1.50 per 1M tokens. Yet, it remains 2X slower than the new, more expensive Gemini 3.5 Flash-Lite, giving those enterprises who value speed more "bang" for their buck. VB Frontier AI Model API Pricing Comparison Chart (Late July 2026 Shortlist) No price was provided yet for the specialty Gemini 3.5 Flash Cyber model, which, as its name would imply, is designed for cybersecurity researchers and red teamers to patch bugs. While the prices are among the middle-low end of all major AI models globally, the fact that Google designed them to use less tokens overall also should drive down costs for enterprises beyond what the sticker price shows (since you'll be paying for fewer total tokens at any rate). Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are available immediately through the Gemini API in Google AI Studio and Android Studio, as well as within the consumer Gemini application and Google Search. According to a separate Google blog post, Gemini 3.5 Flash Cyber will be available "exclusively available to governments and trusted partners via CodeMender soon" -- CodeMender being Google's proprietary AI code bug-fixing agent released last year. As with previous Gemini models, these are all proprietary and "closed source," thus, they can only be obtained through Google's official API and that of its partners, as opposed to an open-source license like MIT or Apache 2.0. One conspicuous omission noted by developers on X and social media: where is the larger, more powerful, flagship Gemini 3.5 Pro model Google previously alluded would be released this summer? After all, Gemini 3.1 Pro, the prior flagship, debuted back in February 2026, and rivals OpenAI and Anthropic have since released several more generations of flagship updates far more powerful than Google's. Google technical staffer Logan Kilpatrick responded to one such inquiry on X, writing: "Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it's ready." Google's release signals that the immediate future of AI lies in agentic capabilities -- systems that operate autonomously over extended periods. If early large language models are akin to massive, fuel-hungry freight trains capable of hauling incredible loads at immense cost, the new Flash series represents a fleet of nimble, hyper-efficient hybrid delivery vans. Efficiency gains ranging from 17% to 65% reduced tokens for strong results on third-party benchmarks Under the hood, Gemini 3.6 Flash achieves significant efficiency gains. The model reduces output token usage by 17% compared to its predecessor, Gemini 3.5 Flash, according to the Artificial Analysis Index maintained by the independent third-party AI benchmarking group of the same name. In specific long-horizon software engineering benchmarks like DeepSWE, which measures how well agents complete multi-step engineering tasks from scratch, the token savings reach up to 65%. This reduction means the model requires fewer reasoning steps and tool calls to complete the exact same multi-step workflow. Think of token efficiency like fuel economy in a vehicle. When an AI model takes a convoluted path to solve a problem, it burns through more computational fuel, driving up the final cost for the developer. By streamlining its internal logic, Gemini 3.6 Flash arrives at the correct answer faster and cheaper. While Google's materials did not specify the exact architectural or algorithmic changes used to achieve this token efficiency, they noted that the model "takes fewer reasoning steps and tool calls to accomplish multi-step workflows" and exhibits reduced "verbosity." The official model cards released by Google reveal that both Gemini 3.6 Flash and Gemini 3.5 Flash-Lite feature a 1-million-token input context window alongside a max output limit of 64,000 tokens, with both models sharing a knowledge cutoff date of March 2026. Respectable benchmark performance at low cost The technological improvements extend to concrete capabilities. Gemini 3.6 Flash scores 49% on the DeepSWE benchmark, a notable increase from the 37% achieved by version 3.5. It also pushes machine learning engineering performance higher, scoring 63.9% on MLE-Bench compared to 49.7% previously. Furthermore, Google integrates computer use as a built-in client-side tool via the Gemini API and Gemini Enterprise, reflecting an OSWorld-Verified score of 83.0%, up from 78.4%. The model also tackles knowledge work with greater proficiency, outperforming its predecessor on benchmarks like GDPval-AA v2 by moving from a score of 1349 to 1421. To ensure safety amidst these capability upgrades, Google deploys enhanced Frontier Safety safeguards. These protections harden the model against jailbreaks and mitigate risks in Chemical, Biological, Radiological, and Nuclear domains, as well as cyber offense misuses. The engineering team trains the model to minimize refusals for beneficial uses, striking a necessary balance between strict security and practical utility. Models for low-cost coding, agentic, and cybersecurity use cases -- respectively Google divided its new offerings into three distinct products tailored for different operational needs. Gemini 3.6 Flash serves as the heavy-duty workhorse of the trio. It handles complex coding, intricate knowledge work, and multimodal processing with improved precision. Enterprise customers utilize it for demanding tasks such as complex document parsing, intricate chart and data analysis, and long-form report drafting. The model executes complex code migrations using multi-agent orchestration frameworks with lower latency and higher quality than earlier iterations. Furthermore, 3.6 Flash aids in developing photographic texture extractors for 3D workflows using canvas interfaces. Gemini 3.5 Flash-Lite targets environments where high throughput and absolute minimal latency are non-negotiable. Google designates it as the fastest model in the 3.5 series. As measured by Artificial Analysis, the model processes 350 output tokens per second, making it highly effective for agentic search and massive document processing workloads. Artificial Analysis notes this is about twice as fast as prior generation model Gemini 3.1 Flash-Lite. Developers can configure 3.5 Flash-Lite to prioritize low-latency execution for high-volume tasks using minimal thinking levels, or engage higher thinking levels to process complex multi-step subagent workloads. Despite its lite designation, it outperforms the standard Gemini 3 Flash on several key agentic and coding evaluations, including SWE-Bench Pro, where it scores 54.2% compared to 49.6%, and OSWorld-Verified, scoring 74.0% versus 65.1%. The model extracts product features from massive datasets, generates interactive web design concepts, and scales receipt translation seamlessly. The third product, Gemini 3.5 Flash Cyber, represents a highly specialized deployment. Google fine-tuned this model specifically to find and fix cybersecurity vulnerabilities. It integrates directly with Google's CodeMender agent. In practice, multiple 3.5 Flash Cyber agents work concurrently to produce a single, comprehensive vulnerability report, achieving competitive performance at the frontier on the CyberGym benchmark. Google did not specify an exact numerical cost for 3.5 Flash Cyber, stating only that it is fine-tuned "at a lower price per token than larger models. Commercial licensing only The licensing framework for the new Gemini models carries profound implications for developers and enterprise users. Google deploys Gemini 3.6 Flash and 3.5 Flash-Lite under a commercial, proprietary API model. Unlike open-source software governed by licenses such as the MIT License or the GNU General Public License, developers do not gain access to the underlying model weights, training data, or source code. An MIT or GPL license grants users the freedom to download the codebase, modify the internal architecture, self-host the deployment, and distribute the software infrastructure independently. In contrast, Google's API approach means developers essentially rent access to the intelligence on a strict metered basis. Every prompt and generated response travels through Google's managed servers, incurring a cost based on the strict pricing structure of $1.50 per million input tokens for 3.6 Flash. This commercial tethering restricts deployment flexibility. Enterprises cannot air-gap the models entirely on their own local secure hardware without establishing specialized, high-tier enterprise agreements with Google Cloud. Developers remain bound by Google's acceptable use policies, arbitrary rate limits, and network requirements, creating a permanent dependency on Google's infrastructure uptime and terms of service. The licensing for Gemini 3.5 Flash Cyber proves even more restrictive. Acknowledging the dual-use nature of cybersecurity AI -- which attackers can weaponize just as easily as defenders can use it to patch systems -- Google is for now making the model only available behind a limited-access pilot program, similar to the trend kicked off by Anthropic's Mythos model with its Project Glasswing program, and continued by OpenAI with its staggered rollout for GPT-5.6. In this case, Google is making 3.5 Flash Cyber exclusively available to governments and trusted partners. This strict gatekeeping prevents open access, prioritizing systemic security over widespread developer innovation. Looking ahead Google DeepMind continues to iterate rapidly, but the gap in its product line remains apparent. While the Flash series excels in speed and economy, the industry eagerly awaits the deployment of Gemini 3.5 Pro to gauge Google's absolute frontier capabilities. Simultaneously, the company confirms that pre-training for Gemini 4 has already commenced. Until the next major flagship release materializes, developers must optimize their systems using the highly efficient, yet purposefully constrained, Flash architecture.
[24]
Google's cheaper, faster Gemini models just arrived
Google barely let its last Flash model settle in before replacing it. Gemini 3.5 Flash showed up in May as the company's speed-focused workhorse. Just two months later, Google is already moving on. The company just launched Gemini 3.6 Flash, along with two more models built for very different jobs. Gemini 3.6 Flash takes over as Google's main Flash model. The pitch here is efficiency. Google says the model uses about 17% fewer tokens than its predecessor. At the same time, it turns in stronger results on coding and general knowledge tasks. That efficiency shows up in the price too. Output tokens now cost $7.50 per million, down from $9 on the outgoing model. Gemini 3.6 Flash also handles computer use tasks better than before. Google's internal benchmark jumped from roughly 78% to 83%. Two more models round out the release The second model, Gemini 3.5 Flash-Lite, trades some capability for raw speed. It's built for high-volume jobs like translating text or sorting through documents. It runs at 350 output tokens per second. Input tokens cost just $0.30 per million. Gemini 3.6 Flash and Flash-Lite are both live now in the Gemini app. Flash-Lite is rolling out to Google Search as well. The third model comes with more restrictions attached. Gemini 3.5 Flash Cyber is built to find and patch security vulnerabilities in code. That's a useful skill, but it's also a risky one to hand out widely. Google is limiting access to governments and select partners for now. Developers can reach the other two models through Google Antigravity, AI Studio, and Android Studio. Google's flagship Gemini 3.5 Pro is still nowhere to be found. The model has been stuck behind an internal delay for months now. For now, the Flash lineup is doing the heavy lifting.
[25]
Google finally gives Gemini a cheaper AI model that actually outperforms competitors
Sara Heritage is a tech and gaming journalist, who's currently making her way up to Master Ball rank in Pokemon Champions. Bylines in IGN, GAMINGbible, The Gamer and more. You can usually find her tinkering with tech, or restoring old consoles, always with one of her 3 cats nearby. Come and talk with her over on Twitter @SHeritageJourno. * Google launched three Gemini Flash models focused on speed, efficiency and lower token costs. * Gemini 3.6 Flash , is cheaper per task, and boosts coding and research accuracy. * 3.5 Flash Cyber is government/partner-only for bug hunting; 3.5 Flash‑Lite is the fastest, cheapest for quick agent tasks. Alphabet just dropped three new Google Gemini AI models: 3.5 Flash Cyber, 3.5 Flash-Lite, and 3.6 Flash. These models are all about making Gemini faster, cheaper, and more efficient -- and come with the added bonus of saving you a little bit of cash on your token use. Neat! Breaking down each new Gemini Model What you need to know about Gemini 3.6 Flash, 3.5 Flash Cyber and 3.5 Flash-Lite Gemini 3.5 Flash Cyber is built to spot and fix software bugs, but for now, only governments and trusted partners get to play with it. It's cheaper per token than the bigger models, and could help Google catch up to Anthropic's Mythos in the cybersecurity race. Gemini 3.5 Flash-Lite is the fastest and cheapest of the 3.5 crew, designed to handle many quick tasks simultaneously within larger AI agent systems. But Gemini 3.6 Flash is the most efficient model, using up to 17% fewer tokens and costing less per token than before. Google says it's cheaper per task than GPT-5.6 Terra Max, Kimi K3, and Qwen 3.7 Max. It's better at coding, handling different types of data, and knowledge work than before, too. Finance teams can process documents faster, retailers can manage catalogs and customer support at lower cost, and healthcare researchers can crunch big datasets without blowing the budget. It's more precise too, making fewer mistakes in code edits (49% vs 37% on DeepSWE) and doing better in machine learning research (63.9% vs 49.7% on MLE Bench). Basically, it's just better at real-world coding and research tasks. If you didn't know, DeepSWE is a benchmark that evaluates the accuracy of AI models' code edits for software engineering tasks, while MLE Bench tests model performance on various machine learning research activities, such as code generation and problem-solving. So what's next for Gemini? Little has been officially confirmed about Google's next flagship Gemini model. Gemini 3.5 Pro is still in testing with partners, but should be out soon and available to all. Meanwhile, Google's already hard at work on Gemini 4, and they sound pretty hyped about it. We don't know much officially about Gemini 4, and it's allegedly been delayed from internal timelines. Early leaks say it'll be smarter, handle bigger chunks of info, and mix text, images, audio, and video all in one go. It should also be faster, more efficient, and more secure. Google is also allegedly working on a special chip to run Gemini up to 10 times more efficiently, all to make AI cheaper to use, which Gemini 4 is all-but-guaranteed to use.
[26]
Gemini 3.6 Flash, 3.5 Flash-Lite Launched With Focus on Faster, Cheaper Agents
Google also revealed that Gemini 4 is currently in pre-training Google on Wednesday expanded its Gemini family with the launch of Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. As per the Mountain View-based tech giant, its latest models are aimed at executing high-volume AI and agentic workloads more efficiently. Gemini 3.6 Flash is claimed to use 17 percent fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index. Meanwhile, Gemini 3.5 Flash-Lite is positioned as the fastest model in the 3.5 series, generating up to 350 output tokens per second. Gemini 3.6 Flash, Gemini 3.5 Flash-Lite Features In a blog post, Google announced that Gemini 3.6 Flash builds upon Gemini 3.5 Flash with improvements to coding, multimodal tasks, and knowledge work, while also requiring fewer output tokens. On the DeepSWE benchmark by Datacurve, the reduction is claimed to reach 65 percent. The model is also advertised to require fewer reasoning steps and tool calls when completing multi-step workflows. Google has priced Gemini 3.6 Flash at $1.50 (roughly Rs. 144) per one million input tokens and $7.50 (roughly Rs. 720) per one million output tokens. Comparing benchmarks, the company revealed that Gemini 3.6 Flash scored 49 percent on DeepSWE, compared to 37 percent for Gemini 3.5 Flash. On MLE Bench, the respective scores were 63.9 percent and 49.7 percent. The newer model is also claimed to have achieved 83 percent on OSWorld-Verified, compared to 78.4 percent for its predecessor. Users can access computer use as a built-in client-side tool through Gemini API and Gemini Enterprise. Further, Gemini 3.6 Flash can handle multimodal workloads such as document parsing, analysing charts and data, and drafting reports, as per the company. Google claims its new model ships with enhanced Frontier Safety safeguards against chemical, biological, radiological, and nuclear (CBRN), as well as cyber-offence misuse. Alongside Gemini 3.6 Flash, Google has introduced Gemini 3.5 Flash-Lite as well. It is claimed to be the fastest and cheapest model in the Gemini 3.5 family, designed for high-throughput and latency-sensitive workloads like agentic search and document processing. Google says Gemini 3.5 Flash-Lite costs $0.30 (roughly Rs. 29) per one million input tokens and $2.50 (roughly Rs. 240) per one million output tokens. With Gemini 3.5 Flash-Lite, developers can select different thinking levels depending on their workload. As per the company, minimal and low thinking levels will prioritise lower latency and cost, while higher thinking levels can be used for more complicated, multi-step sub-agent tasks. Apart from this, a specialised version of Gemini 3.5 Flash, called Gemini 3.5 Flash Cyber, has been introduced. The company says it is fine-tuned to identify, validate, and patch software vulnerabilities. The model works with Google's CodeMender security agent. Due to its specialised cybersecurity-focused nature, Gemini 3.5 Flash Cyber will not receive a general release. Instead, it will be offered exclusively to governments and trusted partners through CodeMender under a limited-access pilot programme. Google says Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are available starting today through the Gemini API via Google AI Studio and Android Studio, as well as the Gemini Enterprise Agent Platform. Gemini 3.6 Flash is also accessible through Google Antigravity and the Gemini Enterprise app. Both models are being made available through the Gemini app, while Gemini 3.5 Flash-Lite is also rolling out to Google Search. The company revealed that Gemini 3.5 Pro is currently in testing with partners, and a broader release is expected "once it is ready". Further, Google said that it has already begun its "most ambitious" pre-training run to date for Gemini 4, although a release timeline has yet to be announced.
[27]
Google Cuts Gemini 3.6 Flash Tokens for Coding Tasks
Google DeepMind's release of Gemini 3.6 Flash brings modest updates focused on improving speed and token efficiency, with particular attention to long-context processing and chart reasoning. These refinements aim to enhance specific technical capabilities but have left some users questioning the broader impact of the release. According to Universe of AI, the continued absence of Gemini 3.5 Pro, originally anticipated in mid-2026, highlights a significant gap in the product lineup, especially for users seeking more adaptable and versatile AI systems. Explore the practical benefits of Gemini 3.6 Flash, including its improvements in token optimization and handling extended datasets, while examining its limitations in areas like coding and complex problem-solving. Gain insight into how this release compares to competitors such as GPT 5.6 Luna and Grok 4.5, and understand the implications of the delayed Gemini 3.5 Pro for Google DeepMind's strategic direction. Key Features of Gemini 3.6 Flash and 3.5 Flash Light The Gemini 3.6 Flash model builds on the strengths of its predecessors, focusing on faster performance and enhanced token efficiency. It is particularly optimized for tasks requiring long-context processing and chart reasoning, areas where the Gemini series has historically performed well. Meanwhile, Gemini 3.5 Flash Light serves as a fine-tuned iteration of Gemini 3.5 Flash, offering minor adjustments rather than significant advancements. Both models are available through AI Studio and the Gemini API, making sure seamless integration for users already operating within the Gemini ecosystem. This accessibility underscores Google DeepMind's commitment to maintaining continuity for its existing user base. Performance Enhancements Gemini 3.6 Flash delivers measurable improvements over earlier models, particularly in technical and analytical tasks. Key performance highlights include: * Enhanced token efficiency, reducing token usage from 276 to 97 in deep software engineering tasks, which lowers computational costs and improves processing speed. * Improved long-context processing, allowing the model to handle complex datasets and extended inputs more effectively. * Consistent strengths in SVG generation and chart reasoning, reinforcing its utility in data visualization and structured reasoning tasks. While these enhancements make Gemini 3.6 Flash a practical tool for specific applications, they do not represent a significant leap forward in AI capabilities. The model remains focused on niche technical tasks rather than broader, innovative innovations. Advance your skills in Gemini Flash by reading more of our detailed content. Market Position and Competitive Challenges Despite its improvements, Gemini 3.6 Flash is positioned as a cost-effective and efficient option rather than a high-performance leader. It faces stiff competition from advanced models like GPT 5.6 Luna, Grok 4.5, and Cloud Sonnet, which excel in areas such as coding, software engineering and advanced problem-solving. This strategic focus on affordability and efficiency suggests that Google DeepMind is targeting a specific segment of the market. However, this approach may limit its appeal to users seeking state-of-the-art AI solutions capable of addressing a wider range of complex challenges. The lack of new advancements further underscores the model's limitations in competing with industry frontrunners. The Absence of Gemini 3.5 Pro The delay of Gemini 3.5 Pro, initially expected in June 2026, has created uncertainty among users and industry analysts. Without a clear timeline for its release, concerns have emerged about Google DeepMind's ability to deliver on its promises. The absence of this advanced model leaves a noticeable gap in the company's product lineup, particularly for users requiring high levels of intelligence, adaptability and versatility. This delay not only impacts user confidence but also raises broader questions about Google DeepMind's capacity to innovate at the pace required to remain competitive in the rapidly evolving AI landscape. Strengths and Limitations Gemini 3.6 Flash offers several notable strengths, including: * Strong performance in long-context tasks, making it suitable for handling extended datasets and complex reasoning. * Improved token efficiency, which reduces computational costs and enhances processing speed. However, these strengths are offset by significant limitations: * Inability to compete with leading AI models in intelligence, adaptability and advanced problem-solving. * Minimal advancements in coding and software engineering tasks, areas where competitors have made significant strides. * Lack of new innovation, which limits its appeal to users seeking innovative AI capabilities. These trade-offs position Gemini 3.6 Flash as a practical choice for specific technical use cases but restrict its broader market appeal. Availability and User Integration Both Gemini 3.6 Flash and 3.5 Flash Light are readily accessible through AI Studio and the Gemini API, making sure ease of integration for existing users of the Gemini ecosystem. This continuity allows users to incorporate the new models into their workflows without significant disruptions. However, the incremental nature of these updates may not be compelling enough to attract new users or those seeking more advanced AI solutions. Looking Ahead Gemini 3.6 Flash represents a modest step forward in AI development, offering improvements in speed and token efficiency. However, its lack of new innovation and the delay of Gemini 3.5 Pro highlight challenges in Google DeepMind's ability to maintain its competitive edge. While the new models provide value for specific technical tasks, they fall short of meeting the broader expectations of an increasingly demanding AI market. As the industry continues to evolve, Google DeepMind will need to address these challenges to remain a relevant and competitive player in the field. Media Credit: Universe of AI Disclosure: Some of our articles include affiliate links. If you buy something through one of these links, Geeky Gadgets may earn an affiliate commission. Learn about our Disclosure Policy.
[28]
Google Just Released Gemini 3.6 Flash, And It Might Be Its Worst Model To-Date
Google could not have picked a worse time to deliver a chronically underpowered Gemini 3.6 Flash, coming right on the heels of stunning open-source AI models from China, such as Moonshot's Kimi K3, which employs novel architectural feats to deliver notable cost savings, all the while offering frontier-level performance. Google's Gemini 3.6 Flash performs worse than Meta Spark 1.1, GLM-5.2, GPT-5.6 Luna, Sonnet 5, Grok 4.5, and GPT-5.6 Terra Google has just released the all-new Gemini 3.6 Flash, sporting a context window of 1 million tokens, which is equivalent to 1,500 A4 pages of text, 30,000 lines of code, or an hour of video, and priced at $7.5 per 1 million tokens of output. It also supports text, image, speech, and video inputs. Critically, according to Google, the model consumes 17 percent fewer output tokens across multi-step workflows. This brings us to the core of today's topic. Google's Gemini 3.6 Flash has achieved a score of 50 on the the Artificial Analysis Intelligence Index, which measures agentic, general, coding, and scientific reasoning performance of AI models. In fact, this score is worse than what Meta Spark 1.1, GLM-5.2, GPT-5.6 Luna, Sonnet 5, Grok 4.5, and GPT-5.6 Terra have earned! A deeper dive reveals that Google's Gemini 3.6 Flash is often superceded by cheaper and older models on most benchmarks. For instance, on SWE-Bench Pro (Public), it is beaten by Grok 4.5, which debuted on July 08, and by Claude Sonnet 5 on MLE Bench. In fact, Gemini 3.6 Flash might be the most underwhelming model that Google has released to-date! Of course, Google has also released Gemini 3.5 Flash-Lite, which is fairly cost effective, managing to deliver 350 output tokens per second, and Gemini 3.5 Flash Cyber, which is geared towards cyber security applications. Follow Wccftech on Google to get more of our news coverage in your feeds.
[29]
Google Gemini Launch Delayed as Tech Falls Short of Internal Goals | PYMNTS.com
The company was widely expected to release 3.5 Pro at its developer conference in May, but it is still working to improve the model's capabilities, especially in coding, according to the report. Google said in a May 19 blog post announcing the launch of Gemini 3.5 Flash: "We're also hard at work on 3.5 Pro. It's already being used internally, and we look forward to rolling it out next month." According to the Bloomberg report, the delay has been caused in part by Google's many layers of stakeholders involved in preparing models for release, the company's efforts to make the 3.5 Pro's skills in writing code more competitive with its rivals, and competing factions within Google each building their own AI coding tools. Asked about the report by Bloomberg, a Google spokesperson said, per the report: "We're shipping quickly across a wide range of models while keeping them highly cost-effective for customers." Google is also working with the U.S. government and its efforts to monitor the most advanced models, according to the report. "We're currently testing 3.5 Pro, an upgraded Flash model, and other models with partners, and we're productively engaged with the U.S. government on model testing and broader frameworks," the Google spokesperson said, per the report. PYMNTS reported in May that Gemini 3.5 Flash had become the default model across the Gemini app and Search's AI Mode; that the Gemini app was serving more than 900 million monthly users across 230 countries; and that daily queries had grown sevenfold. In remarks delivered at a Google event in May, Google CEO Sundar Pichai said: "Today we have 13 products with over a billion users each. Five of those have more than 3 billion users. Our Gemini models are a big reason more people are using our products, and why they're using our products more." Speaking of the company's latest AI models, Pichai said: "Gemini 3.5 Flash is available for everyone today across our products and APIs. We're also excited for Gemini 3.5 Pro. We are using it internally, it's showing great improvements, and it will be coming next month." For all PYMNTS AI and digital transformation coverage, subscribe to the daily AI and Digital Transformation Newsletters.
[30]
Google Releases Three New Gemini Models, Said to be Working on New AI Chip
These changes have come up after a few months during which time rivals OpenAI and Anthropic launched their "powerful" models Google DeepMind released three new Gemini models - 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber - that promises improved capabilities in coding, knowledge work, multimodal performance while also reducing token usage by up to 17%. In parallel, there were also reports of Google working on a new AI chip designed to make Gemini more efficient and potent. "Our newest Gemini models deliver the efficiency, latency, and reliability to build AI agents at scale," says senior director of product management Tulsee Doshi in a blog post. The models are cheaper, faster and optimised for coding, efficiency and cybersecurity. However, the latest update does not include Google's flagship Gemini Pro model, last updated in February. Beyond today's releases, Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it's ready. In parallel, our team is already focusing on building the next generation of models. We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress, the blog said. While Gemini 3.5 Flash-Lite is the most cost-effective model in the class, the 3.5 Flash Cyber is a specialised model that has been fine-tuned for finding and fixing cybersecurity vulnerabilities at a decent price point. The model is being shared exclusively for governments and trusted partners as part of a limited access program. Meanwhile, a report published by The Information, said Alphabet is designing a new server chip to help Gemini models operate more efficiently. The new chip, named "Frozen v2" should be out some time in 2028 and is claimed to be between six to 10 times more efficient than Google's existing AI chips. In fact, the company neither confirmed nor denied the report. They merely said that internal teams were continuously researching and experimenting with new innovations to deliver maximum performance and efficiency for users and customers. Google did acknowledge that by co-designing hardware and software from the ground up, they ensure integration and a high level of optimisation for real-world workouts. In recent times AI companies have sought to produce their own processors in order to make their in-house models run more efficiently. So, we have cases where Nvidia is the main chipmaker for the top brands but that is not prevent companies like OpenAI to announce its first custom chip for inference dubbed as Jalapeno. Reports of Anthropic working with Samsung also emerged. Similarly with the new Gemini models, Google has been off the radar for some time during which OpenAI released both GPT-5.5 and GPT-5.6 Sol while Anthropic came out with Claude Opus 4.8 and Claude Sonnet 5 while also expanding access to its frontier Fable 5 model. Gemini Pro models have been dubbed Google's highest-capability offerings for complex reasoning and coding. Meanwhile, Google DeepMind product lead Logan Kilpatrick took to his X handle to suggest that the company was also testing Gemini 3.5 with partners and hopes to land it soon. He also added that the team has already started its ambitious pre-training for Gemini 4.
[31]
Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite; Next-Gen Gemini 4 Enters Training
Google has expanded its Gemini family with the launch of Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, introducing two AI models designed to deliver faster performance, lower latency, and reduced operating costs for developers and enterprises. Alongside these releases, the company also unveiled Gemini 3.5 Flash Cyber, a specialized model focused on cybersecurity applications. The announcement reflects Google's strategy of making AI models more efficient for real-world workloads while its flagship Gemini 3.5 Pro remains under development.
[32]
Google launches three new Gemini AI models, expands push into cybersecurity By Investing.com
Investing.com -- oogle on Tuesday unveiled three new artificial intelligence models, including its most capable Flash model to date and a cybersecurity-focused system, as the tech giant steps up competition with OpenAI and Anthropic in the rapidly evolving AI market. The announcements come just ahead of Alphabet's quarterly results, as competition in generative AI intensifies. Chinese developers have also been gaining ground, with Moonshot AI temporarily restricting new Kimi K3 subscriptions and API access after surging demand strained computing capacity. Meanwhile, Alibaba has signaled that its upcoming Qwen 3.8 Max model is expected to rank among the industry's top-performing AI systems. The new lineup includes Gemini 3.6 Flash, which Google said delivers stronger performance than its predecessor across tasks such as coding and financial analysis while lowering costs for users. The company also introduced Gemini 3.5 Flash Cyber, a model optimized to identify and patch software security vulnerabilities more efficiently, and Gemini 3.5 Flash-Lite, designed to power lightweight AI agents and autonomous digital assistants. The releases underscore Google's growing focus on cybersecurity, an area where rivals OpenAI and Anthropic have recently introduced advanced AI systems capable of detecting software vulnerabilities and assisting security researchers. Google said the new Flash models are designed to balance performance, speed and cost, aiming to offer developers high-quality AI capabilities without the expense of larger models. Meanwhile, Google's flagship Gemini 3.5 Pro model remains in testing. The company had previously indicated it would launch in June but said it will be made broadly available once it is ready.
[33]
Google announces three new Gemini AI models as competition with OpenAI and Anthropic heats up
The long-awaited Gemini 3.5 Pro update is still missing despite growing competition from OpenAI's GPT-5 and Anthropic's latest Claude models. Google has unveiled three new additions to the Gemini AI models- Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber. These three models are aimed at developers building AI powered applications and autonomous agents. Google claims that these models are faster flagship models, a budget-friendly variant and a cybersecurity-focused offering. Among all the releases, the Gemini 3.6 Flash is the highlight. Google describes it as its primary production model for everyday AI workloads. The company says that the model delivers stronger performance in coding, reasoning, knowledge-intensive tasks and multimodal capabilities while consuming fewer output tokens than its predecessor. As per Google, the model reduces the overall token usage by around 17 percent, with some workloads seeing output token savings of up to 65 percent. The lower token consumption is expected to reduce operating costs for developers using the model at scale. Along with this, Google has also introduced the Gemini 3.5 Flash Lite. The company says that it is the fastest and most affordable AI model. It is designed for lightweight workloads where response speed and lower costs are a priority over advanced reasoning. Also read: Loved Assassin's Creed Resynced? Here are the games you should play next The third one is the Gemini 3.5 Flash Cyber. It is a specialised model which is developed for cybersecurity applications. Unlike two releases, this version will not be widely available. Google stated that it will initially be offered only to government agencies and selected partners through a limited-access pilot programme. The model is designed to identify software vulnerabilities and assist in generating security fixes as part of Google's AI-powered CodeMender initiative. Interestingly, the company did not announce the much anticipated Gemini 3.5 Pro update. The company stated that it was falling short of internal goals specifically in terms of coding, which has emerged as a lucrative use case for enterprise AI. On the other hand, OpenAI has already expanded its GPT-5 lineup, while Anthropic has introduced newer Claude models with enhanced reasoning and coding capabilities. Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are now rolling out to enterprise customers and Gemini app users, with Flash-Lite also expected to power parts of Google Search in the future.
Share
Copy Link
Google DeepMind unveiled three new Gemini models including the efficiency-focused Gemini 3.6 Flash, which uses 17% fewer tokens and costs less than its predecessor. The company also introduced Gemini 3.5 Flash Cyber for cybersecurity and Flash-Lite for high-volume tasks. However, the flagship Gemini 3.5 Pro remains delayed despite promises of a June release, while Google has already begun pre-training for Gemini 4.
Google DeepMind released three new Gemini models on Tuesday, marking a strategic shift toward cost-efficient AI models even as its flagship offering remains conspicuously absent
1
. The star of the announcement is Gemini 3.6 Flash, which Google positions as its "workhorse model" designed to deliver improved coding capabilities, knowledge work performance, and multimodal tasks while using up to 17% fewer tokens than its predecessor2
. This token efficiency translates directly to lower costs for developers, with pricing set at $1.50 per million input tokens and $7.50 per million output tokens, down from $1.50 and $9 respectively for the now-deprecated 3.5 Flash1
.
Source: Analytics Insight
The new Gemini models were built in direct response to user feedback about the 3.5 release, particularly around code generation performance that didn't meet Google's initial promises
1
. In the DeepSWE test for coding, Gemini 3.6 Flash scores 49% compared to 37% for 3.5 Flash, while the OSWorld test for computer use shows improvement to 83% from 78.4%1
. The model now supports computer use as a standard feature in the Gemini API, enabling more sophisticated agentic systems that can complete tasks more accurately and in fewer steps3
.Alongside the flagship update, Google introduced Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber, each targeting distinct use cases within Google's AI strategy
2
. Flash-Lite achieves an impressive 350 output tokens per second, making it Google's fastest and most cost-effective model at $0.30 per million input tokens and $2.50 per million output tokens. This positions it for high-volume workloads and smaller tasks within larger AI-agent systems, with Google already deploying it extensively in Google Search for AI Overviews1
.Gemini 3.5 Flash Cyber represents Google's first LLM tuned specifically for cybersecurity, designed to detect and patch software vulnerabilities
5
. Google claims this specialized model performs nearly as well at finding and fixing cybersecurity issues as the much larger and more expensive Claude Mythos from Anthropic, while maintaining the efficiency of a Flash model1
. The model is already finding and fixing bugs in Google's internal codebases for Android, Chrome, and YouTube3
. However, acknowledging the dual-use nature of vulnerability detection tools, Google is restricting access through a limited pilot program exclusively available to governments and trusted partners via Google DeepMind's CodeMender agent2
.
Source: VentureBeat
The most notable aspect of Tuesday's announcement is what Google didn't release. Gemini 3.5 Pro, the company's flagship reasoning model promised for June at the I/O conference in May, remains in testing with unnamed partners
1
. Google DeepMind product lead Logan Kilpatrick stated the company hopes to "land soon" but provided no concrete timeline2
. Bloomberg reported that Google is facing internal delays as it struggles to meet performance goals, particularly in coding where the model reportedly falls short compared to offerings from OpenAI and Anthropic4
.
Source: 9to5Google
This delay comes at a critical moment as competitors accelerate their release cycles. Since Google last updated its Pro model in February, OpenAI has released GPT-5.5 and begun rolling out GPT-5.6, while Anthropic has launched Claude Opus 4.8, Claude Sonnet 5, and expanded access to its frontier Fable 5 model
2
. Chinese rivals are also gaining momentum, with Moonshot AI's Kimi K3 drawing enough demand to limit new subscriptions due to capacity constraints, and Alibaba teasing Qwen 3.8 Max5
.Related Stories
During Alphabet's quarterly earnings call, CEO Sundar Pichai addressed concerns about the Gemini 3.5 Pro delay by pivoting to future plans, announcing that Google has already begun its most ambitious pre-training run yet for Gemini 4
4
. Pichai also outlined plans to release subsequent AI models at an almost monthly cadence, suggesting a fundamental shift in Google's development and deployment strategy4
.Google's emphasis on price and efficiency may help offset slower timing in key product categories. Artificial Analysis data shows Gemini Flash already undercuts comparable models from Anthropic, OpenAI, and Chinese rivals on cost
5
. The company is also reportedly developing a specialized chip designed to run Google Gemini up to 10 times more efficiently, part of a broader push to lower serving costs through its custom chips, cloud infrastructure, and integrated hardware-software design5
. All new models are available starting today for developers through the Gemini API via Google AI Studio and Android Studio, with Gemini 3.6 Flash also available in the Gemini app3
.Summarized by
Navi
[1]