7 Sources
[1]
Google plans new chip to run Gemini models more efficiently, the Information reports
July 20 (Reuters) - Google is developing a new server chip that would incorporate elements of its Gemini model directly into the hardware, in a bid to serve its AI models more efficiently to users, the Information reported on Monday, citing people familiar with the matter. The Alphabet-owned
[2]
Alphabet stock pops on report it's developing a more efficient AI chip
Alphabet shares climbed 3% on Monday after The Information reported the company is developing a new server chip, internally dubbed "Frozen v2," designed to run Gemini models more efficiently. The chip would permanently embed parts of Gemini's architecture directly into the silicon, reducing the
[3]
Google Frozen chip: Gemini baked into the silicon
A new report says Google is building a server chip with Gemini's design hardwired into the silicon. If the Google Frozen chip works, it could run the model up to 10 times more efficiently, and reshape the economics of AI. Most AI chips are general-purpose. You load a model onto them, and they run
[4]
Google reportedly developing 'Frozen v2' AI chip optimized for Gemini models
Google LLC is reportedly developing a new chip optimized to run its Gemini series of artificial intelligence models. Sources told The Information today that the processor is codenamed "Frozen v2". According to the publication, it's expected to provide between six and ten times better performance
[5]
Google plans new chip to run Gemini models more efficiently: Report
The Alphabet-owned company expects the new chip, informally dubbed "Frozen v2," to help address an AI computing capacity crunch that has fueled internal tensions and prompted Google Cloud to decline deals with outside customers, the report said. Google is developing a new server chip that would
[6]
Google Developing 'Frozen v2' AI Server Chip to Boost Gemini Efficiency
Google is reportedly developing a next-generation custom AI server chip, internally codenamed "Frozen v2," as part of its long-term strategy to improve the efficiency and performance of its Gemini artificial intelligence models. According to media reports, the chip is expected to enter production
[7]
Google reportedly working on a new AI chip, here is how it may improve Gemini models
The chip may be between six and 10 times more efficient than Google's existing AI chips. Google's parent company, Alphabet, is reportedly working on a new AI chip, internally called Frozen v2, that could make Gemini AI models faster and more power-efficient. The chip is expected to launch in 2028.
Share
Copy Link
Google is building a specialized AI chip called Frozen v2 that embeds Gemini's architecture directly into hardware. The chip could deliver 6-10 times better efficiency than current TPUs, addressing severe compute shortages that forced Google Cloud to turn away customers. Deployment targets 2028, but the approach trades flexibility for performance.
Google is developing a new server chip that permanently embeds elements of its Gemini model directly into the hardware, according to a report from The Information citing people familiar with the matter
1
. The project, informally dubbed Frozen v2, represents a shift in how the company approaches custom AI silicon and aims to address a severe AI computing capacity crunch that has fueled internal tensions and forced Google Cloud to decline deals with outside customers2
.
Source: ET
Alphabet shares climbed 3% following the news, reflecting investor confidence in the strategic move
5
. The timing underscores the urgency of Google's compute shortage, which has become severe enough that the company agreed to pay SpaceX nearly $1 billion a month to help bridge the gap and meet enterprise commitments2
.Unlike conventional chips that load models into memory and shuttle data back and forth, the AI chip optimized for Gemini would bake the neural-network structure straight into the circuitry
3
. This approach locks the hardware to the shape of Google's current AI design, with engineers still able to refresh model weights while the underlying architecture stays frozen3
.Google engineers project the chip could serve between six and ten times more tokens per unit of power than the company's newest TPU chips, or tensor processing units
2
. The efficiency gains would come from multiple sources: reducing the number of calculations required to run Gemini models and decreasing data movement between memory and processing cores4
.The Frozen project is aimed at creating a new line of custom AI hardware apart from Google's existing TPUs rather than replacing them entirely
5
. This specialized branch would complement Google's general-purpose chip portfolio, with the company targeting 2028 for data center deployment, though engineers are still finalizing the design and the amount of model information that will be hardwired2
.
Source: SiliconANGLE
The move would deepen Google's self-reliance in silicon, building on efforts that already include designing TPUs to reduce dependence on Nvidia and spreading chip orders across multiple suppliers
3
. For real-time applications like voice assistants where latency matters most, a chip with fixed design can answer queries with minimal delay3
.Related Stories
The primary risk lies in rigidity. AI moves rapidly, and a chip built around today's Gemini architecture could look outdated by 2028
3
. The chip would work with future Gemini models only if Alphabet sticks with the same underlying neural-network structure, according to The Information2
. Google reportedly views Frozen v2 partly as a trial run and does not plan to produce it at the same scale as its TPUs2
.Google is not alone in pursuing this approach. Startup Taalas is already selling chips that print model weights and architecture directly onto silicon in a product called Hardcore
3
. Taalas claims its part serves up to 17,000 tokens per second, compared to roughly 150 per user on a top Nvidia GPU, while eliminating the need for expensive high-bandwidth memory.The race for custom AI silicon is shifting from running any model to fusing one model with the hardware itself
3
. For companies serving one model to billions of users, trading flexibility for speed, cost reduction, and lower power consumption can make strategic sense. Google has not officially confirmed the project, with a spokesperson noting only that teams experiment with high-efficiency ideas and that not every lab project reaches production3
.Summarized by
Navi
[3]
1
Technology

2
Policy and Regulation

3
Technology
