5 Sources
[1]
Google plans new chip to run Gemini models more efficiently, the Information reports
July 20 (Reuters) - Google is developing a new server chip that would incorporate elements of its Gemini model directly into the hardware, in a bid to serve its AI models more efficiently to users, the Information reported on Monday, citing people familiar with the matter. The Alphabet-owned (GOOGL.O), opens new tab company expects the new chip, informally dubbed "Frozen v2," to help address an AI computing capacity crunch that has fueled internal tensions and prompted Google Cloud to decline deals with outside customers, the report said. Shares of Alphabet were up 3% in early trading. Here are some details: Reporting by Rashika Singh in Bengaluru; Editing by Leroy Leo Our Standards: The Thomson Reuters Trust Principles., opens new tab
[2]
Alphabet stock pops on report it's developing a more efficient AI chip
Alphabet shares climbed 3% on Monday after The Information reported the company is developing a new server chip, internally dubbed "Frozen v2," designed to run Gemini models more efficiently. The chip would permanently embed parts of Gemini's architecture directly into the silicon, reducing the number of calculations and amount of data movement required to answer queries, according to the news outlet. Google engineers project it could serve between six and ten times more tokens per unit of power than the company's newest AI chips, called TPUs, or tensor processing units, The Information said. Frozen would become a more specialized branch of Google's custom-chip portfolio rather than replace its general-purpose TPUs. According to the report, the company is targeting 2028 for deployment. The project is aimed at easing a major internal compute shortage that has fueled tensions and reportedly forced Google Cloud to turn away outside business. Just last month, Google agreed to pay SpaceX nearly $1 billion a month to help bridge the gap and meet its enterprise compute commitments. The trade-off is flexibility. The chip would work with future Gemini models only if Google sticks with the same underlying architecture, according to The Information. Google reportedly currently views Frozen v2 partly as a trial run and does not plan to produce it at the same scale as its TPUs. Alphabet did not immediately respond to a request for comment. Read the full story from The Information here.
[3]
Google Frozen chip: Gemini baked into the silicon
A new report says Google is building a server chip with Gemini's design hardwired into the silicon. If the Google Frozen chip works, it could run the model up to 10 times more efficiently, and reshape the economics of AI. Most AI chips are general-purpose. You load a model onto them, and they run it. Google is reportedly trying something stranger: a chip that is the model, with Gemini's blueprint etched into the hardware itself. The project, informally called "Frozen v2," was reported by The Information and picked up by Reuters and Bloomberg Law. Alphabet shares rose as much as 3.7% on the news. Google has not confirmed the project, and the chip is years away. But the idea behind it is a serious bet on where AI infrastructure goes next. What 'frozen' means Today's chips keep the model in memory and shuttle its data back and forth. That flexibility costs power and time. Frozen v2 would bake Gemini's neural-network architecture straight into the circuitry. The hardware locks to the shape of Google's current AI design. Engineers can still refresh the model by loading new weights, but the underlying structure stays fixed, or "frozen." How much of the model gets hardwired is reportedly still being decided. The payoff is efficiency. The Information reports the chip could be 6 to 10 times more efficient than Google's latest custom AI chips, measured by tokens served per unit of power. It would be a new line of silicon, separate from Google's TPUs rather than a replacement. Deployment is targeted for as early as 2028. Why it matters The timing is not random. The Information says Frozen v2 is partly a response to an AI capacity crunch inside Google. The squeeze is severe enough that Google Cloud has turned away some outside customers, and it has stirred internal tensions. Efficiency is the whole game right now. Running AI models is fabulously expensive, and every watt saved at data-centre scale is money. A chip tuned for one model can drop a lot of the overhead a general chip carries. It would also be fast. Because the design is fixed, the chip can answer with very little delay. As one observer noted, that suits real-time uses such as voice assistants, where lag is the enemy. There is a strategic angle too. Google already designs its own TPUs to lower its reliance on Nvidia. A Gemini-specific chip would push that self-reliance further, deepening an effort that has also seen Google spread its chip orders across suppliers. Google is not alone The approach is not unique to Google. A startup called Taalas is already selling the idea, printing a model's weights and architecture directly onto a chip it calls Hardcore. The claimed numbers are eye-catching. Taalas says its part serves up to 17,000 tokens a second, against roughly 150 per user on a top Nvidia GPU. It says it needs no expensive high-bandwidth memory, which would ease the memory crunch squeezing the industry. That is the wider bet. If a model can live in silicon, you trade flexibility for speed, cost and power. For a company serving one model to billions of people, that trade can make sense. It is the same instinct behind efforts to shrink models onto phones. The catch The obvious risk is rigidity. AI moves fast, and a chip built around today's Gemini could look dated by 2028. Google's design tries to soften that by keeping the weights updatable, but the architecture is still set in advance. Then there is the small matter of confirmation. Google has not acknowledged the project. A spokesperson said only that its teams experiment with high-efficiency ideas, and that not every lab project reaches production. So treat Frozen v2 as a signal, not a shipping product. The signal is clear enough. The race for custom AI silicon is moving from running any model to fusing one model with the metal. If Google pulls it off, its rivals will have to answer.
[4]
Google reportedly developing 'Frozen v2' AI chip optimized for Gemini models
Google LLC is reportedly developing a new chip optimized to run its Gemini series of artificial intelligence models. Sources told The Information today that the processor is codenamed "Frozen v2". According to the publication, it's expected to provide between six and ten times better performance per watt than the search giant's current silicon. Shares of Google parent Alphabet Inc. rose 1.5% on the report. Off-the-shelf chips often contain components that customers don't need. If an AI startup buys a graphics card that includes both inference and rendering cores, it may end up leaving the latter modules idle. That can lead to inefficiencies in AI projects. Designing a custom chip addresses the challenge. A company can leave out circuits that aren't needed for its workloads and thereby lower manufacturing costs. Alternatively, it can replace the unnecessary circuits with cores optimized for its use case. Google has long offered custom AI chips called TPUs through its public cloud. The two newest entries in the lineup, the TPU 8t and TPUi, are optimized for training and inference, respectively. According to today's report, Google plans to take that customization a step further by optimizing its upcoming Frozen v2 chip for its Gemini models' architecture. The significant efficiency gains that the company reportedly expects to unlock will be achieved through multiple routes. In particular, the company hopes that Frozen v2 will reduce the number of calculations needed to run Gemini. It will reportedly also decrease data movement. Data movement is an issue when a neural network can't fit into a graphics card's onboard memory. When that happens, neural network elements are stored in off-chip storage. The graphics card must regularly move those elements from the off-chip storage to its logic circuits and vice versa, which slows down processing. Google could address that bottleneck by equipping Frozen v2 with enough memory to run Gemini fully on-chip. Such a design would remove the need to move data to and from off-chip RAM. The fact that Frozen v2 is expected to reduce the number of calculations needed to run Gemini suggests that it will also support a form of operator fusion. That's a widely used technique for speeding up AI models. It works by combining several calculations into a single computation that can be completed faster. Google runs its TPUs in clusters that contain multiple custom components. The TPU 8i, for example, relies on custom devices called optical circuit switches to manage the flow of data. Google will presumably make Frozen v2 compatible with those components to avoid the need for major cluster redesigns. The company reportedly hopes to start rolling out the chip to its data centers in 2028.
[5]
Google plans new chip to run Gemini models more efficiently: Report
The Alphabet-owned company expects the new chip, informally dubbed "Frozen v2," to help address an AI computing capacity crunch that has fueled internal tensions and prompted Google Cloud to decline deals with outside customers, the report said. Google is developing a new server chip that would incorporate elements of its Gemini model directly into the hardware, in a bid to serve its AI models more efficiently to users, the Information reported on Monday, citing people familiar with the matter. The Alphabet-owned company expects the new chip, informally dubbed "Frozen v2," to help address an AI computing capacity crunch that has fueled internal tensions and prompted Google Cloud to decline deals with outside customers, the report said. Shares of Alphabet were up 3% in early trading. Here are some details: Google plans to deploy the chip as soon as 2028, though engineers are still finalizing its design and the amount of model information that will be hardwired, the report said. The chip could be six to 10 times more efficient than Google's latest custom AI chips based on the number of AI tokens served per unit of power, according to the report. The 'Frozen' project is aimed at creating a new set of homegrown chips apart from Google's tensor processing units (TPUs), rather than to replace them, the report said. Google did not immediately respond to a Reuters request for comment. Bloomberg News reported last week that Google delayed the launch of its latest Gemini AI model after it fell short of internal goals, with the company working to improve its capabilities, particularly in coding.
Share
Copy Link
Google is building a specialized AI chip called Frozen v2 that embeds Gemini's architecture directly into hardware. The chip could deliver 6-10 times better efficiency than current TPUs, addressing severe compute shortages that forced Google Cloud to turn away customers. Deployment targets 2028, but the approach trades flexibility for performance.
Google is developing a new server chip that permanently embeds elements of its Gemini model directly into the hardware, according to a report from The Information citing people familiar with the matter
1
. The project, informally dubbed Frozen v2, represents a shift in how the company approaches custom AI silicon and aims to address a severe AI computing capacity crunch that has fueled internal tensions and forced Google Cloud to decline deals with outside customers2
.
Source: ET
Alphabet shares climbed 3% following the news, reflecting investor confidence in the strategic move
5
. The timing underscores the urgency of Google's compute shortage, which has become severe enough that the company agreed to pay SpaceX nearly $1 billion a month to help bridge the gap and meet enterprise commitments2
.Unlike conventional chips that load models into memory and shuttle data back and forth, the AI chip optimized for Gemini would bake the neural-network structure straight into the circuitry
3
. This approach locks the hardware to the shape of Google's current AI design, with engineers still able to refresh model weights while the underlying architecture stays frozen3
.Google engineers project the chip could serve between six and ten times more tokens per unit of power than the company's newest TPU chips, or tensor processing units
2
. The efficiency gains would come from multiple sources: reducing the number of calculations required to run Gemini models and decreasing data movement between memory and processing cores4
.The Frozen project is aimed at creating a new line of custom AI hardware apart from Google's existing TPUs rather than replacing them entirely
5
. This specialized branch would complement Google's general-purpose chip portfolio, with the company targeting 2028 for data center deployment, though engineers are still finalizing the design and the amount of model information that will be hardwired2
.
Source: SiliconANGLE
The move would deepen Google's self-reliance in silicon, building on efforts that already include designing TPUs to reduce dependence on Nvidia and spreading chip orders across multiple suppliers
3
. For real-time applications like voice assistants where latency matters most, a chip with fixed design can answer queries with minimal delay3
.Related Stories
The primary risk lies in rigidity. AI moves rapidly, and a chip built around today's Gemini architecture could look outdated by 2028
3
. The chip would work with future Gemini models only if Alphabet sticks with the same underlying neural-network structure, according to The Information2
. Google reportedly views Frozen v2 partly as a trial run and does not plan to produce it at the same scale as its TPUs2
.Google is not alone in pursuing this approach. Startup Taalas is already selling chips that print model weights and architecture directly onto silicon in a product called Hardcore
3
. Taalas claims its part serves up to 17,000 tokens per second, compared to roughly 150 per user on a top Nvidia GPU, while eliminating the need for expensive high-bandwidth memory.The race for custom AI silicon is shifting from running any model to fusing one model with the hardware itself
3
. For companies serving one model to billions of users, trading flexibility for speed, cost reduction, and lower power consumption can make strategic sense. Google has not officially confirmed the project, with a spokesperson noting only that teams experiment with high-efficiency ideas and that not every lab project reaches production3
.Summarized by
Navi
[3]
Yesterday•Technology

17 Jul 2026•Technology

22 Apr 2026•Technology

1
Policy and Regulation

2
Policy and Regulation

3
Technology
