Google is building an internal server chip called Frozen v2 that embeds the Gemini AI model’s architecture directly into silicon.
Sources cited by The Information say the chip could be 6 to 10 times more efficient at serving AI responses than Google’s current TPU chips. The company plans to deploy it starting in 2028 and views Frozen v2 as a test run for specialised chips, with a smaller production volume than its TPU line.
Unlike Google’s TPUs, which work with many models, Frozen v2 has parts of Gemini’s model structure built right into the hardware. The name follows the same logic as “freezing” parameters in AI models, where you lock values so they stop changing. With Frozen v2, a portion of the model gets permanently frozen into the chip itself, which cuts down on compute steps and speeds up responses.
The original idea reportedly came from Jeff Dean, Google Deepmind’s chief scientist. His first Frozen design called for embedding the model weights directly into the chip. Weights are the specific settings that determine how an AI model responds to queries. Google scrapped that approach because the chip would have only worked with a single Gemini version and would have become outdated too quickly.
Frozen v2 takes a more flexible path by embedding the model architecture instead of weights, meaning the underlying blueprint rather than the tuned parameters. New weights can still be loaded onto the chip. How much of the architecture will actually be hardcoded has not been decided yet, according to The Information.
Squeezing better margins out of AI inference
Because the chip only works as long as Google sticks with the same model architecture, it probably will not become a product for outside customers. Google already leases its TPUs to Meta, offers them to external cloud customers, and positions them through its TPU@Premises program as an alternative to Nvidia with an internal goal of capturing ten percent of Nvidia’s annual revenue. Frozen v2, by contrast, is meant to ease Google’s internal crunch on AI compute capacity.
If the chip delivers on its promise, it could still become a major competitive edge. In the AI business, how well companies optimise inference costs increasingly determines their margins. Google could use Frozen v2 to run powerful models at lower prices and take market share from OpenAI and Anthropic.
What it means
For Google’s internal teams, this shift means the hardware is tuned specifically for Gemini. This reduces the computing power needed to run the model, which lowers the cost per query. It also means the hardware becomes obsolete faster if Google changes its model architecture, limiting the chip’s lifespan. For the wider market, Google is likely keeping this technology internal to protect its competitive position against rivals like OpenAI and Anthropic.




