Alphabet-owned Google is reportedly developing a new custom AI server chip, internally codenamed “Frozen v2,” that could make its Gemini AI models six to ten times more power-efficient than the company’s latest Tensor Processing Units (TPUs). Rather than replacing Google’s existing TPU lineup, the chip is expected to form a new class of highly specialized AI accelerators designed specifically for Gemini inference by embedding parts of the model’s architecture directly into the hardware.
The project reflects Google’s growing focus on optimizing AI infrastructure as demand for Gemini-powered services continues to surge. Reports suggest the company is facing internal compute constraints that have, at times, forced Google Cloud to decline external AI infrastructure deals. By creating hardware tailored to Gemini, Google hopes to dramatically improve inference efficiency while lowering energy consumption and operating costs.
What Is “Frozen v2”?
Unlike general-purpose AI accelerators that are designed to run a wide variety of models, Frozen v2 is reportedly being engineered specifically around the Gemini family of AI models.
The chip would:
- Hardwire portions of Gemini’s architecture into silicon.
- Reduce unnecessary computation during inference.
- Increase the number of AI tokens generated per watt of power.
- Lower latency for Gemini-powered applications.
This approach resembles the concept of model-specific AI hardware, where the processor is optimized for a particular neural network architecture instead of supporting a broad range of models.
Frozen v2 at a Glance
| Feature | Details |
|---|---|
| Chip codename | Frozen v2 |
| Company | Google (Alphabet) |
| Primary purpose | Run Gemini AI models more efficiently |
| Expected efficiency gain | 6–10× more tokens per watt than latest TPUs |
| Deployment target | As early as 2028 (reported) |
| Status | Under development |
Why Google Is Building a Gemini-Specific Chip
AI inference has become one of the largest operating expenses for companies deploying large language models at scale.
Every response generated by Gemini consumes:
- Compute resources.
- Memory bandwidth.
- Electricity.
- Data center capacity.
By embedding parts of Gemini directly into the hardware, Google could eliminate redundant processing steps, allowing the chip to generate more AI output while consuming significantly less power. This would improve the economics of serving billions of AI queries across products such as Search, Workspace, Android, YouTube, and Google Cloud.
Potential Benefits
| Advantage | Impact |
|---|---|
| Higher power efficiency | Lower electricity consumption |
| Faster inference | Reduced response times |
| Lower operating costs | Cheaper AI deployment |
| Greater compute capacity | More users served with existing infrastructure |
Complementing, Not Replacing, Google’s TPUs
Reports indicate that Frozen v2 is not intended to replace Google’s Tensor Processing Units (TPUs).
Instead, Google is expected to maintain two hardware strategies:
- TPUs for flexible AI training and diverse inference workloads.
- Frozen v2 for highly optimized Gemini inference.
This dual-chip approach would allow Google to continue training future Gemini models on programmable TPUs while deploying specialized Frozen chips to serve mature production models at much lower cost.
TPU vs Frozen v2
| TPU | Frozen v2 |
|---|---|
| General-purpose AI accelerator | Gemini-specific AI accelerator |
| Supports multiple AI workloads | Optimized primarily for Gemini |
| Flexible architecture | Hardware tailored to model architecture |
| Training and inference | Primarily inference |
Addressing Google’s AI Compute Crunch
The reported development comes as Google faces rapidly growing demand for AI infrastructure.
According to reports, shortages in AI computing capacity have created internal tensions and even limited Google Cloud’s ability to accept certain external customer workloads. Frozen v2 is intended to alleviate these bottlenecks by dramatically increasing inference efficiency without requiring proportional growth in data center infrastructure.
Improving inference efficiency could also reduce Google’s dependence on third-party AI hardware while maximizing the value of its in-house silicon strategy.
Risks of Specialized AI Hardware
Although specialized chips can deliver significant performance gains, they also carry important trade-offs.
Potential challenges include:
- Gemini’s architecture may evolve before the chip reaches production.
- Less flexibility than general-purpose accelerators.
- High design and manufacturing costs.
- Long development timelines.
Because Frozen v2 reportedly targets deployment around 2028, engineers must ensure that the hardware remains compatible with future Gemini model architectures.
Opportunities and Risks
| Opportunities | Risks |
|---|---|
| Lower AI serving costs | AI models may evolve rapidly |
| Better energy efficiency | Specialized hardware is less flexible |
| Reduced latency | Expensive chip development |
| Increased infrastructure capacity | Longer deployment timelines |
Looking Ahead
Google’s reported Frozen v2 project highlights the next phase of the AI infrastructure race, where companies are moving beyond designing general-purpose AI accelerators toward building chips optimized for specific foundation models. If Frozen v2 delivers its reported six-to-tenfold improvement in tokens per watt, it could significantly reduce the cost of running Gemini while expanding Google’s AI capacity across consumer products and Google Cloud.
Although the project remains under development and is reportedly targeted for deployment as early as 2028, it underscores Google’s long-term commitment to vertically integrating AI software and hardware. As competition intensifies among Google, OpenAI, Microsoft, Meta, and Anthropic, specialized AI silicon may become an increasingly important differentiator in delivering faster, more efficient, and more cost-effective generative AI services.
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.