Google is working on a new AI chip designed to make Gemini more efficient

Six to ten times more efficient. That's the performance leap Google is reportedly targeting with its next-generation AI chip — and it could reshape how every developer on the planet interacts with Gemini. Google is working on a new AI chip designed to make Gemini more efficient at a moment when the

Share
Editorial illustration: A close-up of a silicon wafer under dramatic side-lighting, its intricate circuit pathways and geome — MonstarX

```html

Google is working on a new AI chip designed to make Gemini more efficient

Six to ten times more efficient. That's the performance leap Google is reportedly targeting with its next-generation AI chip — and it could reshape how every developer on the planet interacts with Gemini. Google is working on a new AI chip designed to make Gemini more efficient at a moment when the entire industry is rethinking what AI infrastructure actually costs, and who controls it.

For developers and founders across Asia, this isn't background noise. It's a signal about where AI compute is heading — and what that trajectory means for the products you're building right now.

What Happened

Alphabet, Google's parent company, is designing a new server chip internally codenamed "Frozen v2", according to a TechCrunch report citing The Information. The chip is slated for release sometime in 2028, and the performance target is striking: between six and ten times more efficient than Google's existing AI chips, measured by the number of tokens generated per unit of power.

Google didn't confirm the report directly when TechCrunch reached out, but it didn't deny it either. The company's statement is worth reading carefully:

"Our teams are constantly researching and experimenting with new innovations to deliver maximum performance and efficiency for our users and customers. While not every project moves into production, this rigorous exploration is central to our full stack approach. By co-designing our hardware and software from the ground up, we ensure our systems are integrated and highly optimized for real-world workloads."

That phrase — "co-designing our hardware and software from the ground up" — is the real headline buried inside a carefully worded non-confirmation. Google is describing a vertical integration strategy: owning the full stack from silicon to model to API. That's a fundamentally different posture than buying chips from a third party and optimizing software around them.

The timing also matters. This news lands as investor concerns about AI spending have started to cool market enthusiasm. OpenAI announced its first custom inference chip, dubbed Jalapeño, in June. Anthropic is reportedly in discussions with Samsung about a chipmaking partnership. The hyperscalers aren't just competing on model benchmarks anymore — they're competing on the economics of inference at scale. Frozen v2 is Google's move in that game.

Why It Matters for Asia

Asia is not a passive consumer of AI infrastructure decisions made in Silicon Valley. The region runs some of the world's most demanding AI workloads — from real-time translation across dozens of languages to fraud detection systems processing millions of transactions per second across Southeast Asia's fragmented payments landscape. Compute efficiency isn't an abstract metric here. It directly determines whether a startup in Jakarta or a fintech in Ho Chi Minh City can afford to run Gemini-powered features at production scale.

Right now, AI inference costs are a genuine constraint for most Asian startups. The economics of calling a frontier model API dozens of times per user session are difficult to make work at the margins that Southeast Asian markets demand. Token-per-watt efficiency improvements don't just benefit Google's data centers — they eventually flow downstream as lower API pricing, higher rate limits, and faster response times. A 6–10x efficiency gain on Google's infrastructure is, over time, a 6–10x improvement in what developers can afford to build.

There's a geopolitical layer here too. Asia's AI ambitions have consistently run into one hard constraint: Nvidia's dominance over AI compute and the supply chain pressures that come with it. Every major player — Google, OpenAI, Anthropic, and increasingly domestic champions in China, South Korea, and Japan — is trying to reduce that dependency. For Asian cloud providers and government-backed AI initiatives, watching Google demonstrate that custom silicon can deliver this kind of efficiency improvement provides a concrete blueprint. Expect more regional players to accelerate their own chip programs in response.

The broader Asia tech narrative here is about sovereignty. Countries like India, Singapore, and South Korea have made significant investments in AI infrastructure precisely because they understand that depending entirely on foreign silicon creates strategic vulnerability. Google's Frozen v2 project — whether or not it ships on schedule in 2028 — validates the strategic logic behind those investments.

What This Means for Developers

If you're building AI-powered products today, Frozen v2 won't change your sprint planning for the next 18 months. It ships in 2028, and the downstream effects on API pricing and availability will take longer still to materialize. But there are concrete things to think about now.

Model selection will increasingly be an infrastructure decision. As Google, OpenAI, and Anthropic each develop custom silicon optimized for their own models, the performance characteristics of those models will diverge in ways that benchmark leaderboards don't capture. A model running on purpose-built inference hardware will behave differently in production — lower latency, more consistent throughput, better cost curves at scale — than the same model running on general-purpose GPUs. Developers who understand this will make better architectural choices.

The token economy is getting cheaper, but not uniformly. Efficiency gains tend to concentrate first at the hyperscaler level, then propagate to enterprise tiers, then eventually reach the developer API. If you're building on MonstarX or any AI-native dev platform, watch for pricing changes in Gemini API tiers over the 2027–2028 window. The efficiency improvements from Frozen v2 will likely show up as cost reductions before they show up as new capabilities.

Inference architecture matters more than it used to. The shift toward custom inference chips — Google with Frozen v2, OpenAI with Jalapeño — signals that the industry is optimizing hard for inference rather than training. For developers, this means the gap between "runs in a notebook" and "runs in production at acceptable cost" is narrowing. That's good news for teams building real products rather than demos.

Plan for model lock-in risk. As each major AI provider builds silicon optimized for their own models, switching costs between providers will increase. A product deeply integrated with Gemini's API surface today will face real friction migrating to a different provider in 2028, not just because of API differences but because the underlying performance characteristics will be shaped by hardware that doesn't exist yet. Build with abstraction layers where you can. Use connectors and integration patterns that keep your application logic decoupled from any single model provider.

Watch the token-per-watt metric. This is the number Google used to describe Frozen v2's efficiency target, and it's worth adding to your mental model of AI infrastructure. As sustainability concerns grow — particularly in markets like Singapore and Japan where data center energy consumption is under regulatory scrutiny — token-per-watt will become a procurement criterion, not just a marketing number.

Key Takeaways

  • Frozen v2 is real, or real enough. Google neither confirmed nor denied the report, but its statement about "co-designing hardware and software from the ground up" is a clear signal that custom silicon is central to its AI strategy. A 2028 release date gives the project time to mature, but the direction is set.
  • 6–10x efficiency improvement is the stated target. Measured by tokens per unit of power, this would represent a step-change in how cost-effectively Google can serve Gemini inference at scale. If even half of that gain materializes, it changes the economics of AI APIs meaningfully.
  • The custom chip race is now three-way, at minimum. Google with Frozen v2, OpenAI with Jalapeño, Anthropic in discussions with Samsung. The era of universal Nvidia dependency is ending — not because Nvidia is weakening, but because the largest AI labs have decided vertical integration is worth the investment.
  • For Asian developers, the downstream effect is lower inference costs and faster APIs — eventually. The timeline is 2028 and beyond, but the trajectory is clear. Products that seem economically marginal today may become viable at scale as these efficiency gains propagate through pricing.
  • Build for provider flexibility now. The custom silicon race will deepen model lock-in over time. Architectural decisions you make today — abstraction layers, provider-agnostic interfaces, modular integration patterns — will determine how much optionality you have in 2028 when the hardware landscape looks very different.
  • Token-per-watt is a metric worth tracking. It's not just an engineering benchmark — it will shape API pricing, data center policy, and eventually the competitive dynamics of which AI providers can operate profitably at scale in energy-constrained markets.

The deeper story here isn't about a single chip. It's about the AI industry making a structural bet that the path to sustainable economics runs through vertical integration — owning the silicon, the model, and the API surface simultaneously. For developers building on top of these platforms, that bet has real consequences: better performance curves over time, but also deeper dependencies and higher switching costs. The developers who understand that trade-off now will navigate it more deliberately than those who discover it later.

```