Google is developing "Frozen v2," a custom server chip designed to run Gemini models with exceptional efficiency. The processor bakes Gemini's architecture directly into silicon rather than using general-purpose hardware, a strategy that reportedly delivers 6 to 10 times better performance per watt than current TPUs.

The chip targets 2028 deployment and addresses a critical business problem: AI inference costs. By optimizing silicon specifically for Gemini's computational patterns, Google would slash the expense of serving billions of AI queries. This directly impacts pricing competition against OpenAI and Anthropic, both of which rely on commodity or licensed accelerators for inference workloads.

The approach mirrors strategies from chip leaders. Apple bakes neural engines into iPhones for on-device ML. Amazon designed Trainium chips for training and Inferentia for inference. Microsoft invested in Maia and Cobalt chips. This trend reflects a simple economic reality: general-purpose chips leave performance on the table. Specialization wins.

Frozen v2's efficiency gains matter most at scale. Google runs trillions of inference requests annually across Search, Gmail, and Cloud services. A 6 to 10x efficiency improvement translates to proportional savings in power consumption, cooling, and hardware costs. The company could either pass savings to customers through lower API prices or pocket the margin advantage.

The 2028 timeline gives Google two years to complete design, tape-out, and manufacturing. Custom silicon typically requires 18 to 24 months from design completion to volume production, so the schedule appears aggressive but plausible for a company with Google's resources and expertise.

Internal sources claim Frozen v2 exists, though Google hasn't confirmed details publicly. The secrecy around custom silicon development is standard practice, especially in competitive markets. OpenAI partners with Microsoft on custom chips. Anthropic