Tag: AI inference

  • Google’s ‘Frozen v2’ Chip Targets Gemini AI Efficiency Gains, Deployment Targeted for 2028

    Google’s ‘Frozen v2’ Chip Targets Gemini AI Efficiency Gains, Deployment Targeted for 2028

    Google is reportedly developing a new server chip optimized specifically for its Gemini artificial intelligence models. Internally referred to as ‘Frozen v2’, the chip is designed to integrate core elements of Gemini’s architecture directly into the hardware, according to a report from The Information.

    This initiative underscores Google’s strategy to build bespoke infrastructure for its AI services. As demand for Gemini accelerates, the company aims to increase AI output while reducing computing power and electricity consumption.

    Frozen v2 Could Deliver Major Efficiency Gains

    The proposed chip is expected to minimize data movement and processing decisions during AI inference. These improvements would enhance the efficiency of serving Gemini models to users.

    According to the report, “Frozen v2 could deliver six to 10 times greater efficiency than Google’s latest custom AI chips when measured by the number of AI tokens generated per unit of power. Engineers are still working on the final design and deciding how much of the Gemini model architecture should be hardwired into the chip.”

    Google may deploy the new chip as early as 2028, though the timeline could shift as development continues. The company has not made any official statement. Notably, the Frozen project is intended to complement Google’s existing Tensor Processing Units (TPUs) rather than replace them, potentially creating a family of chips for AI processing.

    AI Computing Shortage Drives Infrastructure Push

    The chip development comes amid mounting pressure on Google’s AI computing capacity. Surging demand for Gemini services has reportedly caused internal strain and forced Google Cloud to turn down some external deals.

    A semiconductor chip based on Gemini would allow Google to perform more AI tasks with fewer computing resources, and also reduce the energy costs associated with generating AI tokens.

    Google’s overall strategy involves custom TPUs and other in-house semiconductors. The Frozen v2 version would be even more specialized by embedding architectural elements into the chip.