Google released three new proprietary models today: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. The company positions them as its most token-efficient models to date, designed to reduce costs for AI agent applications at scale.

Gemini 3.6 Flash delivers the headline performance gain. On long-horizon engineering tasks, the model cuts token consumption by up to 65 percent compared to its predecessors. This efficiency translates directly to lower inference costs. Google prices Gemini 3.6 Flash at $1.50 per million input tokens and $7.50 per million output tokens via its API.

The price tiers span different use cases. Gemini 3.5 Flash-Lite targets budget-conscious builders with rates of just $0.30 per million input tokens and $2.50 per million output tokens. That undercuts the standard Gemini 3.5 Flash, which costs $1.50/$9.00 per million tokens. Gemini 3.5 Flash Cyber appears positioned for specialized applications, though Google provided limited details on that variant.

The efficiency gains matter because token costs compound fast for AI agents. These systems make multiple sequential reasoning steps, each consuming tokens. A 65 percent reduction on complex tasks like code generation or multi-step engineering workflows reduces both operational expenses and latency. Google also announced Gemini 3.5 Pro is coming, suggesting the company plans to expand its model lineup further.

The release targets competition from Anthropic, OpenAI, and other AI labs racing to optimize token efficiency. Cheaper inference enables broader deployment of AI agents in production environments. It also lowers the barrier for smaller companies to build agent-based applications without absorbing massive compute costs.