---
title: "Google Releases Gemini 3.6 Flash and 3.5 Flash-Lite: 17% Fewer Output Tokens, Flash-Lite Positioned as a High-Volume Subagent"
url: https://xcube.enlightcorp.com.tw/en/news/google-gemini-3-6-flash-3-5-flash-lite-token-efficiency
category: "Frontier Models"
source_type: official
published: 2026-07-21
updated: 2026-07-21
lang: en
---

# Google Releases Gemini 3.6 Flash and 3.5 Flash-Lite: 17% Fewer Output Tokens, Flash-Lite Positioned as a High-Volume Subagent

Published 2026-07-21 · Updated 2026-07-21 · Official release · Source: [Google Blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/)

**Key answer:** Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are two Google models released on July 21, 2026: 3.6 Flash uses 17% fewer output tokens than 3.5 Flash, and 3.5 Flash-Lite is priced at $0.30 per million input tokens as a high-volume subagent. They show cloud-model competition shifting from headline scores toward the actual tokens and cost each task consumes.

On July 21, 2026, Tulsee Doshi, Senior Director of Product Management, announced Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber on behalf of the Gemini team. 3.6 Flash focuses on token efficiency: per the Artificial Analysis Index it reduces output token usage by 17% compared with 3.5 Flash, while improving coding, knowledge work, and multimodal performance.

Google's published comparisons for 3.6 Flash versus 3.5 Flash include DeepSWE 49% vs. 37%, MLE Bench 63.9% vs. 49.7%, and OSWorld-Verified 83.0% vs. 78.4%. It costs $1.50 per million input tokens and $7.50 per million output tokens, and is available through the Gemini API (Google AI Studio, Android Studio), Google Antigravity, Gemini Enterprise Agent Platform, and the Gemini app.

3.5 Flash-Lite is positioned as a low-latency, cost-effective subagent option for high-volume automation. Google cites about 350 output tokens per second per Artificial Analysis, 54% on Terminal-Bench 2.1 (vs. 31% for 3.1 Flash-Lite), and 54.2% on SWE-Bench Pro (vs. 49.6% for 3 Flash). Pricing is $0.30 per million input and $2.50 per million output tokens, via the Gemini API and Gemini Enterprise Agent Platform, and it is rolling into the Gemini app and Google Search. The security-focused 3.5 Flash Cyber is limited to governments and trusted partners in a pilot, and 3.5 Pro is still being tested with partners.

The change with the most direct engineering impact is API behavior: the Gemini API changelog noted the same day that the temperature, top_p, and top_k sampling parameters are now deprecated. Applications that rely on these parameters for output style or reproducibility need to be revalidated. The performance figures are measured by Google or Artificial Analysis, and real token savings will depend on prompts and task types.

In practice, gateway teams can start with a few checks: inventory existing requests that set temperature, top_p, or top_k and confirm how those parameters are handled when routed to Gemini 3.6 Flash or 3.5 Flash-Lite; split the workhorse model (3.6 Flash at $1.50 input / $7.50 output per million tokens) and the subagent model (3.5 Flash-Lite at $0.30 / $2.50) into separate routes and quotas; and measure output tokens on the same set of real tasks to check whether the claimed 17% savings hold on their own prompts. These numbers directly shape the total cost of a multi-agent architecture.

Google followed with 3.7 Flash in August and 3.8 Flash in September, so the Flash release cadence has clearly accelerated. Worth watching: how low-cost models like Flash-Lite hold up as subagents in multi-agent systems, and whether deprecating sampling parameters becomes a direction other vendors follow.

## X Cube view

Once vendors deprecate sampling parameters such as temperature, an OpenAI-compatible API or model gateway cannot assume every model accepts the same parameter set. Enterprises need the gateway to record and handle per-model parameter differences, and should compare workhorse and subagent models on total token cost per task rather than on unit price.

Tags: Gemini, Gemini API, Model Pricing, Token Efficiency, Subagents, Google, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, Gemini 3.5 Flash Cyber, Gemini Enterprise Agent Platform
