Twitter/X

Gemini 3.6 Flash reduces token usage by up to 65% on complex coding tasks versus…

Brief

Josh Woodward summarized launches on 2026-07-21: Gemini 3.6 Flash reduces token usage up to 65% for complex coding while matching cost and improving quality; Gemini 3.5 Flash-Lite hits 350 output tokens/sec for fast, low-cost document and agentic tasks; Gemini 3.5 Flash Cyber targets vulnerability detection/patching, and Gemini 3.5 Pro is in partner testing.

Why it matters

Gemini 3.6 Flash reduces token usage by up to 65% on complex coding tasks versus prior versions, claims higher quality at the same cost, and went live in the Gemini app on 2026-07-21.

Key details

  • Gemini 3.5 Flash-Lite reaches speeds of 350 output tokens/sec, positioned as a fast, cost-effective option for document processing and agentic search, and is live in the Gemini app.
  • Google DeepMind also launched Gemini 3.5 Flash Cyber to find and patch critical software vulnerabilities, and announced Gemini 3.5 Pro has entered partner testing.
Source evidence

Today’s launches are all about better performance, lower latency, and a smaller bill.

  • 3.6 Flash cuts token usage by up to 65% on complex coding
  • 3.5 Flash-Lite reaches speeds of 350 output tokens/sec

Both are live in the Gemini app today!

Next up: Gemini 3.5 Pro, which has officially entered partner testing.

Google DeepMind (@GoogleDeepMind)

We’re rolling out three new models to make AI agents faster, smarter, and cheaper at scale:

🔵 Gemini 3.6 Flash: It uses fewer tokens than 3.5 Flash to deliver higher quality work at the exact same cost.

🔵 Gemini 3.5 Flash-Lite: A fast, cost-effective option for everyday tasks like processing documents and agentic search.

🔵 Gemini 3.5 Flash Cyber: A cybersecurity model built to find and patch critical software vulnerabilities.

— https://nitter.net/GoogleDeepMind/status/2079589698490572961#m