Skip to main content
Gemini Flash models from Google DeepMind are available in Tess for fast, cost-efficient chat and agent workflows. This page covers Gemini 3.6 Flash (production workhorse) and Gemini 3.5 Flash-Lite (highest throughput / lowest cost in the 3.5 Flash class).

When to use which

Gemini 3.6 Flash

  • Stronger coding and knowledge-work quality with lower verbosity than 3.5 Flash (up to ~17% fewer output tokens on comparable tasks)
  • Native multimodal: text, image, audio, and video
  • Reasoning, tools (function calling / MCP), and computer-use support
  • Large context for long-horizon agent workflows (1M input / up to 64K output)

Gemini 3.5 Flash-Lite

  • Fastest and most economical option in the Gemini 3.5 Flash class (up to ~350 output tokens/s)
  • Configurable thinking levels: prioritize latency/cost or deeper reasoning
  • Ideal as a cheap worker model in multi-agent pipelines

Pricing (Tess credits)

Values follow Models and Costs (credits per 100 tokens):
Screenshot placeholder — model picker: Capture the chat model selector with Gemini 3.6 Flash and Gemini 3.5 Flash-Lite visible (and selected once each).
Best practices
  • Prefer 3.6 Flash for coding agents and complex multimodal jobs.
  • Prefer 3.5 Flash-Lite for high-volume extraction, classification, and summarization.
  • Keep prompts tight — Flash models reward clear, scoped instructions and burn fewer tokens when you constrain output length.
See also: Models and Costs · Google DeepMind Gemini Flash.