Gemini 3.5 Transcribe (and Transcribe Live)
Changelog lists both ids GA on 2026-08-26, while the launch blog says public preview in AI Studio and Gemini Enterprise Agent Platform. Limits: 1 h per file request (30 min with diarization/timestamps), 10 min per live session. Diarization: docs say up to 8 speakers, blog says up to three - unresolved. Powers Rambler on Android and the Gemini app on macOS; coming to Chrome and Gboard. Press quotes $0.005/min (file) and $0.009/min (live) all-in. Model card (read 2026-09-29): https://deepmind.google/models/model-cards/gemini-3-5-audio/ - covers Gemini 3.5 Live Translate, Transcribe and Transcribe Live (card dated 2026-08-26); no numeric evals in the card itself; knowledge cutoff January 2025; did not reach any Tracked or Critical Capability Levels under the Frontier Safety Framework. Card lists hallucinations and occasional slowness/timeouts as limitations; surfaces: Antigravity, Gboard, Gemini app, Vertex AI, Google Workspace.
- Input
- audio, text
- Output
- text
- License
- proprietary
- Pricing
- audio input: $2 · output: $12 (per 1M tokens (USD) for gemini-3.5-transcribe (~$0.003/min audio in + ~$0.002/min text out); gemini-3.5-transcribe-live $3.50 in / $21.00 out (~$0.005 + ~$0.004 per min)) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Gemini API (Interactions API, files) | gemini-3.5-transcribe | — | docs |
| Gemini Live API (WebSocket streaming) | gemini-3.5-transcribe-live | — | docs |
| Google AI Studio | — | aistudio.google.com | — |
Notable capabilities (3)
- Smart transcription: Handles self-corrections, removes filler words and auto-formats text; custom vocabulary biasing up to 1,000 terms. source
- Low word error rate: Google cites Artificial Analysis WER of 2.6% (non-streaming) and 4.0% (streaming); 70% faster time-to-final than Chirp 3. source
- 85+ languages with code-switching, diarization, word timestamps: Utterance-level language detection across 85+ languages; speaker diarization; word-level timestamps (not combinable with custom vocabulary). source
Google's Gemini-based speech-to-text, successor in practice to Cloud Chirp 3 for developers.
Sources: model page, pricing, changelog, launch blog.
Other Google DeepMind models
Gemini 3.8 Flash TTS · Gemini 3.8 Live · Gemini 3.8 Flash · Lyria 3.5 · Gemini 3.5 Flash-Lite · Nano Banana 2 Lite (Gemini 3.1 Flash-Lite Image) · Gemini Omni Flash (Omni 1.1 Flash) · Gemini Embedding 2 · Gemma 4 · Nano Banana 2 (Gemini 3.1 Flash Image) · Nano Banana Pro (Gemini 3 Pro Image) · Gemini Robotics 2 · Gemini Robotics ER 2 · Gemini Robotics On-Device 2 · Gemini 3.5 Live Translate · Gemini 3.1 Pro · Veo 3.1 · Genie 3 · Lyria RealTime · Gemini 3.7 Flash · Gemini 3.6 Flash · Gemini 3.5 Flash · Gemini 3.1 Flash TTS (preview) · Gemini 3.1 Flash Live (preview) · Lyria 3 (Clip / Pro) · Lyria 2 · Gemini 2.5 Flash Native Audio (Live, preview) · Gemini 2.5 Flash-Lite · Gemini 2.5 Flash · Gemini 2.5 Pro · Gemini 2.5 Flash TTS / Pro TTS · Gemini 3.1 Flash-Lite · Nano Banana (Gemini 2.5 Flash Image) · Gemini Robotics-ER 1.5 / 1.6 · Imagen 4