Gemma 4
Sizes E2B, E4B (128K context), 12B, 26B A4B MoE, 31B dense (256K context); base and -it variants plus QAT/GGUF quantized repos. 12B released later (HF repo 2026-05-23). Audio input on E2B/E4B/12B only. 140+ languages. pricing is third-party (OpenRouter), not Google.
- Context window
- 262,144 tokens
- Input
- text, image, audio, video
- Output
- text
- License
- apache-2.0
- Pricing
- input: $0.09 · output: $0.34 (per 1M tokens, OpenRouter price for google/gemma-4-31b-it (26B-A4B: $0.09 / $0.30; free variants exist). Weights free to download) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | google/gemma-4-31B-it | huggingface.co/google/gemma-4-31B-it | — |
| Hugging Face | google/gemma-4-26B-A4B-it | huggingface.co/google/gemma-4-26B-A4B-it | — |
| Hugging Face | google/gemma-4-12B-it | huggingface.co/google/gemma-4-12B-it | — |
| Hugging Face | google/gemma-4-E4B-it | huggingface.co/google/gemma-4-E4B-it | — |
| Hugging Face | google/gemma-4-E2B-it | huggingface.co/google/gemma-4-E2B-it | — |
| OpenRouter | google/gemma-4-31b-it | — | — |
| OpenRouter | google/gemma-4-26b-a4b-it | — | — |
| Google docs | — | — | docs |
Notable capabilities (4)
- First Apache-2.0 Gemma: First Gemma generation under the permissive Apache 2.0 license instead of Google's custom Gemma terms. source
- Intelligence per parameter: 31B dense ranked #3 and 26B A4B MoE #6 among open models on Arena at launch. source
- On-device agentic models: E2B/E4B edge models with native audio+vision, function calling and structured JSON, running offline on phones/Raspberry Pi/Jetson. source
- Encoder-free unified 12B (found after launch): Gemma 4 12B, added later, is a unified encoder-free multimodal model with native audio. source
Latest open-weights family from Google DeepMind (built from Gemini 3 research); run locally or self-host.
from transformers import pipeline
pipe = pipeline("image-text-to-text", model="google/gemma-4-E4B-it")
print(pipe(text=[{"role":"user","content":[{"type":"text","text":"Hello!"}]}]))
Sources: Gemma docs, launch blog, HF repo.
Other Google DeepMind models
Gemini 3.8 Flash TTS · Gemini 3.8 Live · Gemini 3.8 Flash · Gemini 3.5 Transcribe (and Transcribe Live) · Lyria 3.5 · Gemini 3.5 Flash-Lite · Nano Banana 2 Lite (Gemini 3.1 Flash-Lite Image) · Gemini Omni Flash (Omni 1.1 Flash) · Gemini Embedding 2 · Nano Banana 2 (Gemini 3.1 Flash Image) · Nano Banana Pro (Gemini 3 Pro Image) · Gemini Robotics 2 · Gemini Robotics ER 2 · Gemini Robotics On-Device 2 · Gemini 3.5 Live Translate · Gemini 3.1 Pro · Veo 3.1 · Genie 3 · Lyria RealTime · Gemini 3.7 Flash · Gemini 3.6 Flash · Gemini 3.5 Flash · Gemini 3.1 Flash TTS (preview) · Gemini 3.1 Flash Live (preview) · Lyria 3 (Clip / Pro) · Lyria 2 · Gemini 2.5 Flash Native Audio (Live, preview) · Gemini 2.5 Flash-Lite · Gemini 2.5 Flash · Gemini 2.5 Pro · Gemini 2.5 Flash TTS / Pro TTS · Gemini 3.1 Flash-Lite · Nano Banana (Gemini 2.5 Flash Image) · Gemini Robotics-ER 1.5 / 1.6 · Imagen 4