Gemma 4
Run text, image, and supported audio workloads on your own infrastructure with Gemma 4. Choose an edge, dense, or mixture-of-experts variant to match your hardware and task.
The verdict on Gemma 4: Developers building local or self-hosted assistants who can manage deployment, evaluation, and data handling Gemma 4 is worth evaluating when you want downloadable weights, a permissive license, and a choice between edge-sized and larger models. Pricing: Free Apache 2.0 model weights. Hardware, cloud compute, managed hosting, and serving costs are separate.. Last reviewed: September 2026.
Best For
Developers building local or self-hosted assistants who can manage deployment, evaluation, and data handling
Standout Feature
Five deployment options: E2B, E4B, 12B Unified, 26B A4B MoE, and 31B Dense, with audio input on the three smaller dense variants
TL;DR
Choose Gemma 4 when control over model deployment matters and you can support the infrastructure; it is not a managed assistant subscription.
Alternatives
Overview
Run and fine-tune Google's current Gemma open-weight models on infrastructure you choose, from edge devices to GPU workstations and cloud deployments. Gemma 4 includes E2B, E4B, 12B Unified, 26B A4B mixture-of-experts, and 31B dense variants. Google publishes model weights and deployment documentation, with downloads through Hugging Face and Kaggle.
All five variants accept text and images and generate text. Audio input is supported on E2B, E4B, and 12B Unified; the 26B A4B and 31B models do not accept audio. E2B and E4B support 128K-token context windows, while 12B Unified, 26B A4B, and 31B support 256K. Google's model card distinguishes pretraining across more than 140 languages from out-of-the-box support for more than 35 languages. Native function calling supports applications that connect the model to external tools.
The weights are distributed under Apache 2.0, including commercial use subject to the license. Free weights do not mean free operation: budget for hardware or cloud compute, serving, monitoring, and any hosted provider's charges. Memory needs depend on the variant, quantization, context length, and workload. Validate generated answers and tool calls before using them in consequential decisions or actions.
Our Take
The important choice is not just parameter count: audio support, memory use, context length, and serving costs differ by variant. Start with a representative task and a realistic hardware budget, then test output quality and failure cases before scaling. This assessment is based on Google's current model documentation, not a new hands-on benchmark. Self-hosting gives you control over deployment, but does not automatically make an application private, secure, or compliant.
- • Permissive licensing and downloadable weights give developers deployment flexibility
- • Multiple architectures and sizes support different hardware budgets
- • Text, vision, and selected audio input can support several tasks in one deployment
- • Official model cards and deployment documentation explain variant-specific trade-offs
- • Serving, updates, evaluation, and access controls remain your responsibility
- • Larger variants and long contexts can require substantial memory and compute
- • Generated facts, interpretations, and tool calls still need validation
- • Audio input is not available on every variant, and output is text only
Key Features
- Apache 2.0 downloadable model weights
- E2B, E4B, 12B Unified, 26B A4B MoE, and 31B Dense variants
- Text and image input with text output across the family
- Audio input on E2B, E4B, and 12B Unified only
- 128K context on E2B/E4B; 256K on 12B/26B A4B/31B
- Native function calling for tool-connected applications
- Pretraining in 140+ languages and 35+ languages supported out of the box
- Pre-trained and instruction-tuned weights, with documented fine-tuning options
- Official quantized formats and deployment guidance for local and cloud environments
Trust & Data
Verified 2026-09
- Training on your data
- Not publishedGemma 4 is a downloadable model family, not one hosted service. Local inference does not require sending prompts to Google. API access and data-use policies depend on the runtime or hosting provider you choose.
- Compliance
- Not published
- Retention
- Determined by the deployment, logging configuration, and hosting provider. Downloadable weights do not define a hosted data-retention policy.
- Export
- Not published
Sources: ai.google.dev/model card 4, ai.google.dev/apache 2, cloud.google.com
As published by the vendor. Verify independently before purchase decisions.
People Also Use
Other Models tools builders reach for alongside Gemma 4.
DeepSeek
Frontier-level coding and reasoning with a 1M-token context window at a fraction of Western competitor cost. Now on DeepSeek V4 (Flash and Pro tiers, both mixture-of-experts); verify current per-token pricing directly before budgeting on cached historical rates.
Mistral
Access Mistral Large 3, an open-weight, multilingual, multimodal flagship model at a fraction of the cost of closed competitors, from cloud API to edge deployment.
Open WebUI
Run a self-hosted chat interface for local or API-based LLMs like Ollama behind your own login and controls, free at any scale if you keep default branding.
Azure OpenAI Service
Access GPT and other OpenAI models through Azure with enterprise compliance, networking, and regional data controls. Now offers Global, Data Zone, and Regional deployment types.
Poe
Chat with and compare many AI models and community-built bots through one subscription instead of juggling separate accounts. Pricing now transparent against per-model USD/token rates.
HuggingChat
Chat with today's best open-weight AI models for free, with no API key, subscription, or vendor lock-in.