Back to Directory
Gemma 4 logo

Gemma 4

Run text, image, and supported audio workloads on your own infrastructure with Gemma 4. Choose an edge, dense, or mixture-of-experts variant to match your hardware and task.

Models
4.5free

The verdict on Gemma 4: Developers building local or self-hosted assistants who can manage deployment, evaluation, and data handling Gemma 4 is worth evaluating when you want downloadable weights, a permissive license, and a choice between edge-sized and larger models. Pricing: Free Apache 2.0 model weights. Hardware, cloud compute, managed hosting, and serving costs are separate.. Last reviewed: September 2026.

Best For

Developers building local or self-hosted assistants who can manage deployment, evaluation, and data handling

Standout Feature

Five deployment options: E2B, E4B, 12B Unified, 26B A4B MoE, and 31B Dense, with audio input on the three smaller dense variants

TL;DR

Choose Gemma 4 when control over model deployment matters and you can support the infrastructure; it is not a managed assistant subscription.

Alternatives

Overview

Run and fine-tune Google's current Gemma open-weight models on infrastructure you choose, from edge devices to GPU workstations and cloud deployments. Gemma 4 includes E2B, E4B, 12B Unified, 26B A4B mixture-of-experts, and 31B dense variants. Google publishes model weights and deployment documentation, with downloads through Hugging Face and Kaggle.

All five variants accept text and images and generate text. Audio input is supported on E2B, E4B, and 12B Unified; the 26B A4B and 31B models do not accept audio. E2B and E4B support 128K-token context windows, while 12B Unified, 26B A4B, and 31B support 256K. Google's model card distinguishes pretraining across more than 140 languages from out-of-the-box support for more than 35 languages. Native function calling supports applications that connect the model to external tools.

The weights are distributed under Apache 2.0, including commercial use subject to the license. Free weights do not mean free operation: budget for hardware or cloud compute, serving, monitoring, and any hosted provider's charges. Memory needs depend on the variant, quantization, context length, and workload. Validate generated answers and tool calls before using them in consequential decisions or actions.

Our Take

The important choice is not just parameter count: audio support, memory use, context length, and serving costs differ by variant. Start with a representative task and a realistic hardware budget, then test output quality and failure cases before scaling. This assessment is based on Google's current model documentation, not a new hands-on benchmark. Self-hosting gives you control over deployment, but does not automatically make an application private, secure, or compliant.

Was this useful?
Pros
  • Permissive licensing and downloadable weights give developers deployment flexibility
  • Multiple architectures and sizes support different hardware budgets
  • Text, vision, and selected audio input can support several tasks in one deployment
  • Official model cards and deployment documentation explain variant-specific trade-offs
Cons
  • Serving, updates, evaluation, and access controls remain your responsibility
  • Larger variants and long contexts can require substantial memory and compute
  • Generated facts, interpretations, and tool calls still need validation
  • Audio input is not available on every variant, and output is text only

Key Features

  • Apache 2.0 downloadable model weights
  • E2B, E4B, 12B Unified, 26B A4B MoE, and 31B Dense variants
  • Text and image input with text output across the family
  • Audio input on E2B, E4B, and 12B Unified only
  • 128K context on E2B/E4B; 256K on 12B/26B A4B/31B
  • Native function calling for tool-connected applications
  • Pretraining in 140+ languages and 35+ languages supported out of the box
  • Pre-trained and instruction-tuned weights, with documented fine-tuning options
  • Official quantized formats and deployment guidance for local and cloud environments

Trust & Data

Verified 2026-09

Training on your data
Not publishedGemma 4 is a downloadable model family, not one hosted service. Local inference does not require sending prompts to Google. API access and data-use policies depend on the runtime or hosting provider you choose.
Compliance
Not published
Retention
Determined by the deployment, logging configuration, and hosting provider. Downloadable weights do not define a hosted data-retention policy.
Export
Not published

Sources: ai.google.dev/model card 4, ai.google.dev/apache 2, cloud.google.com

As published by the vendor. Verify independently before purchase decisions.

Other Models tools builders reach for alongside Gemma 4.