Official releaseOpen Models / On-prem AI

Google Releases Gemma 4 12B: an Encoder-Free Multimodal Open Model and the First Mid-Size Gemma with Native Audio

Gemma 4 12B processes text, images, and audio in a single decoder-only architecture, runs locally on devices with 16GB of VRAM or unified memory, and works with inference frameworks such as vLLM and SGLang.

Key answer

Gemma 4 12B is a 12-billion-parameter open multimodal model released by Google on June 3, 2026, with a unified encoder-free architecture that natively ingests text, images, and audio; Google's Gemma release notes list up to 256K context. It matters because on-prem and local deployments can now cover speech, vision, and text with one mid-size model instead of stitching together several specialized ones.

On June 3, 2026, Google published the Gemma 4 12B developer guide on the Google Developers Blog. It is a 12-billion-parameter dense multimodal model using the same single decoder-only transformer structure as Gemma 4 31B Dense, with a unified, encoder-free design that ingests raw 48x48 pixel patches and 16 kHz audio directly; the vision embedder is only about 35M parameters.

Google notes that audio input in the Gemma family was previously limited to small edge architectures such as E4B, and that Gemma 4 12B is the first medium-sized model to natively ingest audio. Listed uses include automatic speech recognition, diarization, video understanding, agentic reasoning, and coding, and it can process multi-minute videos at 1 FPS with synchronized audio. The Gemma release notes list it as "Gemma 4 12B Unified" with up to 256K context.

On deployment, Google says it is small enough to run locally on dedicated GPU laptops with 16GB of VRAM or unified memory, with optimizations for Apple Silicon on macOS and a multi-token prediction variant for faster local inference. It is distributed via Hugging Face, Kaggle, LM Studio, and Ollama, and works with Hugging Face Transformers, llama.cpp, MLX, vLLM, SGLang, and Unsloth for fine-tuning.

Caveats: the developer guide does not publish quantitative benchmark scores, so capability claims are Google's own. The Gemma 4 family is released under the Apache 2.0 license, per Google's Gemma 4 announcement on the Google Open Source Blog. Running in 16GB also does not mean throughput is sufficient for team use at long context and concurrency.

For on-prem deployment teams, the official information supports an evaluation checklist. Licensing: the Gemma 4 family is released under Apache 2.0 per Google's Gemma 4 announcement on the Google Open Source Blog, so internal legal review can start from that. Inference stack: test on vLLM or SGLang first, confirming that text, image, and 16 kHz audio inputs are all served correctly. Capacity: starting from Google's 16GB VRAM baseline, measure memory use and throughput at long context and under concurrency. Performance: compare latency between the standard and multi-token prediction variants. Finally, validate quality on in-house transcription, document-image, and support-conversation data to decide whether it can replace several existing specialized models.

What to watch: how mature vLLM and SGLang support is for its audio and video inputs, actual memory requirements at 256K context, and the quality of community quantizations and fine-tuned variants. These determine whether it moves from a local demo into an enterprise on-prem inference service.

X Cube view

For enterprise platforms serving open models on on-prem nodes such as DGX Spark / GB10, Gemma 4 12B suggests one mid-size model could cover transcription, image, and text understanding, reducing the number of specialized models to deploy and operate. Before adding it to a production model catalog, teams should validate multimodal input support, long-context memory use, and concurrent throughput with vLLM or SGLang on their own hardware.