Gemma 4 E2B — Multimodal Language Models AI Model
The compact Gemma 4, E2B, runs in bfloat16 on about 10GB of VRAM and still handles visual Q&A, captioning and image comparison with a 128K context. Pick it when the 31B model is more than the job needs.
Details
- CategoryMultimodal Language Models
- Year2026
- LicenseApache-2.0
Compliance & Provenance
- ProviderGoogle (open weights) · GPAI
- EU AI Act RiskLimited Risk
- Art. 50 TransparencyRequired — AI-generated outputs are marked
Inputs & Outputs
- TextInput · string
Text prompt or question
- Image 1Input · image · optional
Optional reference image for visual understanding. When provided, the model analyzes the image content together with the text prompt. Supports up to 4096x4096 resolution. Connect from any image-producing node (e.g. super-resolution, deblur, inpainting) to analyze processed results.
- Image 2Input · image · optional
Optional second reference image. When provided alongside Image 1, the model can compare, contrast, or jointly reason over both images with the text prompt. Supports up to 4096x4096 resolution.
- TextOutput · string
Generated text response
Tags
Alternatives in Multimodal Language Models
- Gemma 4 31B
Gemma 4 31B Dense flagship VLM, 256K context, thinking mode.
- Qwen3.6-35B-MoE
Qwen 3.6 35B/A3B MoE VLM, 256K context. fast inference.