Gemma 4 E2B — Multimodal Language Models AI Model

The compact Gemma 4, E2B, runs in bfloat16 on about 10GB of VRAM and still handles visual Q&A, captioning and image comparison with a 128K context. Pick it when the 31B model is more than the job needs.

Details

  • CategoryMultimodal Language Models
  • Year2026
  • LicenseApache-2.0

Compliance & Provenance

  • ProviderGoogle (open weights) · GPAI
  • EU AI Act RiskLimited Risk
  • Art. 50 TransparencyRequired — AI-generated outputs are marked

Inputs & Outputs

  • TextInput · string

    Text prompt or question

  • Image 1Input · image · optional

    Optional reference image for visual understanding. When provided, the model analyzes the image content together with the text prompt. Supports up to 4096x4096 resolution. Connect from any image-producing node (e.g. super-resolution, deblur, inpainting) to analyze processed results.

  • Image 2Input · image · optional

    Optional second reference image. When provided alongside Image 1, the model can compare, contrast, or jointly reason over both images with the text prompt. Supports up to 4096x4096 resolution.

  • TextOutput · string

    Generated text response

Tags

  • multimodal
  • visual-qa
  • image-understanding
  • reasoning
  • dense
  • gpu-medium

Alternatives in Multimodal Language Models

  • Gemma 4 31B

    Gemma 4 31B Dense flagship VLM, 256K context, thinking mode.

  • Qwen3.6-35B-MoE

    Qwen 3.6 35B/A3B MoE VLM, 256K context. fast inference.

Resources