ViLT VQAv2 — Image Q&A AI Model

Ask a picture a question — what color is the car, how many people are there — and ViLT answers in a few words. Small and quick; for open-ended questions reach for a full vision-language model instead. CPU-capable.

Details

  • CategoryImage Q&A
  • Year2021
  • LicenseApache-2.0

Compliance & Provenance

  • ProviderOpen-source · Specialized
  • EU AI Act RiskMinimal Risk
  • Art. 50 TransparencyNot applicable

Inputs & Outputs

  • ImageInput · image

    Source image to answer questions about

  • TextInput · string

    Question about the image

  • TextOutput · string

    Answer to the question about the image

Tags

  • visual-qa
  • image-understanding
  • question-answering
  • multimodal
  • no-gpu
  • fast

Alternatives in Image Understanding

Resources