ViLT VQAv2 — Image Q&A AI Model
Ask a picture a question — what color is the car, how many people are there — and ViLT answers in a few words. Small and quick; for open-ended questions reach for a full vision-language model instead. CPU-capable.
Details
- CategoryImage Q&A
- Year2021
- LicenseApache-2.0
Compliance & Provenance
- ProviderOpen-source · Specialized
- EU AI Act RiskMinimal Risk
- Art. 50 TransparencyNot applicable
Inputs & Outputs
- ImageInput · image
Source image to answer questions about
- TextInput · string
Question about the image
- TextOutput · string
Answer to the question about the image
Tags
Alternatives in Image Understanding
- BLIP (Image Description)
Auto image captioning (BLIP Large). Fast.