Bildverständnis und visuelle Fragen
Ein Bild ansehen und freie Fragen dazu beantworten
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Takes an image, optionally with a text question, and outputs a natural-language description or answer — what the person is doing, where the scene might be. Unlike OCR it targets meaning rather than characters, and unlike classification it answers open-ended questions instead of choosing from a fixed label set.
Wie sie technisch umgesetzt wird
The mainstream is a multimodal large model: a vision encoder splits the image into patches and encodes them as visual tokens, which are fed with text tokens into one Transformer so the language side produces the answer. Alignment training uses image-caption pairs with contrastive learning, and later instruction tuning teaches question answering. Small text and fine detail are preserved by splitting the image into higher-resolution patches.
Repräsentative Produkte
9GPT-4o
2024Ein von Grund auf multimodales Allzweckmodell: Text, Bild und Audio über einen Zugang
Gemini
2023Ein von Grund auf multimodales Allzweckmodell für sehr langen Kontext
Claude
2023Ein allgemeines Dialogmodell, bekannt für langen Kontext und Sicherheitsausrichtung
Qwen
2023Eine Open-Weights-Familie über viele Größen hinweg, mit multimodalen Versionen
Doubao
2023ByteDances allgemeines Dialogmodell und App
Kimi
2023Ein chinesischer Chat-Assistent, bekannt für langen Kontext
GLM
2023Ein chinesisches Allzweckmodell, das mit autoregressivem Lückenfüllen begann
Gemini
2024Ein Chat-Eingang, der Suche, Office-Apps und ein multimodales Modell bündelt
ChatGPT
2022Das Chatfenster, das ein großes Sprachmodell für alle zugänglich machte
Beteiligte Organisationen
Typische Verwendungen
- Accessibility and image narration
- Product and listing description from photos
- Assisted reading of industrial and medical images
- Photo management and content search
Wie sie bewertet wird
- VQA accuracy
- Share answered correctly on annotated question sets
- Captioning CIDEr / SPICE
- Semantic match between captions and references
- Hallucination rate
- Share of captions mentioning objects not in the image
Grenzen und schwierige Punkte
- Exact counting and spatial relations are unreliable; miscounting people and flipping left-right are common
- Small text, chart readings and dense tables in the image are frequently misread
- The model fills gaps with common sense, describing what should be there as if it were seen
Konzepte dahinter
Multimodale Generierung
Ein Modell, das sprechen, zeichnen, sich bewegen — und sogar die 3D-Welt modellieren lernt
Attention-Mechanismus
Jede Position kann direkt auf alle anderen blicken und ihre Aufmerksamkeit nach Relevanz verteilen
Transformer-Architektur
Statt Wort-für-Wort-Stafette ein Raum, in dem alle zugleich sprechen — und weite Abhängigkeiten sind nur einen Schritt entfernt