이미지 이해와 시각 질의응답
이미지를 보고 자유로운 질문에 답한다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
이 능력이 뜻하는 것
Takes an image, optionally with a text question, and outputs a natural-language description or answer — what the person is doing, where the scene might be. Unlike OCR it targets meaning rather than characters, and unlike classification it answers open-ended questions instead of choosing from a fixed label set.
기술적으로 구현하는 방법
The mainstream is a multimodal large model: a vision encoder splits the image into patches and encodes them as visual tokens, which are fed with text tokens into one Transformer so the language side produces the answer. Alignment training uses image-caption pairs with contrastive learning, and later instruction tuning teaches question answering. Small text and fine detail are preserved by splitting the image into higher-resolution patches.
대표 제품
9GPT-4o
2024네이티브 멀티모달 범용 모델. 텍스트·이미지·오디오를 한 창구에서 다룬다
Gemini
2023네이티브 멀티모달에 초장문 문맥을 다루는 범용 모델
Claude
2023긴 문맥과 안전 정렬로 알려진 범용 대화 모델
Qwen
2023여러 규모와 멀티모달 버전을 아우르는 오픈웨이트 모델 계열
Doubao
2023바이트댄스의 범용 대화 모델과 앱
Kimi
2023긴 문맥 처리에 강한 중국어 대화 어시스턴트
GLM
2023자기회귀 빈칸 채우기 사전학습에서 출발한 중국어 범용 모델
Gemini
2024검색·오피스·멀티모달 모델을 하나의 대화 창구로 모았다
ChatGPT
2022대규모 언어 모델을 누구나 쓰는 대화 창으로 만들었다
관련 기관
대표적 용도
- Accessibility and image narration
- Product and listing description from photos
- Assisted reading of industrial and medical images
- Photo management and content search
성능을 평가하는 방법
- VQA accuracy
- Share answered correctly on annotated question sets
- Captioning CIDEr / SPICE
- Semantic match between captions and references
- Hallucination rate
- Share of captions mentioning objects not in the image
경계와 난점
- Exact counting and spatial relations are unreliable; miscounting people and flipping left-right are common
- Small text, chart readings and dense tables in the image are frequently misread
- The model fills gaps with common sense, describing what should be there as if it were seen