본문으로 건너뛰기
AI 도감

이미지 이해와 시각 질의응답

이미지를 보고 자유로운 질문에 답한다

시각 이해입문 #17
입력이미지텍스트텍스트

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

이 능력이 뜻하는 것

Takes an image, optionally with a text question, and outputs a natural-language description or answer — what the person is doing, where the scene might be. Unlike OCR it targets meaning rather than characters, and unlike classification it answers open-ended questions instead of choosing from a fixed label set.

기술적으로 구현하는 방법

The mainstream is a multimodal large model: a vision encoder splits the image into patches and encodes them as visual tokens, which are fed with text tokens into one Transformer so the language side produces the answer. Alignment training uses image-caption pairs with contrastive learning, and later instruction tuning teaches question answering. Small text and fine detail are preserved by splitting the image into higher-resolution patches.

대표 제품

9

관련 기관

대표적 용도

  • Accessibility and image narration
  • Product and listing description from photos
  • Assisted reading of industrial and medical images
  • Photo management and content search

성능을 평가하는 방법

VQA accuracy
Share answered correctly on annotated question sets
Captioning CIDEr / SPICE
Semantic match between captions and references
Hallucination rate
Share of captions mentioning objects not in the image

경계와 난점

  • Exact counting and spatial relations are unreliable; miscounting people and flipping left-right are common
  • Small text, chart readings and dense tables in the image are frequently misread
  • The model fills gaps with common sense, describing what should be there as if it were seen

뒤에 있는 개념