Chuyển đến nội dung
Bản đồ AI

Hiểu hình ảnh và hỏi đáp thị giác

Nhìn ảnh và trả lời câu hỏi tự do

Hiểu thị giácCơ bản #17
đầu vàoHình ảnhVăn bảnVăn bản

Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.

NĂNG LỰC NÀY NGHĨA LÀ GÌ

Takes an image, optionally with a text question, and outputs a natural-language description or answer — what the person is doing, where the scene might be. Unlike OCR it targets meaning rather than characters, and unlike classification it answers open-ended questions instead of choosing from a fixed label set.

Làm ra sao về mặt kỹ thuật

The mainstream is a multimodal large model: a vision encoder splits the image into patches and encodes them as visual tokens, which are fed with text tokens into one Transformer so the language side produces the answer. Alignment training uses image-caption pairs with contrastive learning, and later instruction tuning teaches question answering. Small text and fine detail are preserved by splitting the image into higher-resolution patches.

Sản phẩm tiêu biểu

9

GPT-4o

2024
OpenAI

Mô hình đa phương thức gốc, xử lý văn bản, hình ảnh và âm thanh qua một cửa vào

Mô hình Đóng
Văn bảnHình ảnhÂm thanhVăn bảnÂm thanh

Gemini

2023
Google DeepMind

Mô hình đa phương thức gốc, xử lý ngữ cảnh siêu dài

Mô hình Đóng
Văn bảnHình ảnhÂm thanhVideoVăn bản

Claude

2023
Anthropic

Mô hình hội thoại phổ thông nổi bật với ngữ cảnh dài và căn chỉnh an toàn

Mô hình Đóng
Văn bảnHình ảnhVăn bản

Qwen

2023
Alibaba (Qwen)

Dòng trọng số mở với nhiều kích cỡ và bản đa phương thức

Mô hình Trọng số mở
Văn bảnHình ảnhVăn bản

Doubao

2023
ByteDance (Seed)

Mô hình hội thoại phổ thông và ứng dụng của ByteDance

Mô hình Đóng
Văn bảnHình ảnhVăn bản

Kimi

2023
Moonshot AI

Trợ lý hội thoại tiếng Trung nổi bật với xử lý ngữ cảnh dài

Mô hình Đóng
Văn bảnVăn bản

GLM

2023
Zhipu AI

Mô hình phổ thông tiếng Trung khởi đầu bằng tiền huấn luyện điền khuyết tự hồi quy

Mô hình Đóng
Văn bảnVăn bản

Gemini

2024
Google DeepMind

Một cửa vào trò chuyện gom tìm kiếm, ứng dụng văn phòng và mô hình đa phương thức

Ứng dụng Freemium
Văn bảnHình ảnhÂm thanhVăn bảnHình ảnhÂm thanh

ChatGPT

2022
OpenAI

Cửa sổ trò chuyện đưa mô hình ngôn ngữ lớn đến tay mọi người

Ứng dụng Freemium
Văn bảnHình ảnhÂm thanhVăn bảnHình ảnhÂm thanh

Tổ chức liên quan

Cách dùng tiêu biểu

  • Accessibility and image narration
  • Product and listing description from photos
  • Assisted reading of industrial and medical images
  • Photo management and content search

Đánh giá nó tốt hay không thế nào

VQA accuracy
Share answered correctly on annotated question sets
Captioning CIDEr / SPICE
Semantic match between captions and references
Hallucination rate
Share of captions mentioning objects not in the image

Ranh giới và điểm khó

  • Exact counting and spatial relations are unreliable; miscounting people and flipping left-right are common
  • Small text, chart readings and dense tables in the image are frequently misread
  • The model fills gaps with common sense, describing what should be there as if it were seen

Các khái niệm đằng sau