Hiểu hình ảnh và hỏi đáp thị giác
Nhìn ảnh và trả lời câu hỏi tự do
Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.
NĂNG LỰC NÀY NGHĨA LÀ GÌ
Takes an image, optionally with a text question, and outputs a natural-language description or answer — what the person is doing, where the scene might be. Unlike OCR it targets meaning rather than characters, and unlike classification it answers open-ended questions instead of choosing from a fixed label set.
Làm ra sao về mặt kỹ thuật
The mainstream is a multimodal large model: a vision encoder splits the image into patches and encodes them as visual tokens, which are fed with text tokens into one Transformer so the language side produces the answer. Alignment training uses image-caption pairs with contrastive learning, and later instruction tuning teaches question answering. Small text and fine detail are preserved by splitting the image into higher-resolution patches.
Sản phẩm tiêu biểu
9GPT-4o
2024Mô hình đa phương thức gốc, xử lý văn bản, hình ảnh và âm thanh qua một cửa vào
Gemini
2023Mô hình đa phương thức gốc, xử lý ngữ cảnh siêu dài
Claude
2023Mô hình hội thoại phổ thông nổi bật với ngữ cảnh dài và căn chỉnh an toàn
Qwen
2023Dòng trọng số mở với nhiều kích cỡ và bản đa phương thức
Doubao
2023Mô hình hội thoại phổ thông và ứng dụng của ByteDance
Kimi
2023Trợ lý hội thoại tiếng Trung nổi bật với xử lý ngữ cảnh dài
GLM
2023Mô hình phổ thông tiếng Trung khởi đầu bằng tiền huấn luyện điền khuyết tự hồi quy
Gemini
2024Một cửa vào trò chuyện gom tìm kiếm, ứng dụng văn phòng và mô hình đa phương thức
ChatGPT
2022Cửa sổ trò chuyện đưa mô hình ngôn ngữ lớn đến tay mọi người
Tổ chức liên quan
Cách dùng tiêu biểu
- Accessibility and image narration
- Product and listing description from photos
- Assisted reading of industrial and medical images
- Photo management and content search
Đánh giá nó tốt hay không thế nào
- VQA accuracy
- Share answered correctly on annotated question sets
- Captioning CIDEr / SPICE
- Semantic match between captions and references
- Hallucination rate
- Share of captions mentioning objects not in the image
Ranh giới và điểm khó
- Exact counting and spatial relations are unreliable; miscounting people and flipping left-right are common
- Small text, chart readings and dense tables in the image are frequently misread
- The model fills gaps with common sense, describing what should be there as if it were seen
Các khái niệm đằng sau
Sinh đa phương thức
Một mô hình học nói, vẽ, chuyển động — thậm chí mô hình hóa thế giới ba chiều
Cơ chế chú ý (Attention)
Mọi vị trí đều có thể nhìn thẳng vào mọi vị trí khác và phân bổ chú ý theo mức liên quan
Kiến trúc Transformer
Thay cách truyền từng từ bằng một phòng họp nơi mọi từ cùng lên tiếng, để phụ thuộc xa chỉ còn cách một bước