本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
これは何か
GPT-4o is a general-purpose model OpenAI released in May 2024; the “o” stands for omni, meaning text, image and audio are handled inside one model. Vision and speech understanding share a single forward pass, and spoken conversation runs at near-real-time latency. Where earlier systems stitched several models into a pipeline, GPT-4o trains multimodality end to end and simplifies the interaction entry point. It is offered through both the API and ChatGPT, and serves as OpenAI’s main multimodal general model.
なぜ覚えておく価値があるか
It moved multimodality from “several models stitched into a pipeline” to “one model end to end” and cut spoken-dialogue latency to near human-conversation levels; general models have since treated native multimodality as the default target.
主な仕様
- Context window
- 128K tokens
- Modality
- Text, image, audio in; text, audio out
- Released
- 2024-05
- Open weights
- No
対応する能力
関連する概念
同種の製品
Claude
2023長い文脈と安全性の調整で知られる汎用対話モデル
Gemini
2023ネイティブにマルチモーダルで、超長文脈を扱う汎用モデル
Llama
2023オープンウェイト路線を主流にした汎用モデル群
Grok
2023ソーシャルプラットフォームのデータと結びついた対話モデル
Mistral Large
2024欧州のオープンウェイト研究所による商用フラッグシップモデル
Command R
2024検索拡張とツール呼び出しのために設計された商用モデル
Kimi
2023長い文脈の処理を得意とする中国語の対話アシスタント