मुख्य सामग्री पर जाएँ

GPT-4o

मूल रूप से बहुविध सामान्य मॉडल — पाठ, चित्र और ऑडियो एक ही द्वार से

OpenAI मॉडल बंद स्रोत
इनपुटटेक्स्टइमेजऑडियोटेक्स्टऑडियो

यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।

यह क्या है

GPT-4o is a general-purpose model OpenAI released in May 2024; the “o” stands for omni, meaning text, image and audio are handled inside one model. Vision and speech understanding share a single forward pass, and spoken conversation runs at near-real-time latency. Where earlier systems stitched several models into a pipeline, GPT-4o trains multimodality end to end and simplifies the interaction entry point. It is offered through both the API and ChatGPT, and serves as OpenAI’s main multimodal general model.

यह क्यों महत्वपूर्ण है

It moved multimodality from “several models stitched into a pipeline” to “one model end to end” and cut spoken-dialogue latency to near human-conversation levels; general models have since treated native multimodality as the default target.

मुख्य विशिष्टताएँ

Context window
128K tokens
Modality
Text, image, audio in; text, audio out
Released
2024-05
Open weights
No

संबंधित क्षमताएँ

संबंधित अवधारणाएँ

समान उत्पाद

Claude

2023
Anthropic

लंबे संदर्भ और सुरक्षा-संरेखण के लिए जाना जाने वाला सामान्य संवाद मॉडल

मॉडल बंद स्रोत
टेक्स्टइमेजटेक्स्ट

Gemini

2023
Google DeepMind

मूल रूप से बहुविध, अति-लंबे संदर्भ के लिए बना सामान्य मॉडल

मॉडल बंद स्रोत
टेक्स्टइमेजऑडियोवीडियोटेक्स्ट

Llama

2023
Meta AI (FAIR)

खुले भार के रास्ते को मुख्यधारा बनाने वाला मॉडल परिवार

मॉडल खुले वेट
टेक्स्टटेक्स्ट

Grok

2023
xAI

सोशल प्लेटफ़ॉर्म के डेटा से जुड़ा संवाद मॉडल

मॉडल बंद स्रोत
टेक्स्टइमेजटेक्स्ट

Mistral Large

2024
Mistral AI

यूरोपीय ओपन-वेट लैब का फ़्लैगशिप वाणिज्यिक मॉडल

मॉडल बंद स्रोत
टेक्स्टटेक्स्ट

Command R

2024
Cohere

पुनर्प्राप्ति-संवर्धन और टूल उपयोग के लिए बना वाणिज्यिक मॉडल

मॉडल खुले वेट
टेक्स्टटेक्स्ट

Kimi

2023
Moonshot AI

लंबे संदर्भ के लिए जाना जाने वाला चीनी संवाद सहायक

मॉडल बंद स्रोत
टेक्स्टटेक्स्ट