टेक्स्ट जनरेशन
पहले के पाठ को शब्द-दर-शब्द आगे बढ़ाता है
यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।
यह क्षमता क्या है
Given a passage as context, the model writes what comes next. Both input and output are text and no particular transformation is targeted: it can continue a story, finish an explanation or draft an email. It differs from conversation in that no instruction format is required, and from summarisation or translation in that neither compression nor a language switch is the goal — both ends stay in the same modality.
तकनीकी रूप से कैसे
The dominant route is an autoregressive Transformer language model: text is split into tokens, the model predicts the probability of the next token position by position, and training maximises likelihood over the corpus. GPT, Llama and Qwen all follow it, differing in scale, data and post-training recipe. At inference, temperature, top-k and top-p shape diversity, while long outputs reuse cached keys and values to cut cost.
प्रतिनिधि उत्पाद
27GPT-4o
2024मूल रूप से बहुविध सामान्य मॉडल — पाठ, चित्र और ऑडियो एक ही द्वार से
Llama
2023खुले भार के रास्ते को मुख्यधारा बनाने वाला मॉडल परिवार
Qwen
2023अनेक आकारों और बहुविध संस्करणों वाला ओपन-वेट परिवार
DeepSeek-V3
2024671B पैरामीटर वाला ओपन-वेट MoE, प्रति टोकन केवल 37B सक्रिय
Mistral Large
2024यूरोपीय ओपन-वेट लैब का फ़्लैगशिप वाणिज्यिक मॉडल
Gemini
2023मूल रूप से बहुविध, अति-लंबे संदर्भ के लिए बना सामान्य मॉडल
Claude
2023लंबे संदर्भ और सुरक्षा-संरेखण के लिए जाना जाने वाला सामान्य संवाद मॉडल
MiniMax-M
2025हाइब्रिड अटेंशन और दस लाख टोकन संदर्भ वाला ओपन-वेट तर्क मॉडल
DeepSeek-R1
2025तर्क-श्रृंखला पर RL से प्रशिक्षित तर्क मॉडल, भार MIT लाइसेंस के तहत खुले
Apple Intelligence
2024सिस्टम-स्तरीय एआई जो काम को डिवाइस और निजी क्लाउड में बाँटता है
Command R
2024पुनर्प्राप्ति-संवर्धन और टूल उपयोग के लिए बना वाणिज्यिक मॉडल
Jamba
2024स्टेट-स्पेस मॉडल और Transformer को मिलाने वाला ओपन-वेट मॉडल
DBRX
2024Databricks का ओपन-वेट MoE भाषा मॉडल
Gemini
2024खोज, ऑफ़िस और बहुविध मॉडल को एक चैट द्वार में समेटता है
Grok
2023सोशल प्लेटफ़ॉर्म के डेटा से जुड़ा संवाद मॉडल
Yi
2023चीनी-अंग्रेज़ी द्विभाषी ओपन-वेट मॉडल, अति-लंबे संदर्भ संस्करणों के साथ
Kimi
2023लंबे संदर्भ के लिए जाना जाने वाला चीनी संवाद सहायक
Hunyuan
2023ओपन-वेट संस्करणों सहित टेनसेंट का सामान्य मॉडल परिवार
Microsoft Copilot
2023ऑपरेटिंग सिस्टम और ऑफ़िस ऐप्स में बुनी हुई संवादात्मक एआई
Doubao
2023बाइटडांस का सामान्य संवाद मॉडल और ऐप
Phi
2023चुने हुए डेटा से बने छोटे और कुशल मॉडल
Step
2023बहुविधता और ऑन-डिवाइस उपयोग पर लक्षित सामान्य मॉडल परिवार
Baichuan
2023चीनी-भाषा उपयोग के लिए ओपन-वेट सामान्य मॉडल
GLM
2023ऑटोरेग्रेसिव ब्लैंक-भरने वाले प्रीट्रेनिंग से शुरू हुआ चीनी सामान्य मॉडल
ChatGPT
2022वह चैट विंडो जिसने बड़े भाषा मॉडल को हर किसी तक पहुँचाया
Character.AI
2022अपने बनाए पात्रों के साथ लंबी बातचीत वाला रोल-प्ले चैट
ERNIE
2019ज्ञान-संवर्धित प्रीट्रेनिंग से शुरू हुआ चीनी मॉडल, शुरुआती प्रतिनिधि संस्करण
संबंधित संस्थान
सामान्य उपयोग
- First drafts and rewriting
- Drafting email, documents and copy
- Code comments and API docs
- Test data and synthetic corpora
इसका मूल्यांकन कैसे होता है
- Perplexity
- Predictive uncertainty on held-out text; lower is better
- Human-preference Elo
- Preference ranking from pairwise comparison
- ROUGE/BLEU on constrained tasks
- Reference overlap when an answer key exists
सीमाएँ और कठिनाइयाँ
- Fluency is not truth: the model states wrong facts with confidence — hallucination
- Over long outputs, earlier setup drifts and the text contradicts itself
- Poor sampling settings can trap the model in repetitive or degenerate loops
इसके पीछे की अवधारणाएँ
Transformer आर्किटेक्चर
शब्द-दर-शब्द relay की जगह वह कक्ष जहाँ सब एक साथ बोलते हैं, जिससे दूर की निर्भरताएँ एक कदम पर आ जाती हैं
प्री-ट्रेनिंग और फाइन-ट्यूनिंग
पहले विशाल अलेबल पाठ से भाषा सीखना, फिर थोड़े डेटा से विशेषज्ञ बनना — आधुनिक एआई का सबसे डेटा-कुशल प्रतिमान
प्रॉम्प्ट इंजीनियरिंग और अलाइनमेंट
मॉडल को उपयोगी, ईमानदार और हानिरहित बनाना उसे केवल बड़ा करने से कठिन है