सूचना निष्कर्षण और एनईआर
मुक्त पाठ से नाम, स्थान और संबंध निकालना
यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।
यह क्षमता क्या है
Takes natural-language text and returns structured fragments: typed entities (person, organisation, date, amount) plus the relations or events among them. Unlike classification it does not give one label to the whole text but locates spans; unlike schema-driven structured output, the extracted targets come from the text itself rather than a fully prescribed field list.
तकनीकी रूप से कैसे
Early systems used conditional random fields or rule-based sequence labelling, tagging tokens BIO-style; pre-trained encoders with a tagging head then became standard. Relation and event extraction are often framed as entity-pair classification or as generation of triples directly. More recently, large models perform few-shot or zero-shot extraction against a given schema, avoiding per-type annotation.
प्रतिनिधि उत्पाद
5GPT-4o
2024मूल रूप से बहुविध सामान्य मॉडल — पाठ, चित्र और ऑडियो एक ही द्वार से
Qwen
2023अनेक आकारों और बहुविध संस्करणों वाला ओपन-वेट परिवार
ERNIE
2019ज्ञान-संवर्धित प्रीट्रेनिंग से शुरू हुआ चीनी मॉडल, शुरुआती प्रतिनिधि संस्करण
Doubao
2023बाइटडांस का सामान्य संवाद मॉडल और ऐप
GLM
2023ऑटोरेग्रेसिव ब्लैंक-भरने वाले प्रीट्रेनिंग से शुरू हुआ चीनी सामान्य मॉडल
संबंधित संस्थान
सामान्य उपयोग
- Clause and amount extraction from contracts and filings
- Résumé parsing and talent-pool building
- Drug and symptom recognition in clinical notes
- News events and knowledge-graph construction
इसका मूल्यांकन कैसे होता है
- Span-level F1
- A hit requires both boundary and type to be correct
- Relation F1
- Share of triples (head, relation, tail) matched exactly
- Exact-match rate
- Share of records whose fields are all correct
सीमाएँ और कठिनाइयाँ
- Nested and overlapping entities are flattened by token-level tagging schemes
- Cross-sentence coreference is hard; pronouns bind to the wrong antecedent
- Domain terms and novel words are missed when the type was unseen in training
इसके पीछे की अवधारणाएँ
टोकनाइज़ेशन
मॉडल अक्षर नहीं, टोकन पढ़ते हैं — विभाजन का तरीका चुपचाप क्षमता और लागत तय करता है
प्री-ट्रेनिंग और फाइन-ट्यूनिंग
पहले विशाल अलेबल पाठ से भाषा सीखना, फिर थोड़े डेटा से विशेषज्ञ बनना — आधुनिक एआई का सबसे डेटा-कुशल प्रतिमान
पर्यवेक्षित अधिगम
‘प्रश्न–उत्तर’ के जोड़े मॉडल को स्वयं उत्तर देना सिखाते हैं