मुख्य सामग्री पर जाएँ

सूचना निष्कर्षण और एनईआर

मुक्त पाठ से नाम, स्थान और संबंध निकालना

भाषा और ज्ञानमध्यवर्ती #06
इनपुटटेक्स्टटेक्स्ट

यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।

यह क्षमता क्या है

Takes natural-language text and returns structured fragments: typed entities (person, organisation, date, amount) plus the relations or events among them. Unlike classification it does not give one label to the whole text but locates spans; unlike schema-driven structured output, the extracted targets come from the text itself rather than a fully prescribed field list.

तकनीकी रूप से कैसे

Early systems used conditional random fields or rule-based sequence labelling, tagging tokens BIO-style; pre-trained encoders with a tagging head then became standard. Relation and event extraction are often framed as entity-pair classification or as generation of triples directly. More recently, large models perform few-shot or zero-shot extraction against a given schema, avoiding per-type annotation.

प्रतिनिधि उत्पाद

5

संबंधित संस्थान

सामान्य उपयोग

  • Clause and amount extraction from contracts and filings
  • Résumé parsing and talent-pool building
  • Drug and symptom recognition in clinical notes
  • News events and knowledge-graph construction

इसका मूल्यांकन कैसे होता है

Span-level F1
A hit requires both boundary and type to be correct
Relation F1
Share of triples (head, relation, tail) matched exactly
Exact-match rate
Share of records whose fields are all correct

सीमाएँ और कठिनाइयाँ

  • Nested and overlapping entities are flattened by token-level tagging schemes
  • Cross-sentence coreference is hard; pronouns bind to the wrong antecedent
  • Domain terms and novel words are missed when the type was unseen in training

इसके पीछे की अवधारणाएँ