استخراج المعلومات وتعرّف الكيانات
استخراج الأسماء والأماكن والعلاقات من نص حر
يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
ما الذي تعنيه هذه القدرة
Takes natural-language text and returns structured fragments: typed entities (person, organisation, date, amount) plus the relations or events among them. Unlike classification it does not give one label to the whole text but locates spans; unlike schema-driven structured output, the extracted targets come from the text itself rather than a fully prescribed field list.
كيف تُنفَّذ تقنيًا
Early systems used conditional random fields or rule-based sequence labelling, tagging tokens BIO-style; pre-trained encoders with a tagging head then became standard. Relation and event extraction are often framed as entity-pair classification or as generation of triples directly. More recently, large models perform few-shot or zero-shot extraction against a given schema, avoiding per-type annotation.
منتجات تمثيلية
5GPT-4o
2024نموذج عام متعدد الوسائط بطبيعته، يجمع النص والصورة والصوت في مدخل واحد
Qwen
2023عائلة بأوزان مفتوحة تغطي أحجامًا متعددة، مع نسخ متعددة الوسائط
ERNIE
2019نموذج صيني بدأ بالتدريب المسبق المعزّز بالمعرفة، نسخة مبكرة ممثلة
Doubao
2023نموذج المحادثة العام وتطبيق بايت دانس
GLM
2023نموذج عام صيني بدأ بالتدريب المسبق لملء الفراغات ذاتيًا
المؤسسات ذات الصلة
الاستخدامات الشائعة
- Clause and amount extraction from contracts and filings
- Résumé parsing and talent-pool building
- Drug and symptom recognition in clinical notes
- News events and knowledge-graph construction
كيف يُقاس مدى جودتها
- Span-level F1
- A hit requires both boundary and type to be correct
- Relation F1
- Share of triples (head, relation, tail) matched exactly
- Exact-match rate
- Share of records whose fields are all correct
الحدود والصعوبات
- Nested and overlapping entities are flattened by token-level tagging schemes
- Cross-sentence coreference is hard; pronouns bind to the wrong antecedent
- Domain terms and novel words are missed when the type was unseen in training
المفاهيم الكامنة وراءها
الترميز إلى رموز (Tokenization)
لا تقرأ النماذج الحروف بل الرموز، وطريقة التقسيم تحدّد القدرة والكلفة بصمت
التدريب المسبق والضبط الدقيق
تعلّم اللغة أولاً من نصوص ضخمة بلا وسوم ثم التخصص ببيانات قليلة — أكثر النماذج كفاءةً في البيانات
التعلّم بالإشراف
تُعلِّم أزواجُ السؤال والجواب النموذجَ أن يجيب من تلقاء نفسه