정보 추출과 개체명 인식
자유 텍스트에서 인명·지명·관계를 뽑아낸다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
이 능력이 뜻하는 것
Takes natural-language text and returns structured fragments: typed entities (person, organisation, date, amount) plus the relations or events among them. Unlike classification it does not give one label to the whole text but locates spans; unlike schema-driven structured output, the extracted targets come from the text itself rather than a fully prescribed field list.
기술적으로 구현하는 방법
Early systems used conditional random fields or rule-based sequence labelling, tagging tokens BIO-style; pre-trained encoders with a tagging head then became standard. Relation and event extraction are often framed as entity-pair classification or as generation of triples directly. More recently, large models perform few-shot or zero-shot extraction against a given schema, avoiding per-type annotation.
대표 제품
5GPT-4o
2024네이티브 멀티모달 범용 모델. 텍스트·이미지·오디오를 한 창구에서 다룬다
Qwen
2023여러 규모와 멀티모달 버전을 아우르는 오픈웨이트 모델 계열
ERNIE
2019지식 강화 사전학습에서 출발한 중국어 모델, 초기 대표 버전
Doubao
2023바이트댄스의 범용 대화 모델과 앱
GLM
2023자기회귀 빈칸 채우기 사전학습에서 출발한 중국어 범용 모델
관련 기관
대표적 용도
- Clause and amount extraction from contracts and filings
- Résumé parsing and talent-pool building
- Drug and symptom recognition in clinical notes
- News events and knowledge-graph construction
성능을 평가하는 방법
- Span-level F1
- A hit requires both boundary and type to be correct
- Relation F1
- Share of triples (head, relation, tail) matched exactly
- Exact-match rate
- Share of records whose fields are all correct
경계와 난점
- Nested and overlapping entities are flattened by token-level tagging schemes
- Cross-sentence coreference is hard; pronouns bind to the wrong antecedent
- Domain terms and novel words are missed when the type was unseen in training