Extração de informação e NER
Extrair nomes, lugares e relações de texto livre
O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.
O QUE ESTA CAPACIDADE SIGNIFICA
Takes natural-language text and returns structured fragments: typed entities (person, organisation, date, amount) plus the relations or events among them. Unlike classification it does not give one label to the whole text but locates spans; unlike schema-driven structured output, the extracted targets come from the text itself rather than a fully prescribed field list.
Como é feita tecnicamente
Early systems used conditional random fields or rule-based sequence labelling, tagging tokens BIO-style; pre-trained encoders with a tagging head then became standard. Relation and event extraction are often framed as entity-pair classification or as generation of triples directly. More recently, large models perform few-shot or zero-shot extraction against a given schema, avoiding per-type annotation.
Produtos representativos
5GPT-4o
2024Um modelo geral nativamente multimodal: texto, imagem e áudio por uma única porta
Qwen
2023Uma família de pesos abertos com muitos tamanhos e versões multimodais
ERNIE
2019Um modelo chinês que começou com pré-treinamento enriquecido por conhecimento, versão inicial representativa
Doubao
2023O modelo de conversa geral e o aplicativo da ByteDance
GLM
2023Um modelo geral chinês que começou com pré-treinamento de preenchimento de lacunas autorregressivo
Organizações relacionadas
Usos típicos
- Clause and amount extraction from contracts and filings
- Résumé parsing and talent-pool building
- Drug and symptom recognition in clinical notes
- News events and knowledge-graph construction
Como avaliar se funciona bem
- Span-level F1
- A hit requires both boundary and type to be correct
- Relation F1
- Share of triples (head, relation, tail) matched exactly
- Exact-match rate
- Share of records whose fields are all correct
Limites e dificuldades
- Nested and overlapping entities are flattened by token-level tagging schemes
- Cross-sentence coreference is hard; pronouns bind to the wrong antecedent
- Domain terms and novel words are missed when the type was unseen in training
Conceitos por trás
Tokenização
Os modelos não leem caracteres, leem tokens; e como você divide define silenciosamente a capacidade e o custo
Pré-treinamento e ajuste fino
Aprender a língua primeiro com texto não rotulado e depois especializar com poucos dados: o paradigma mais eficiente em dados da IA moderna
Aprendizagem supervisionada
Pares de pergunta e resposta ensinam o modelo a responder sozinho