본문으로 건너뛰기
AI 도감

텍스트 요약

긴 글을 더 짧고 정확한 형태로 줄인다

언어와 지식입문 #04
입력텍스트텍스트

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

이 능력이 뜻하는 것

Takes a long document and returns a shorter version that keeps the key information. It splits into extractive (selecting sentences) and abstractive (rewriting in new words) styles, the latter now dominant. Unlike free generation it has an explicit compression target, and unlike question answering it is not aimed at one query but should cover the whole thread.

기술적으로 구현하는 방법

The classic approach is a sequence-to-sequence attention model producing abstractive summaries, with rewriting ability coming from large-scale pre-training. Long documents are usually handled hierarchically or by a map-reduce scheme that summarises chunks and then combines them. Controlled summarisation passes length, angle or audience constraints through the prompt or light fine-tuning to steer style.

대표 제품

8

관련 기관

대표적 용도

  • Quick reads of news and reports
  • Meeting and call minutes
  • Literature triage and paper skims
  • Rolling summaries of tickets and email

성능을 평가하는 방법

ROUGE
N-gram overlap with reference summaries
BERTScore
Semantic-embedding similarity, more tolerant than literal overlap
Factual consistency
Share of statements conflicting with the source, as in FactCC-style evaluation

경계와 난점

  • Middle sections of long documents are often dropped; models favour the start and end
  • Abstractive models splice facts from different sentences into claims the source never made
  • On ambiguous or multi-sided texts, minority views are often reported as the majority

뒤에 있는 개념