자연어 처리와 대규모 언어 모델
토큰화에서 트랜스포머, 그리고 말하는 대형 모델까지
OVERVIEW
- 전체 항목
- 6
- 입문
- 2
- 중급
- 3
- 고급
- 1
Language is cut into tokens, compressed into vectors, and then made to “look at” itself through attention — a loop so plain it seems excessive, yet it lifted next-token prediction all the way to a general assistant. This domain traces that path in order: bag of words, embeddings, RNNs, attention, Transformers, pre-training and alignment.
이 영역이 답하는 질문
- Q1
What exactly does a model see when it reads a word?
- Q2
Why does attention beat recurrence?
- Q3
Why does next-token prediction yield conversation?
이 영역의 항목
- 01토큰화입문모델은 글자가 아니라 토큰을 읽는다. 분할 방식이 성능과 비용을 조용히 결정한다
- 02단어 임베딩입문단어를 좌표로 바꾸면 유의어가 스스로 모이고, 의미를 처음으로 더하고 뺄 수 있다
- 03어텐션 메커니즘중급모든 위치가 다른 모든 위치를 직접 보고 관련도에 따라 주의를 동적으로 배분한다
- 04Transformer 아키텍처중급한 단어씩 전달하는 대신 모두가 동시에 말하는 회의로 바꾸어, 멀리 떨어진 의존성을 한 번에 연결한다
- 05사전학습과 미세조정중급방대한 무표지 텍스트로 언어를 먼저 익히고 소량 데이터로 전문화한다 — 현대 AI에서 가장 데이터 효율적인 패러다임
- 06프롬프트 엔지니어링과 정렬고급모델을 유용하고 정직하며 무해하게 만드는 일은 단순히 더 크게 만드는 것보다 어렵다