自然言語処理と大規模言語モデル
分かち書きから Transformer、そして「話せる」大規模モデルへ
OVERVIEW
- 全項目
- 6
- 初級
- 2
- 中級
- 3
- 上級
- 1
Language is cut into tokens, compressed into vectors, and then made to “look at” itself through attention — a loop so plain it seems excessive, yet it lifted next-token prediction all the way to a general assistant. This domain traces that path in order: bag of words, embeddings, RNNs, attention, Transformers, pre-training and alignment.
この領域が答える問い
- Q1
What exactly does a model see when it reads a word?
- Q2
Why does attention beat recurrence?
- Q3
Why does next-token prediction yield conversation?
この領域の項目
- 01トークン化初級モデルは文字ではなくトークンを読む。分割の仕方が能力とコストを静かに決める
- 02単語埋め込み初級「単語」を「座標」に変える。類義語は自然に集まり、意味が初めて足し引きできる
- 03アテンション機構中級各位置が他のすべての位置を直接「見て」、関連度に応じて注意を動的に配分する
- 04Transformer アーキテクチャ中級「一語ずつの伝言」を「全員が同時に話す会議」に置き換え、長距離依存を一跳で届かせる
- 05事前学習とファインチューニング中級大量のラベルなしテキストで言語を学び、少量データで専門化する。現代 AI で最もデータ効率のよい范式
- 06プロンプト設計とアラインメント上級モデルを「役に立ち、正直で、無害」にすることは、ただ大きくするより難しい