Tokenização
Os modelos não leem caracteres, leem tokens; e como você divide define silenciosamente a capacidade e o custo
O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.
DEFINIÇÃO
Tokenization is the process of splitting a continuous piece of text into the smallest units a model can handle — tokens. Each token maps to an integer index in a fixed vocabulary, and the model reads only the sequence of those indices, never the characters themselves. The granularity may be characters, whole words, or the subword units in between.
Intuição
Feeding text to a model is a bit like moving house: a standard sofa (a common word) goes out in one piece, while an odd-shaped item (a rare word or unusual spelling) must be taken apart and carried in several trips. The finer the split, the more trips, and the more you pay. A sentence lands in fewer slots in a language the vocabulary covers well, and in more fragments otherwise — more tokens means a slower and more expensive model.
From raw text to embeddings: the tokenizer is the only entrance and the boundary of what the model can see
Under one BPE vocabulary, the average characters each token covers per language: poorer coverage means the same content costs more tokens
Como funciona
- 01
Normalise and pre-split
First unify the Unicode form and settle rules for case, whitespace and punctuation, cutting the text into candidate pieces. This step decides whether punctuation, digits and spaces become standalone tokens, and it shapes everything that follows.
- 02
Learn subwords statistically (BPE / SentencePiece)
BPE starts from single characters and repeatedly merges the most frequent adjacent symbol pair in the corpus into a new unit, until the vocabulary reaches its target size. SentencePiece runs directly on the raw byte or character stream without pre-splitting on spaces, which suits Chinese and Japanese far better.
- 03
Map to integer IDs and embeddings
Each token looks up its index in the vocabulary, then passes through an embedding layer into the network. Once the vocabulary is fixed, indices and embeddings are bound together: swapping the tokenizer swaps the entire input interface, so the weights must be retrained.
Onde é usado
- LLM input and output: context limits and pricing are counted in tokens, not characters
- Multilingual and code modelling: subwords let a model spell out unseen words, rare inflections and identifiers
- Retrieval and indexing: token counts drive index size and retrieval cost
- Multimodal extension: audio frames and image patches are also chopped and encoded into token sequences
Equívocos comuns
- A token is neither a character nor a word. One Chinese token often covers less than a single character, while an English token may carry a leading space — so estimating token counts from character counts is markedly unreliable.
- Different models use different tokenizers, and the same text can differ by more than a factor of two in token count. When comparing price or context usage across models, always measure with each one’s own tokenizer.
- The tokenizer defines the model’s input boundary and is usually frozen before pretraining. Replacing it later means retraining every embedding and weight — a very expensive move.
Termos-chave
- Vocabulary
- The fixed set of all tokens and their indices
- BPE
- Byte-Pair Encoding: bottom-up merging of frequent symbol pairs
- SentencePiece
- A subword toolkit that runs directly on the character/byte stream
- Out-of-vocabulary (OOV)
- A word absent from the vocabulary, spelled out from subwords