本文へスキップ
AI図鑑

GLOSSARY

用語集

各項目に散らばる重要用語をひとつの索引にまとめました。知らない言葉はここで引けます。

全 193 語

3 1

3D Gaussian Splatting

Representing a scene with many 3D Gaussian ellipsoids for fast rendering

出典マルチモーダル生成

A 9

Accelerator

A high-throughput parallel unit such as a GPU or TPU

出典学習・推論インフラストラクチャ
Action A

What the agent can do; either discrete or continuous

出典マルコフ決定過程
Action value Q(s, a)

Expected discounted return after forcing the first action to be a

出典価値関数と Q 学習
Activation function

A function that applies a nonlinear transform to the weighted sum

出典ニューロンとパーセプトロン
Actor / Critic

The policy network and the value network: one acts, one scores

出典方策勾配法
Advantage A(s, a)

How much better an action is than the average at that state

出典方策勾配法
Alignment

Making model behaviour match human intent and values

出典安全性・アラインメント・プロンプトインジェクション
Anomaly detection

Finding the few samples that deviate from the bulk distribution

出典教師なし学習
Automatic differentiation

Letting a framework compute exact gradients automatically, not by numerical approximation

出典誤差逆伝播法

B 8

Batch size

How many samples estimate the gradient per step

出典勾配と勾配降下法
Bias

How far the model’s average prediction departs from the true regularity

出典バイアス・バリアンス分解
Bias

A learnable offset applied to the threshold

出典ニューロンとパーセプトロン
Bit depth

How many bits encode each channel; 8 bits give 256 levels

出典画像のデジタル表現
Bottleneck

The low-dimensional layer holding the latent code, limiting its bandwidth

出典オートエンコーダと変分オートエンコーダ
Bounding box

A rectangle represented as (x, y, w, h) or corner points

出典物体検出
BPE

Byte-Pair Encoding: bottom-up merging of frequent symbol pairs

出典トークン化
BPTT

Backpropagation through time after unrolling

出典リカレントニューラルネットワーク

C 20

Catastrophic forgetting

Rapid loss of old abilities while learning a new task

出典事前学習とファインチューニング
Chain rule

The derivative of a composition is the product of the local derivatives

出典誤差逆伝播法
Chain-of-thought (CoT)

Making the model write out intermediate reasoning steps

出典プロンプト設計とアラインメント
Channel

A distinct measurement at the same location, such as R/G/B or alpha

出典画像のデジタル表現
Chunking

Splitting long documents into retrievable pieces

出典検索拡張生成
Classifier-free guidance

Extrapolating between conditional and unconditional predictions to control prompt fidelity

出典拡散モデル
Clustering

Grouping samples by similarity (k-means, hierarchical clustering)

出典教師なし学習
Colour space

A coordinate system for colour values, such as sRGB, HSV or Lab

出典画像のデジタル表現
Computation graph

A computation expressed as nodes and directed edges over which derivatives propagate

出典誤差逆伝播法
Confusion matrix

A cross-tabulation of true versus predicted classes

出典モデル評価と交差検証
Continuous batching

Re-forming the batch every step to keep the GPU busy

出典推論最適化とサービング
Contrastive learning

Learning representations by pulling positives together and pushing negatives apart

出典自己教師あり視覚とマルチモーダル対照学習
Contrastive loss

A loss that pulls same-class embeddings together and pushes different-class ones apart

出典損失関数
ControlNet

A bypass network guiding structure from a condition map

出典潜在空間拡散と条件制御
Cosine similarity

The alignment of two vector directions, from −1 to 1

出典単語埋め込み
Cross-attention

Query from one sequence, Key/Value from another

出典アテンション機構
Cross-attention

The attention mechanism letting image features query text vectors

出典潜在空間拡散と条件制御
Cross-entropy

The information needed to encode data from P using distribution Q

出典エントロピーと情報理論
Cross-entropy

Negative log-probability of the correct class; the default classification loss

出典損失関数
Cross-entropy loss

The standard objective for classification training

出典画像分類

D 12

DDIM

Deterministic sampling achieving comparable quality in a few dozen steps

出典拡散モデル
DDPM

Discrete Markov diffusion, typically needing a thousand sampling steps

出典拡散モデル
Degradation problem

Deeper networks with higher training error, and not from overfitting

出典正規化と残差接続
Density estimation

Estimating the probability distribution the data follows

出典教師なし学習
Dimension

The number of entries in a vector

出典ベクトルとベクトル空間
Dimensionality reduction

Compressing high-dimensional data to fewer dimensions while preserving structure (PCA, t-SNE, UMAP)

出典教師なし学習
Discount factor γ

Between 0 and 1; how much future rewards are valued

出典マルコフ決定過程
Discriminator

The network judging real versus fake, serving as the loss

出典敵対的生成ネットワーク
Double descent

The modern counterexample where test error falls again past the interpolation point

出典バイアス・バリアンス分解
DPO

Direct preference optimisation without an explicit reward model

出典プロンプト設計とアラインメント
Dropout

Randomly silencing units during training to prevent co-adaptation

出典過学習と正則化
Dying ReLU

A neuron stuck in the negative region with zero gradient, no longer updating

出典活性化関数

E 10

Early stopping

Halting training before validation loss turns upward

出典過学習と正則化
Eigenvector / eigenvalue

A vector whose direction is unchanged by the map, and the factor by which it is scaled

出典行列演算と線形写像
ELBO

A lower bound on the log-likelihood: the reconstruction term minus the KL term; a VAE’s actual objective

出典オートエンコーダと変分オートエンコーダ
Embedding

The layer, or its output, that maps a discrete object into a continuous vector

出典ベクトルとベクトル空間
Embedding model

A model that encodes text into vectors

出典検索拡張生成
Empirical risk

The model’s average loss on the training samples

出典教師あり学習
Equivariance

When the input shifts, the output shifts accordingly rather than changing

出典畳み込みニューラルネットワーク
Evidence

The total probability of the data across all hypotheses; it normalises the result

出典ベイズの定理
Experience replay

Store past transitions and sample randomly to break correlation

出典深層強化学習
Explicit density

A model that writes down or approximates p(x), e.g. autoregressive or diffusion

出典生成モデルの全体像

F 3

F1

The harmonic mean of precision and recall

出典モデル評価と交差検証
FID

Fréchet distance between generated and real distributions in Inception feature space; lower is better

出典生成モデルの全体像
Function calling

The model emitting structured arguments to invoke an external function

出典エージェントとツール利用

G 5

Gating

Using 0–1 coefficients from Sigmoid to control how much information passes

出典リカレントニューラルネットワーク
Generalisation

Performance on data the model has not seen

出典過学習と正則化
Generator

The network mapping noise to samples

出典敵対的生成ネットワーク
Gradient flow

The magnitude and stability of gradients as they propagate layer by layer

出典誤差逆伝播法
Guardrail

Checks and constraints bounding what an agent may do

出典エージェントとツール利用

H 3

Hidden state

A continuously updated "summary so far" vector

出典リカレントニューラルネットワーク
Hinge loss

Requires the correct class to win by a margin; the heart of the SVM

出典損失関数
Hypothesis space

The set of all functions the model can represent

出典教師あり学習

I 9

Identity shortcut

The path in a residual connection that adds the input straight back to the output

出典正規化と残差接続
Implicit density

A model that offers only a sampler, not a probability, e.g. a GAN

出典生成モデルの全体像
In-context learning

Solving a task from prompt examples without updating parameters

出典プロンプト設計とアラインメント
InfoNCE

The standard contrastive loss; essentially a multi-class cross-entropy

出典自己教師あり視覚とマルチモーダル対照学習
Inner product

Element-wise product summed over entries; the numerator of cosine similarity

出典ベクトルとベクトル空間
Input x

The feature vector fed to the model

出典教師あり学習
Internal covariate shift

The shifting distribution of inputs to later layers during training

出典正規化と残差接続
IoU

The ratio of the intersection to the union of two boxes

出典物体検出
Irreducible error

The unavoidable error floor caused by label noise

出典バイアス・バリアンス分解

J 1

K 7

Kernel / filter

A set of learnable weights that slides over the input

出典畳み込みニューラルネットワーク
Kernel / filter

The small weight matrix that is learned

出典畳み込み演算
KL divergence

Cross-entropy minus true entropy; non-negative and asymmetric

出典エントロピーと情報理論
KL divergence

Measures how far the encoded distribution deviates from a standard normal; acts as a regulariser

出典オートエンコーダと変分オートエンコーダ
KL penalty

Penalises divergence from the reference policy to prevent degeneration

出典人間のフィードバックによる強化学習
Knowledge distillation

Training a small model on a large model’s soft outputs

出典モデル圧縮
KV cache

Caching past tokens’ keys and values to avoid recomputation

出典推論最適化とサービング

L 12

Label y

The correct output for each sample; the source of supervision

出典教師あり学習
Latent space

The low-dimensional representation space produced by the autoencoder

出典潜在空間拡散と条件制御
Learning rate η

How far each step moves

出典勾配と勾配降下法
Likelihood

The probability of the observed data under given parameters

出典確率と確率分布
Likelihood

The probability of observed data given that the hypothesis is true

出典ベイズの定理
Linearly separable

A hyperplane exists that separates the two classes perfectly

出典ニューロンとパーセプトロン
Log-derivative trick

Turns the gradient of an expectation into a weighted sum of log-probabilities

出典方策勾配法
Long-range dependency

Influence between elements far apart in a sequence

出典リカレントニューラルネットワーク
LoRA

Low-rank adapters training only a tiny number of new parameters

出典事前学習とファインチューニング
LoRA

Low-rank adaptation increments for low-cost customisation

出典潜在空間拡散と条件制御
Loss surface

The high-dimensional terrain of loss values over parameter space

出典勾配と勾配降下法
Low-rank factorisation

Approximating a large matrix by a product of two smaller ones

出典モデル圧縮

M 9

mAP

Mean average precision across classes and IoU thresholds

出典物体検出
Masked language modelling

Hide random words and recover them, a bidirectional objective

出典事前学習とファインチューニング
MCTS

An algorithm that evaluates moves via sampled rollouts to guide search

出典深層強化学習
Mean squared error (MSE)

The average squared difference between prediction and label; the default regression loss

出典損失関数
mIoU

The mean of per-class IoU, the primary segmentation metric

出典セマンティックセグメンテーション
Mode collapse

When a generator covers only a few modes of the data distribution

出典生成モデルの全体像
Mode collapse

The generator covers few modes and loses diversity

出典敵対的生成ネットワーク
Multi-armed bandit

The simplest sequential model: unknown reward distributions, one pull per round

出典探索と活用
Multi-head attention

Several attentions in parallel, each learning a different focus

出典アテンション機構

N 5

Negative sampling

Replacing full-vocabulary softmax with a few random negatives

出典単語埋め込み
NeRF

A neural network representing a scene’s radiance field for novel-view synthesis

出典マルチモーダル生成
NMS

Non-maximum suppression, removing duplicate boxes

出典物体検出
Noise schedule

The timetable of noise added per step, described by βₜ or ᾱₜ

出典拡散モデル
Norm

A function measuring a vector’s "length"; L2 is the common choice

出典ベクトルとベクトル空間

O 3

Off-policy

The behaviour policy may differ from the policy being learned

出典価値関数と Q 学習
One-hot

A sparse vector with a single 1; distinct words are fully orthogonal

出典単語埋め込み
Out-of-vocabulary (OOV)

A word absent from the vocabulary, spelled out from subwords

出典トークン化

P 16

Padding

Adding zeros at the border to control output size

出典畳み込み演算
Perplexity

The exponential of the cross-entropy; the effective number of options the model hesitates among per step

出典エントロピーと情報理論
Pipeline parallelism

Placing different layers on different devices and filling bubbles with micro-batches

出典学習・推論インフラストラクチャ
Pixel

The smallest sampling unit of an image, carrying one or more channel values

出典画像のデジタル表現
Policy π

A mapping from states to actions, or to a distribution over actions

出典マルコフ決定過程
Positional encoding

An explicit order signal, sinusoidal or RoPE

出典Transformer アーキテクチャ
Posterior

The updated degree of belief after incorporating the evidence

出典ベイズの定理
Pre-activation

A layout placing normalisation before the convolution, which trains more stably

出典正規化と残差接続
Pre-LN

Placing layer norm before each sublayer for stability

出典Transformer アーキテクチャ
Precision & recall

Precision asks how many alerts are real; recall asks how many real cases were caught

出典モデル評価と交差検証
Preference pair

Two candidate outputs for one input plus the human’s choice between them

出典人間のフィードバックによる強化学習
Prior

The degree of belief in a hypothesis before seeing data

出典ベイズの定理
Probability density

The "thickness" of probability for a continuous variable; its integral over an interval is the probability

出典確率と確率分布
Projection head

The MLP the contrastive loss is applied to, usually discarded after training

出典自己教師あり視覚とマルチモーダル対照学習
Prompt injection

Smuggling malicious instructions as data for the model to follow

出典安全性・アラインメント・プロンプトインジェクション
Pruning

Removing low-impact weights or whole structures

出典モデル圧縮

Q 2

Quantisation

Representing float weights and activations with low-bit integers

出典モデル圧縮
Query / Key / Value

The three vector roles: what you seek, what is on offer, what is carried

出典アテンション機構

R 14

Random variable

A function mapping outcomes of a random experiment to numbers

出典確率と確率分布
Rank

The number of independent directions the map actually spans; at most rows or columns

出典行列演算と線形写像
ReAct

A prompting paradigm alternating reasoning and action

出典エージェントとツール利用
Receptive field

The region of the original input that a given output covers

出典畳み込みニューラルネットワーク
Receptive field

The input region that one output pixel depends on

出典畳み込み演算
Red teaming

Actively hunting for failure and misuse paths

出典安全性・アラインメント・プロンプトインジェクション
Regret

The gap between realised cumulative reward and always picking the best arm

出典探索と活用
Reparameterisation

Writing sampling as a deterministic transform plus external noise so gradients flow

出典オートエンコーダと変分オートエンコーダ
Reranking

Rescoring candidate passages with a more accurate model

出典検索拡張生成
Residual connection

Adding the input past a sublayer to ease vanishing gradients in depth

出典Transformer アーキテクチャ
Reward hacking

Exploiting the proxy reward instead of genuinely completing the task

出典人間のフィードバックによる強化学習
Reward model

A model fitting human preferences and emitting a differentiable score

出典人間のフィードバックによる強化学習
ROC-AUC

Area under the ROC curve, measuring ranking ability across all thresholds

出典モデル評価と交差検証
RoPE

Rotary Position Embedding: relative position with better extrapolation

出典Transformer アーキテクチャ

S 19

Saturation

A function whose derivative tends to 0 at the extremes, blocking gradients

出典活性化関数
Self-attention

Attention whose Q, K and V all come from one sequence

出典アテンション機構
Self-information

The information of a single event, −log p

出典エントロピーと情報理論
Self-play

Generating training data by having an agent play against its past selves

出典深層強化学習
Self-supervised

Labels manufactured from the data itself, no manual annotation

出典事前学習とファインチューニング
Semantic / instance / panoptic

Class → class + instance → the two unified

出典セマンティックセグメンテーション
SentencePiece

A subword toolkit that runs directly on the character/byte stream

出典トークン化
SFT

Supervised fine-tuning on instruction–response pairs

出典プロンプト設計とアラインメント
SGD

Approximating the full gradient with a mini-batch

出典勾配と勾配降下法
Shannon entropy

The average information, or uncertainty, of a random variable

出典エントロピーと情報理論
Singular value decomposition

Writing any matrix as the product "rotate · stretch · rotate"

出典行列演算と線形写像
Skip connection

Routing shallow high-resolution features into deep layers to preserve boundaries

出典セマンティックセグメンテーション
Skip-gram

A training objective that predicts surrounding words from the centre

出典単語埋め込み
Softmax

Turns a set of real scores into a probability distribution summing to 1

出典確率と確率分布
Spatiotemporal patch

A local unit spanning frames in video, used to model motion

出典マルチモーダル生成
Speculative decoding

A small model drafts and the large model verifies in parallel to speed up generation

出典推論最適化とサービング
State S

The variables describing the present situation; must satisfy the Markov property

出典マルコフ決定過程
State value V(s)

Expected discounted return from s under policy π

出典価値関数と Q 学習
Stride

How many pixels the window jumps each step

出典畳み込み演算

T 11

Target network

A slowly updated copy of the network providing stable bootstrap targets

出典深層強化学習
TD error

The gap between the fresh target and the old estimate

出典価値関数と Q 学習
Tensor parallelism

Splitting a single layer’s large matrices across devices

出典学習・推論インフラストラクチャ
Thompson sampling

Sample from the posterior and pick the max, auto-directing exploration to uncertainty

出典探索と活用
Time to first token (TTFT)

Time from sending a request to receiving the first token

出典推論最適化とサービング
Tool

One external capability an agent may invoke

出典エージェントとツール利用
Top-1 / top-5 error

Whether the top prediction / top five include the true label

出典画像分類
Transfer learning

Pre-train on a large dataset, then fine-tune on a small task

出典画像分類
Transpose

Flip a matrix across its diagonal so rows become columns

出典行列演算と線形写像
Transposed convolution

An upsampling operation common in segmentation decoders

出典セマンティックセグメンテーション
Trust region / KL constraint

Bounds how far the new policy may drift from the old

出典方策勾配法

V 6

Vanishing gradient

Gradients shrinking exponentially as they are multiplied across layers

出典活性化関数
Variance

How sensitive the model is to perturbations of the training set

出典バイアス・バリアンス分解
Vector database

A store providing nearest-neighbour search over high-dimensional vectors

出典検索拡張生成
ViT

An architecture that applies a Transformer to image patches

出典画像分類
Vocabulary

The fixed set of all tokens and their indices

出典トークン化
Vocoder

The component that turns acoustic features back into a waveform

出典マルチモーダル生成

W 4

Wasserstein distance

An earth-mover distance between distributions, better behaved for training than JS divergence

出典敵対的生成ネットワーク
Weight

How strongly an input influences the output; may be positive or negative

出典ニューロンとパーセプトロン
Weight decay (L2)

Adding a squared-weight penalty to the loss to suppress large weights

出典過学習と正則化
Weight sharing

Reusing one set of weights across all spatial positions

出典畳み込みニューラルネットワーク

Z 3

ZeRO

Sharding optimiser states, gradients and parameters to cut per-device memory

出典学習・推論インフラストラクチャ
Zero-centred

Outputs symmetric about 0, which aids optimisation

出典活性化関数
Zero-shot classification

Classifying directly with text prompts, without fine-tuning

出典自己教師あり視覚とマルチモーダル対照学習

Ε 1

ε-greedy

Explore at random with probability ε, exploit the current best otherwise

出典探索と活用