تخطٍّ إلى المحتوى
أطلس الذكاء الاصطناعي

GLOSSARY

مسرد المصطلحات

يجمع كل مصطلح أساسي متفرق في المداخل ضمن فهرس واحد، فتراجع منه كلما صادفت كلمة غريبة.

193 مصطلحًا

3 1

3D Gaussian Splatting

Representing a scene with many 3D Gaussian ellipsoids for fast rendering

منالتوليد متعدد الوسائط

A 9

Accelerator

A high-throughput parallel unit such as a GPU or TPU

منالبنية التحتية للتدريب والاستدلال
Action A

What the agent can do; either discrete or continuous

منعملية قرار ماركوف
Action value Q(s, a)

Expected discounted return after forcing the first action to be a

مندوال القيمة وتعلّم Q
Activation function

A function that applies a nonlinear transform to the weighted sum

منالعصبون والبرسبترون
Actor / Critic

The policy network and the value network: one acts, one scores

منتدرّجات السياسة
Advantage A(s, a)

How much better an action is than the average at that state

منتدرّجات السياسة
Alignment

Making model behaviour match human intent and values

منالسلامة والمواءمة وحقن الأوامر
Anomaly detection

Finding the few samples that deviate from the bulk distribution

منالتعلّم بدون إشراف
Automatic differentiation

Letting a framework compute exact gradients automatically, not by numerical approximation

منالانتشار العكسي

B 8

Batch size

How many samples estimate the gradient per step

منالتدرّج والانحدار التدريجي
Bias

How far the model’s average prediction departs from the true regularity

منالمقايضة بين التحيّز والتباين
Bias

A learnable offset applied to the threshold

منالعصبون والبرسبترون
Bit depth

How many bits encode each channel; 8 bits give 256 levels

منالتمثيل الرقمي للصورة
Bottleneck

The low-dimensional layer holding the latent code, limiting its bandwidth

منالمرمّزات الذاتية والمرمّزات الذاتية التباينية
Bounding box

A rectangle represented as (x, y, w, h) or corner points

منكشف الأجسام
BPE

Byte-Pair Encoding: bottom-up merging of frequent symbol pairs

منالترميز إلى رموز (Tokenization)
BPTT

Backpropagation through time after unrolling

منالشبكات العصبية المتكررة

C 20

Catastrophic forgetting

Rapid loss of old abilities while learning a new task

منالتدريب المسبق والضبط الدقيق
Chain rule

The derivative of a composition is the product of the local derivatives

منالانتشار العكسي
Chain-of-thought (CoT)

Making the model write out intermediate reasoning steps

منهندسة الأوامر والمواءمة
Channel

A distinct measurement at the same location, such as R/G/B or alpha

منالتمثيل الرقمي للصورة
Chunking

Splitting long documents into retrievable pieces

منالتوليد المعزّز بالاسترجاع
Classifier-free guidance

Extrapolating between conditional and unconditional predictions to control prompt fidelity

مننماذج الانتشار
Clustering

Grouping samples by similarity (k-means, hierarchical clustering)

منالتعلّم بدون إشراف
Colour space

A coordinate system for colour values, such as sRGB, HSV or Lab

منالتمثيل الرقمي للصورة
Computation graph

A computation expressed as nodes and directed edges over which derivatives propagate

منالانتشار العكسي
Confusion matrix

A cross-tabulation of true versus predicted classes

منتقييم النموذج والتحقق المتقاطع
Continuous batching

Re-forming the batch every step to keep the GPU busy

منتحسين الاستدلال وتقديم الخدمة
Contrastive learning

Learning representations by pulling positives together and pushing negatives apart

منالرؤية ذاتية الإشراف والتعلّم التقابلي متعدد الوسائط
Contrastive loss

A loss that pulls same-class embeddings together and pushes different-class ones apart

مندوال الخسارة
ControlNet

A bypass network guiding structure from a condition map

منالانتشار في الفضاء الكامن والتحكم الشرطي
Cosine similarity

The alignment of two vector directions, from −1 to 1

منالتضمينات اللفظية
Cross-attention

Query from one sequence, Key/Value from another

منآلية الانتباه
Cross-attention

The attention mechanism letting image features query text vectors

منالانتشار في الفضاء الكامن والتحكم الشرطي
Cross-entropy

The information needed to encode data from P using distribution Q

منالإنتروبيا ونظرية المعلومات
Cross-entropy

Negative log-probability of the correct class; the default classification loss

مندوال الخسارة
Cross-entropy loss

The standard objective for classification training

منتصنيف الصور

D 12

DDIM

Deterministic sampling achieving comparable quality in a few dozen steps

مننماذج الانتشار
DDPM

Discrete Markov diffusion, typically needing a thousand sampling steps

مننماذج الانتشار
Degradation problem

Deeper networks with higher training error, and not from overfitting

منالتطبيع والوصلات المتبقية
Density estimation

Estimating the probability distribution the data follows

منالتعلّم بدون إشراف
Dimension

The number of entries in a vector

منالمتجهات والفضاءات المتجهية
Dimensionality reduction

Compressing high-dimensional data to fewer dimensions while preserving structure (PCA, t-SNE, UMAP)

منالتعلّم بدون إشراف
Discount factor γ

Between 0 and 1; how much future rewards are valued

منعملية قرار ماركوف
Discriminator

The network judging real versus fake, serving as the loss

منالشبكات التوليدية التنافسية
Double descent

The modern counterexample where test error falls again past the interpolation point

منالمقايضة بين التحيّز والتباين
DPO

Direct preference optimisation without an explicit reward model

منهندسة الأوامر والمواءمة
Dropout

Randomly silencing units during training to prevent co-adaptation

منالإفراط في التعلّم والتنظيم
Dying ReLU

A neuron stuck in the negative region with zero gradient, no longer updating

مندوال التنشيط

E 10

Early stopping

Halting training before validation loss turns upward

منالإفراط في التعلّم والتنظيم
Eigenvector / eigenvalue

A vector whose direction is unchanged by the map, and the factor by which it is scaled

منعمليات المصفوفات والتحويلات الخطية
ELBO

A lower bound on the log-likelihood: the reconstruction term minus the KL term; a VAE’s actual objective

منالمرمّزات الذاتية والمرمّزات الذاتية التباينية
Embedding

The layer, or its output, that maps a discrete object into a continuous vector

منالمتجهات والفضاءات المتجهية
Embedding model

A model that encodes text into vectors

منالتوليد المعزّز بالاسترجاع
Empirical risk

The model’s average loss on the training samples

منالتعلّم بالإشراف
Equivariance

When the input shifts, the output shifts accordingly rather than changing

منالشبكات العصبية الالتفافية
Evidence

The total probability of the data across all hypotheses; it normalises the result

منمبرهنة بايز
Experience replay

Store past transitions and sample randomly to break correlation

منالتعلّم المعزّز العميق
Explicit density

A model that writes down or approximates p(x), e.g. autoregressive or diffusion

مننظرة عامة على النماذج التوليدية

F 3

F1

The harmonic mean of precision and recall

منتقييم النموذج والتحقق المتقاطع
FID

Fréchet distance between generated and real distributions in Inception feature space; lower is better

مننظرة عامة على النماذج التوليدية
Function calling

The model emitting structured arguments to invoke an external function

منالوكلاء واستخدام الأدوات

G 5

Gating

Using 0–1 coefficients from Sigmoid to control how much information passes

منالشبكات العصبية المتكررة
Generalisation

Performance on data the model has not seen

منالإفراط في التعلّم والتنظيم
Generator

The network mapping noise to samples

منالشبكات التوليدية التنافسية
Gradient flow

The magnitude and stability of gradients as they propagate layer by layer

منالانتشار العكسي
Guardrail

Checks and constraints bounding what an agent may do

منالوكلاء واستخدام الأدوات

H 3

Hidden state

A continuously updated "summary so far" vector

منالشبكات العصبية المتكررة
Hinge loss

Requires the correct class to win by a margin; the heart of the SVM

مندوال الخسارة
Hypothesis space

The set of all functions the model can represent

منالتعلّم بالإشراف

I 9

Identity shortcut

The path in a residual connection that adds the input straight back to the output

منالتطبيع والوصلات المتبقية
Implicit density

A model that offers only a sampler, not a probability, e.g. a GAN

مننظرة عامة على النماذج التوليدية
In-context learning

Solving a task from prompt examples without updating parameters

منهندسة الأوامر والمواءمة
InfoNCE

The standard contrastive loss; essentially a multi-class cross-entropy

منالرؤية ذاتية الإشراف والتعلّم التقابلي متعدد الوسائط
Inner product

Element-wise product summed over entries; the numerator of cosine similarity

منالمتجهات والفضاءات المتجهية
Input x

The feature vector fed to the model

منالتعلّم بالإشراف
Internal covariate shift

The shifting distribution of inputs to later layers during training

منالتطبيع والوصلات المتبقية
IoU

The ratio of the intersection to the union of two boxes

منكشف الأجسام
Irreducible error

The unavoidable error floor caused by label noise

منالمقايضة بين التحيّز والتباين

J 1

Jailbreak

Inducing a model past its safety training

منالسلامة والمواءمة وحقن الأوامر

K 7

Kernel / filter

A set of learnable weights that slides over the input

منالشبكات العصبية الالتفافية
Kernel / filter

The small weight matrix that is learned

منعمليات الالتفاف
KL divergence

Cross-entropy minus true entropy; non-negative and asymmetric

منالإنتروبيا ونظرية المعلومات
KL divergence

Measures how far the encoded distribution deviates from a standard normal; acts as a regulariser

منالمرمّزات الذاتية والمرمّزات الذاتية التباينية
KL penalty

Penalises divergence from the reference policy to prevent degeneration

منالتعلّم المعزّز من التغذية الراجعة البشرية
Knowledge distillation

Training a small model on a large model’s soft outputs

منضغط النماذج
KV cache

Caching past tokens’ keys and values to avoid recomputation

منتحسين الاستدلال وتقديم الخدمة

L 12

Label y

The correct output for each sample; the source of supervision

منالتعلّم بالإشراف
Latent space

The low-dimensional representation space produced by the autoencoder

منالانتشار في الفضاء الكامن والتحكم الشرطي
Learning rate η

How far each step moves

منالتدرّج والانحدار التدريجي
Likelihood

The probability of the observed data under given parameters

منالاحتمال والتوزيعات الاحتمالية
Likelihood

The probability of observed data given that the hypothesis is true

منمبرهنة بايز
Linearly separable

A hyperplane exists that separates the two classes perfectly

منالعصبون والبرسبترون
Log-derivative trick

Turns the gradient of an expectation into a weighted sum of log-probabilities

منتدرّجات السياسة
Long-range dependency

Influence between elements far apart in a sequence

منالشبكات العصبية المتكررة
LoRA

Low-rank adapters training only a tiny number of new parameters

منالتدريب المسبق والضبط الدقيق
LoRA

Low-rank adaptation increments for low-cost customisation

منالانتشار في الفضاء الكامن والتحكم الشرطي
Loss surface

The high-dimensional terrain of loss values over parameter space

منالتدرّج والانحدار التدريجي
Low-rank factorisation

Approximating a large matrix by a product of two smaller ones

منضغط النماذج

M 9

mAP

Mean average precision across classes and IoU thresholds

منكشف الأجسام
Masked language modelling

Hide random words and recover them, a bidirectional objective

منالتدريب المسبق والضبط الدقيق
MCTS

An algorithm that evaluates moves via sampled rollouts to guide search

منالتعلّم المعزّز العميق
Mean squared error (MSE)

The average squared difference between prediction and label; the default regression loss

مندوال الخسارة
mIoU

The mean of per-class IoU, the primary segmentation metric

منالتجزئة الدلالية
Mode collapse

When a generator covers only a few modes of the data distribution

مننظرة عامة على النماذج التوليدية
Mode collapse

The generator covers few modes and loses diversity

منالشبكات التوليدية التنافسية
Multi-armed bandit

The simplest sequential model: unknown reward distributions, one pull per round

منالاستكشاف والاستغلال
Multi-head attention

Several attentions in parallel, each learning a different focus

منآلية الانتباه

N 5

Negative sampling

Replacing full-vocabulary softmax with a few random negatives

منالتضمينات اللفظية
NeRF

A neural network representing a scene’s radiance field for novel-view synthesis

منالتوليد متعدد الوسائط
NMS

Non-maximum suppression, removing duplicate boxes

منكشف الأجسام
Noise schedule

The timetable of noise added per step, described by βₜ or ᾱₜ

مننماذج الانتشار
Norm

A function measuring a vector’s "length"; L2 is the common choice

منالمتجهات والفضاءات المتجهية

O 3

Off-policy

The behaviour policy may differ from the policy being learned

مندوال القيمة وتعلّم Q
One-hot

A sparse vector with a single 1; distinct words are fully orthogonal

منالتضمينات اللفظية
Out-of-vocabulary (OOV)

A word absent from the vocabulary, spelled out from subwords

منالترميز إلى رموز (Tokenization)

P 16

Padding

Adding zeros at the border to control output size

منعمليات الالتفاف
Perplexity

The exponential of the cross-entropy; the effective number of options the model hesitates among per step

منالإنتروبيا ونظرية المعلومات
Pipeline parallelism

Placing different layers on different devices and filling bubbles with micro-batches

منالبنية التحتية للتدريب والاستدلال
Pixel

The smallest sampling unit of an image, carrying one or more channel values

منالتمثيل الرقمي للصورة
Policy π

A mapping from states to actions, or to a distribution over actions

منعملية قرار ماركوف
Positional encoding

An explicit order signal, sinusoidal or RoPE

منمعمارية Transformer
Posterior

The updated degree of belief after incorporating the evidence

منمبرهنة بايز
Pre-activation

A layout placing normalisation before the convolution, which trains more stably

منالتطبيع والوصلات المتبقية
Pre-LN

Placing layer norm before each sublayer for stability

منمعمارية Transformer
Precision & recall

Precision asks how many alerts are real; recall asks how many real cases were caught

منتقييم النموذج والتحقق المتقاطع
Preference pair

Two candidate outputs for one input plus the human’s choice between them

منالتعلّم المعزّز من التغذية الراجعة البشرية
Prior

The degree of belief in a hypothesis before seeing data

منمبرهنة بايز
Probability density

The "thickness" of probability for a continuous variable; its integral over an interval is the probability

منالاحتمال والتوزيعات الاحتمالية
Projection head

The MLP the contrastive loss is applied to, usually discarded after training

منالرؤية ذاتية الإشراف والتعلّم التقابلي متعدد الوسائط
Prompt injection

Smuggling malicious instructions as data for the model to follow

منالسلامة والمواءمة وحقن الأوامر
Pruning

Removing low-impact weights or whole structures

منضغط النماذج

Q 2

Quantisation

Representing float weights and activations with low-bit integers

منضغط النماذج
Query / Key / Value

The three vector roles: what you seek, what is on offer, what is carried

منآلية الانتباه

R 14

Random variable

A function mapping outcomes of a random experiment to numbers

منالاحتمال والتوزيعات الاحتمالية
Rank

The number of independent directions the map actually spans; at most rows or columns

منعمليات المصفوفات والتحويلات الخطية
ReAct

A prompting paradigm alternating reasoning and action

منالوكلاء واستخدام الأدوات
Receptive field

The region of the original input that a given output covers

منالشبكات العصبية الالتفافية
Receptive field

The input region that one output pixel depends on

منعمليات الالتفاف
Red teaming

Actively hunting for failure and misuse paths

منالسلامة والمواءمة وحقن الأوامر
Regret

The gap between realised cumulative reward and always picking the best arm

منالاستكشاف والاستغلال
Reparameterisation

Writing sampling as a deterministic transform plus external noise so gradients flow

منالمرمّزات الذاتية والمرمّزات الذاتية التباينية
Reranking

Rescoring candidate passages with a more accurate model

منالتوليد المعزّز بالاسترجاع
Residual connection

Adding the input past a sublayer to ease vanishing gradients in depth

منمعمارية Transformer
Reward hacking

Exploiting the proxy reward instead of genuinely completing the task

منالتعلّم المعزّز من التغذية الراجعة البشرية
Reward model

A model fitting human preferences and emitting a differentiable score

منالتعلّم المعزّز من التغذية الراجعة البشرية
ROC-AUC

Area under the ROC curve, measuring ranking ability across all thresholds

منتقييم النموذج والتحقق المتقاطع
RoPE

Rotary Position Embedding: relative position with better extrapolation

منمعمارية Transformer

S 19

Saturation

A function whose derivative tends to 0 at the extremes, blocking gradients

مندوال التنشيط
Self-attention

Attention whose Q, K and V all come from one sequence

منآلية الانتباه
Self-information

The information of a single event, −log p

منالإنتروبيا ونظرية المعلومات
Self-play

Generating training data by having an agent play against its past selves

منالتعلّم المعزّز العميق
Self-supervised

Labels manufactured from the data itself, no manual annotation

منالتدريب المسبق والضبط الدقيق
Semantic / instance / panoptic

Class → class + instance → the two unified

منالتجزئة الدلالية
SentencePiece

A subword toolkit that runs directly on the character/byte stream

منالترميز إلى رموز (Tokenization)
SFT

Supervised fine-tuning on instruction–response pairs

منهندسة الأوامر والمواءمة
SGD

Approximating the full gradient with a mini-batch

منالتدرّج والانحدار التدريجي
Shannon entropy

The average information, or uncertainty, of a random variable

منالإنتروبيا ونظرية المعلومات
Singular value decomposition

Writing any matrix as the product "rotate · stretch · rotate"

منعمليات المصفوفات والتحويلات الخطية
Skip connection

Routing shallow high-resolution features into deep layers to preserve boundaries

منالتجزئة الدلالية
Skip-gram

A training objective that predicts surrounding words from the centre

منالتضمينات اللفظية
Softmax

Turns a set of real scores into a probability distribution summing to 1

منالاحتمال والتوزيعات الاحتمالية
Spatiotemporal patch

A local unit spanning frames in video, used to model motion

منالتوليد متعدد الوسائط
Speculative decoding

A small model drafts and the large model verifies in parallel to speed up generation

منتحسين الاستدلال وتقديم الخدمة
State S

The variables describing the present situation; must satisfy the Markov property

منعملية قرار ماركوف
State value V(s)

Expected discounted return from s under policy π

مندوال القيمة وتعلّم Q
Stride

How many pixels the window jumps each step

منعمليات الالتفاف

T 11

Target network

A slowly updated copy of the network providing stable bootstrap targets

منالتعلّم المعزّز العميق
TD error

The gap between the fresh target and the old estimate

مندوال القيمة وتعلّم Q
Tensor parallelism

Splitting a single layer’s large matrices across devices

منالبنية التحتية للتدريب والاستدلال
Thompson sampling

Sample from the posterior and pick the max, auto-directing exploration to uncertainty

منالاستكشاف والاستغلال
Time to first token (TTFT)

Time from sending a request to receiving the first token

منتحسين الاستدلال وتقديم الخدمة
Tool

One external capability an agent may invoke

منالوكلاء واستخدام الأدوات
Top-1 / top-5 error

Whether the top prediction / top five include the true label

منتصنيف الصور
Transfer learning

Pre-train on a large dataset, then fine-tune on a small task

منتصنيف الصور
Transpose

Flip a matrix across its diagonal so rows become columns

منعمليات المصفوفات والتحويلات الخطية
Transposed convolution

An upsampling operation common in segmentation decoders

منالتجزئة الدلالية
Trust region / KL constraint

Bounds how far the new policy may drift from the old

منتدرّجات السياسة

V 6

Vanishing gradient

Gradients shrinking exponentially as they are multiplied across layers

مندوال التنشيط
Variance

How sensitive the model is to perturbations of the training set

منالمقايضة بين التحيّز والتباين
Vector database

A store providing nearest-neighbour search over high-dimensional vectors

منالتوليد المعزّز بالاسترجاع
ViT

An architecture that applies a Transformer to image patches

منتصنيف الصور
Vocabulary

The fixed set of all tokens and their indices

منالترميز إلى رموز (Tokenization)
Vocoder

The component that turns acoustic features back into a waveform

منالتوليد متعدد الوسائط

W 4

Wasserstein distance

An earth-mover distance between distributions, better behaved for training than JS divergence

منالشبكات التوليدية التنافسية
Weight

How strongly an input influences the output; may be positive or negative

منالعصبون والبرسبترون
Weight decay (L2)

Adding a squared-weight penalty to the loss to suppress large weights

منالإفراط في التعلّم والتنظيم
Weight sharing

Reusing one set of weights across all spatial positions

منالشبكات العصبية الالتفافية

Z 3

ZeRO

Sharding optimiser states, gradients and parameters to cut per-device memory

منالبنية التحتية للتدريب والاستدلال
Zero-centred

Outputs symmetric about 0, which aids optimisation

مندوال التنشيط
Zero-shot classification

Classifying directly with text prompts, without fine-tuning

منالرؤية ذاتية الإشراف والتعلّم التقابلي متعدد الوسائط

Ε 1

ε-greedy

Explore at random with probability ε, exploit the current best otherwise

منالاستكشاف والاستغلال