Saltar al contenido
Atlas de IA

Probabilidad y distribuciones

Un modelo nunca te da una respuesta; te da un grado de creencia sobre cada respuesta posible

01 Fundamentos de matemáticas y estadísticaPrincipianteEntrada 4 de este dominio

El texto completo se presenta en inglés; el título y el resumen están traducidos.

DEFINICIÓN

Probability quantifies how likely an uncertain event is, on a scale from 0 to 1. A random variable maps the outcomes of a random experiment to numbers, and a probability distribution says how likely each number is. Machine learning models prediction as a distribution: rather than asserting a single answer, a model assigns a probability to every possible answer.

Intuición

The forecast says "a 70% chance of rain tomorrow". It does not decide for you whether to carry an umbrella — it hands you a whole set of possibilities with weights. Your job is to weigh the nuisance of the umbrella against the risk of getting soaked using those weights. A model does the same: it turns "what is the next word" into a table of probabilities and lets the downstream decision bear the consequences.

Fig. 1

Discrete versus continuous distributions: the former puts probability on individual points, the latter describes density across an interval

Fig. 2

Sampling 800 heights from a normal distribution and counting frequencies: the more samples, the closer the bars hug the smooth bell curve

Cómo funciona

  1. 01

    Random variables: turn outcomes into numbers

    The result of a coin flip is "heads/tails", not a number; a random variable X maps it to 1 or 0. A discrete random variable takes finitely or countably many values, while a continuous one takes any value in an interval and is described by a probability density rather than a point probability.

  2. 02

    Common distributions: ready-made templates for different settings

    The Bernoulli distribution describes one success/failure trial; the categorical distribution — exactly what a softmax emits — describes one choice among many; the multinomial counts the outcomes of n such trials; the Gaussian is the limiting shape when many small independent effects add up. Choosing the right distribution often matters more than choosing the right model.

  3. 03

    Conditional probability and independence: how information changes belief

    P(A|B) is the probability of A given that B happened. If P(A|B)=P(A), then B carries no information about A and the two are independent. Conditional probability is the starting point of all "see evidence, then update" reasoning, and the skeleton of Bayes’ theorem.

  4. 04

    Maximum likelihood: the parameters that make the data least surprising

    Given a model, we ask: which parameters were most likely to produce the data in front of us? Multiply the probabilities of all samples to get the likelihood, take the log to turn products into sums, then maximise it. Cross-entropy loss is exactly negative log-likelihood — the seam where probability theory meets deep learning.

Fig. 4

A language model’s prediction of the next word at one moment: it does not pick one word but spreads probability across the whole vocabulary (a local illustration)

  • the · 31%
  • a · 18%
  • this · 12%
  • one · 7%
  • all other words · 32%

Dónde se usa

  • Next-word prediction in language models: a probability distribution over the entire vocabulary
  • A classifier’s softmax output: turning scores into probabilities that carry confidence
  • Medical testing, A/B tests and risk pricing: describing and managing uncertainty with probabilities
  • Sampling in generative models: drawing one sample from a distribution yields a concrete image or sentence

Errores comunes

  • High probability is not causation. Two events may be strongly correlated because a third factor drives both, or by pure coincidence; probability describes co-occurrence, not who caused whom.
  • Choosing the wrong distribution shape is more damaging than choosing the wrong model. Forcing long-tailed data (income, word frequency) into a Gaussian badly underestimates extreme events, systematically misstating risk and loss.
  • High confidence is not correctness. Softmax probabilities are often overconfident — a well-trained network can assign 99% to a wrong answer — so probabilities need calibration before they can be trusted as confidence.

Términos clave

Random variable
A function mapping outcomes of a random experiment to numbers
Probability density
The "thickness" of probability for a continuous variable; its integral over an interval is the probability
Softmax
Turns a set of real scores into a probability distribution summing to 1
Likelihood
The probability of the observed data under given parameters

Lecturas complementarias