Chuyển đến nội dung
Bản đồ AI

Mạng nơ-ron hồi tiếp

Cho mạng một “bộ nhớ”: một đơn vị được dùng lại theo thời gian để xử lý chuỗi dài tùy ý

03 Học sâuTrung cấpMục từ thứ 5 trong lĩnh vực

Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.

ĐỊNH NGHĨA

A recurrent neural network processes sequences. At each time step it reads an input and receives the previous hidden state, producing a new hidden state and an output. Because the same unit is reused over time and all steps share one set of weights, it handles variable-length sequences and carries history into the current decision.

Trực giác

Reading a sentence is like listening to someone speak: you do not replay the whole thing, you keep updating an in-head summary of "what has been said so far" — that continuously rewritten summary is the hidden state. Sadly, a vanilla RNN’s memory is short: information is multiplied at every step, so after dozens of steps earlier content either fades (vanishing gradient) or blows up (exploding gradient). Gating is, at heart, the act of fitting that summary with a few valves: what to store, what to drop, what to reveal.

Hình 1

A recurrent network unrolled over time: the same unit reads xₜ, updates hₜ and emits yₜ at each step, with hidden states linked by one shared weight Wₕ

Hình 2

The longer the sequence, the harder it is for a vanilla RNN to retain early information, and accuracy collapses; gated LSTM/GRU stay far more stable on long sequences (a long-range-dependency benchmark; magnitudes from typical reported results)

  • Vanilla RNN
  • LSTM
  • GRU

Cách hoạt động

  1. 01

    Hidden state and unrolling over time

    hₜ = f(W hₜ₋₁ + U xₜ + b). Although called "recurrent", training unrolls it into a very deep chain over time, with all steps sharing a single set of weights.

  2. 02

    Vanishing and exploding gradients

    Errors are propagated back through time (BPTT), multiplied by the weight matrix and activation derivative at each step. This repeated product makes the gradient decay or grow exponentially with the number of steps, so long-range dependencies are hard to learn.

  3. 03

    Gating: LSTM and GRU

    LSTM uses input, forget and output gates to decide what to store, drop and reveal, protecting gradients through a near-identity cell-state path; GRU achieves a similar effect with fewer gates.

  4. 04

    Superseded by attention

    Even with gating, an RNN must compute step by step, cannot parallelise like a Transformer, and struggles to retain very long context. Attention replaces "pass it along word by word" with "look at all positions directly", which is why it became the mainstream.

Ứng dụng

  • Language modelling and machine translation (the mainstream before attention became widespread)
  • Frame-by-frame tasks such as speech and handwriting recognition
  • Time-series forecasting: sensor, weather and financial data
  • Encoders and decoders for sequence-to-sequence tasks

Hiểu lầm thường gặp

  • "RNNs have memory" only to a degree: a vanilla RNN’s effective memory is typically a few dozen steps, and information beyond that is hard to retain.
  • LSTM mitigates rather than solves long-range dependencies. Gating adds computation without changing the fundamental sequential bottleneck.
  • Training with teacher forcing but generating freely at inference creates a distribution mismatch that accumulates and amplifies errors.

Thuật ngữ chính

Hidden state
A continuously updated "summary so far" vector
BPTT
Backpropagation through time after unrolling
Gating
Using 0–1 coefficients from Sigmoid to control how much information passes
Long-range dependency
Influence between elements far apart in a sequence

Đọc thêm