Recurrent Neural Network (RNN)

A Recurrent Neural Network (RNN) is a neural network specifically designed for processing sequential data (such as text, speech, and time series).

Unlike traditional feedforward neural networks, RNNs have a "memory" capability that can retain information from previous steps.

RNNs use the hidden state from the previous step to influence the output of the current step, thereby capturing temporal dependencies in sequences.

Core Idea of RNN

The core of RNN lies inrecurrent connections(Recurrent Connection), meaning the network's output depends not only on the current input but also on the inputs from all previous time steps. This structure enables RNNs to process sequential data of arbitrary length.

Traditional neural networks: Inputs and outputs are independent (e.g., image classification, where individual images are unrelated).

RNN: Throughrecurrent connections(Recurrent Connection), the hidden state from the previous step is passed to the next step, forming a "memory".

  • Input at each step = current data + hidden state from the previous step.

  • The output depends not only on the current input but also on the context from all previous steps.

Just as when reading a sentence, understanding the current word depends on previously read content (e.g., "He opened the __", you would predict "door" or "book").

Example

# Simple RNN cell implementation example
import numpy as np

class SimpleRNN:
    def __init__(self, input_size, hidden_size):
        self.Wx = np.random.randn(hidden_size, input_size)  # Input weights
        self.Wh = np.random.randn(hidden_size, hidden_size)  # Hidden state weights
        self.b = np.zeros((hidden_size, 1))  # Bias term
   
    def forward(self, x, h_prev):
        h_next = np.tanh(np.dot(self.Wx, x) + np.dot(self.Wh, h_prev) + self.b)
        return h_next

How RNN Works

At each time step t, the RNN performs the following computations:

  1. Receives the current input xₜ and the hidden state hₜ₋₁ from the previous time step
  2. Computes the new hidden state hₜ = f(Wₕₕ·hₜ₋₁ + Wₓₕ·xₜ + b)
  3. Produces the output yₜ = g(Wₕᵧ·hₜ + c)

Where f and g are typically activation functions (such as tanh or softmax).

Advantages and Disadvantages of RNN

Advantages:

  • Can handle variable-length sequences
  • Theoretically can remember historical information of arbitrary length
  • Parameter sharing (the same set of weights is used for all time steps)

Disadvantages:

  • Vanishing/exploding gradient problem (difficult to learn long-term dependencies)
  • Low computational efficiency (cannot process time steps in parallel)

Long Short-Term Memory (LSTM)

LSTM (Long Short-Term Memory) is an improved architecture of RNN, specifically designed to solve the long-term dependency problem of standard RNNs.

2.1 Core Structure of LSTM

LSTM introduces three gating mechanisms and a memory cell:

Component Function
Input gate Controls the flow of new information
Forget gate Decides which old information to discard
Output gate Controls the amount of information to output
Memory cell Stores long-term state

Example

# Basic implementation of an LSTM cell
class LSTMCell:
    def __init__(self, input_size, hidden_size):
        # Combine the weights of all gates
        self.W = np.random.randn(4*hidden_size, input_size+hidden_size)
        self.b = np.random.randn(4*hidden_size, 1)
   
    def forward(self, x, h_prev, c_prev):
        combined = np.vstack((h_prev, x))
        gates = np.dot(self.W, combined) + self.b
       
        # Split to obtain each gate
        f_gate = sigmoid(gates[:hidden_size])  # Forget gate
        i_gate = sigmoid(gates[hidden_size:2*hidden_size])  # Input gate
        o_gate = sigmoid(gates[2*hidden_size:3*hidden_size])  # Output gate
        c_candidate = np.tanh(gates[3*hidden_size:])  # Candidate memory
       
        # Update memory and hidden state
        c_next = f_gate * c_prev + i_gate * c_candidate
        h_next = o_gate * np.tanh(c_next)
       
        return h_next, c_next

How LSTM Solves the Long-Term Dependency Problem

  1. Selective memory: The forget gate can decide to retain or discard specific information
  2. Gradient pathway: The memory cell provides a relatively direct path for gradient propagation
  3. Information protection: The stored memory content is not directly modified by the operations at every time step

Gated Recurrent Unit (GRU)

GRU (Gated Recurrent Unit) is a simplified version of LSTM that reduces the number of parameters while maintaining similar performance.

Core Structure of GRU

GRU merges certain components of LSTM:

Component Function
Update gate Decides how much old information to retain
Reset gate Decides how to combine new and old information
Candidate activation New state computed based on the reset gate

Example

# Implementation of a GRU cell
class GRUCell:
    def __init__(self, input_size, hidden_size):
        self.W = np.random.randn(3*hidden_size, input_size+hidden_size)
        self.b = np.random.randn(3*hidden_size, 1)
   
    def forward(self, x, h_prev):
        combined = np.vstack((h_prev, x))
        gates = np.dot(self.W, combined) + self.b
       
        # Split gating signals
        z = sigmoid(gates[:hidden_size])  # Update gate
        r = sigmoid(gates[hidden_size:2*hidden_size])  # Reset gate
        h_candidate = np.tanh(np.dot(self.W[2*hidden_size:],
                              np.vstack((r*h_prev, x))) + self.b[2*hidden_size:]
       
        # Update hidden state
        h_next = (1-z)*h_prev + z*h_candidate
        return h_next

GRU vs LSTM

Feature GRU LSTM
Number of parameters Fewer More
Training speed Faster Slower
Memory cell None Yes
Number of gates 2 3
Performance Better for small datasets May be better for large datasets

Bidirectional RNN (Bi-RNN)

Bidirectional RNN enhances sequence modeling capability by considering both past and future context information simultaneously.

Bidirectional RNN Architecture

Bi-RNN consists of two independent RNN layers:

  1. Forward layer: processes the sequence in chronological order
  2. Backward layer: processes the sequence in reverse chronological order

The final output is a combination of the outputs from both directions (usually concatenation or summation).

Application Scenarios of Bidirectional RNN

  1. Natural language processing: Part-of-speech tagging, named entity recognition
  2. Speech recognition: Uses surrounding context to improve accuracy
  3. Bioinformatics: Protein structure prediction
  4. Time series forecasting: Considers historical and future trends

Bidirectional LSTM/GRU

In modern applications, bidirectional RNNs typically use LSTM or GRU as the basic unit:

Example

from tensorflow.keras.layers import Bidirectional, LSTM

model.add(Bidirectional(LSTM(64)))  # Create a bidirectional LSTM layer

Practical Exercises

Exercise 1: Implement a Simple RNN

Use Python and NumPy to implement a simple RNN capable of character-level text generation.

Exercise 2: LSTM Sentiment Analysis

Use Keras to build an LSTM-based movie review sentiment classifier.

Exercise 3: Bidirectional GRU Named Entity Recognition

Implement a bidirectional GRU model to identify entities such as person names and locations in text.

Exercise 4: Comparative Experiment

Compare the performance differences of Vanilla RNN, LSTM, and GRU on the same dataset.


Summary and Further Learning

RNN and its variants are powerful tools for processing sequential data. To master them deeply:

  1. Understand how gradients propagate in RNNs
  2. Learn how the attention mechanism enhances RNNs
  3. Explore the relationship between the Transformer architecture and RNNs
  4. Practice various sequence modeling tasks (machine translation, speech synthesis, etc.)
Other Extensions