How AI Works

You might ask: I just want to use AI, why should I understand how it works?

The reason is simple: if you know where AI's capabilities come from, you can use it better.

  • You will know when to trust it and when to question it.

  • You will know what it is good at and what it is not good at.

  • You will know where hallucinations come from and how to reduce their impact.


Intuitive Understanding of Neural Networks

The core of modern AI isneural networks, a name that comes from the neuron structure of the human brain.

Analogy to Brain Neurons

The human brain has about 86 billion neurons that are interconnected and transmit signals. Each neuron receives input from other neurons, processes it, and then outputs to other neurons. The learning process is the process of adjusting the strength of these connections.

Artificial neural networks borrow this idea but make it greatly simplified.

Three Basic Layers

A typical neural network is divided into three layers: input layer, hidden layer, and output layer.

神经网络基本结构示意图

We use the example of recognizing whether an image is a cat or a dog:

LayerFunctionIn this example
Input layerReceives raw dataEach pixel and color information of the image
Hidden layerExtract features layer by layerEdges → Textures → Organs such as ears and eyes
Output layerGive the final result"85% probability it's a cat", "15% probability it's a dog"

Each layer has many neurons. Each neuron receives the output from the previous layer, does a simple calculation, and passes it to the next layer.

  • The first layer might recognize "there is a vertical line here" or "there is a circle here".

  • The second layer combines these: a vertical line plus a circle might be an ear.

  • The third layer continues combining: two pointed ears, whiskers, cat eyes — this is very likely a cat.

The amazing part is that these features are not designed by humans; the model learns them from data on its own.

The Simplest Neuron

Let's use a few lines of Python code to show what a neuron does:

Example

# ============================================
# Basic computation logic of an artificial neuron
# No complex math, just weighted sum + activation
# ============================================

def simple_neuron(inputs: list, weights: list, bias: float) -> float:
    """
The simplest neuron
inputs: input values (from neurons in the previous layer)
weights: weights (importance of each input, learned during training)
bias: bias (threshold, learned during training)
    """

    # Step 1: Weighted sum
    # Multiply each input by its corresponding weight, then sum them all up
    weighted_sum = 0.0
    for input_value, weight in zip(inputs, weights):
        weighted_sum += input_value * weight

    # Add bias
    weighted_sum += bias

    # Step 2: Activation function (makes the output non-linear)
    # Here we use the simplest ReLU: negative numbers become 0, positive numbers stay unchanged
    output = max(0.0, weighted_sum)

    return output


# Simulation: a neuron that determines "is this a cat's ear"
# Input: [pointiness, height position, has fur]
inputs = [0.8, 0.9, 0.7]  # These three features are all quite obvious

# Weights: learned after training (in the example example, we assume these values are already well-learned)
weights = [0.5, 0.4, 0.3]

# Bias: threshold
bias = -0.6

result = simple_neuron(inputs, weights, bias)
print(f"Neuron output: {result:.3f}")
print(f"Judgment: {'likely a cat ear' if result > 0 else 'not likely'}")
# Output: Neuron output: 0.660
# Output: Judgment: likely a cat ear

What this neuron does is very simple: take a weighted sum of the inputs, pass it through an activation function, and output the result.

But when thousands of such neurons are connected together, with each layer learning different features, the whole system produces astonishing intelligence.

Remember this intuition: a neural network = many simple computing units connected together, learning by adjusting connection weights.


Training vs. Inference: Two Different Stages

The AI lifecycle is divided into two completely distinct stages: training and inference.

Understanding the difference between these two stages can help you understand many things—such as why training is so expensive, while inference is relatively cheap.

Training: Let AI Learn

Training is a process: show the model large amounts of data, let it continuously adjust parameters, and make predictions more and more accurately.

For example, training a model to recognize cats and dogs:

  • 1. Prepare millions of labeled images (this one is a cat, that one is a dog)

  • 2. Let the model guess "what is this"; at first it will guess wrong a lot

  • 3. Tell it "you guessed wrong, it should be a cat", and let it adjust the weights in the network

  • 4. Repeat millions of times until the model predicts more and more accurately

The training phase requires enormous computational power and data.A large model may need thousands of GPUs trained for several months, costing millions of dollars.

Inference: Let AI Use

Inference is a process: use the trained model, give new input, get output.

You send ChatGPT a message, it replies to you—this is inference.

You use your phone to take a photo to identify a plant—this is also inference.

The characteristics of inference:

  • No need to adjust parameters, only use the trained weights for computation.

  • Usually only one GPU or even a phone chip is needed.

  • The cost is much lower than training.

Comparison of the Two

DimensionTrainingInference
GoalLearn knowledge, adjust weightsApply learned knowledge, provide answers
Data volumeNeeds massive amounts of dataA single input is enough
Compute requirementsExtremely high (thousands of GPUs)Relatively low (single GPU or phone)
CostExtremely high (millions of dollars)Relatively low (a few cents per time)
FrequencyA few times or dozens of timesMillions of times per second
Who does itCompanies like OpenAI, AnthropicOrdinary users or applications

As an analogy: training is like "studying hard for ten years", inference is like "solving problems in an exam".

Studying requires a lot of time and energy, but once you learn it, solving problems becomes fast.

When you use ChatGPT, you are doing "inference"—the model does not "learn" or "become smarter" from your conversations. Its knowledge ends at the moment training is completed.


Introduction to the Transformer Architecture

In 2017, Google published a paper titled "Attention Is All You Need", proposing the Transformer architecture.

This paper changed the entire AI field. Today's large language models are almost all based on Transformer.

Why is Transformer So Important?

Before Transformer, sequence data (such as sentences) was processed using RNN or LSTM.

Their problem is: they can only process one word at a time, making it difficult to capture long-distance dependencies.

For example: "I left my wallet at a café in Beijing. When I went back to look for it the next day, ____ was still there." — fill in "it" in the blank. You know "it" refers to "wallet" because you remember the earlier content.

When old models processed "still there", they might have already forgotten the existence of "wallet".

Transformer's breakthrough lies in:It can see the entire sentence at once, and through the attention mechanism, know which words to focus on.

Encoder and Decoder

A complete Transformer is divided into two parts:

ComponentFunctionTypical application scenarios
EncoderUnderstands input and converts text into vector representationsText classification, sentiment analysis, semantic search
DecoderGenerates output text based on understandingWriting, translation, dialogue generation

Some models use only the Encoder (such as BERT), some use only the Decoder (such as GPT), and some use both (such as T5).

The GPT series, Claude, and Llama all use "Decoder-only" architectures — their strength is generating fluent text.


What is the Attention Mechanism?

Attention is the core of Transformer and the key to its ability to handle long texts.

Analogy: Human Attention While Reading

When you read a sentence, you don't expend the same effort on every word.

For example: "It has a long tail and triangular ears" — when you see "it", you focus on "tail" and "ears" to determine what animal it is.

What the attention mechanism does is exactly this:Based on the current position being processed, it dynamically decides which words in the input to focus on.

注意力机制示意图

Animation demo:

The large model is thinking…

Each step does only one thing: predict the next word based on the previous context

Candidates for the next word Top-5 probabilities

How is Attention Computed?

In simple terms, each word generates three vectors:

  • Query:What information am I looking for?

  • Key:What information do I contain?

  • Value:What is my actual content?

Compute the similarity between the Query at each position and the Keys at all positions to obtain attention weights, then use these weights to compute a weighted sum of the Values, yielding the output for that position.

You don't need to remember the details, just remember this intuition:Attention = assigning a weight to each word in the sentence, indicating its importance to the current position.

Why is Attention So Important?

Attention brings several key advantages:

  • First, it can capture long-distance dependencies. No matter how far apart two words are, as long as they are related, attention can connect them.

  • Second, it supports parallel computation. Unlike RNNs, which process word by word, it can process the entire sentence simultaneously.

  • Third, a degree of interpretability. You can look at the attention weights and know what the model is focusing on.

For example, when translating "It has a long tail", you can see that when the model translates "it", the attention is mainly focused on "tail".


Pre-training and Fine-tuning

Today's large models typically use two-stage training: pre-training + fine-tuning.

Pre-training: Learning General Knowledge on Massive Data

Pre-training works like this: take large amounts of text from the internet (Wikipedia, books, web pages, code, etc.) and have the model do one thing — predict the next word.

  • You input "The weather today is really" and let it guess what the next word is.

  • You input "1 + 1 = " and let it guess what the next word is.

This process has no specific task; it simply lets the model broadly learn language, knowledge, logic, code, and so on.

Pre-training is the most expensive part: lots of data, massive computational power, and long duration.

Fine-tuning: Optimizing for Specific Tasks

The model after pre-training is very knowledgeable, but not necessarily obedient.

If you ask it a question, it might keep talking endlessly instead of answering directly.

If you ask it to write code, it might write halfway and then start writing something else.

Fine-tuning uses high-quality dialogue data to further train the model, teaching it to:

  • Understand instructions and output as required

  • Maintain conversational style and refuse harmful requests

  • Output safer and more useful content

  • The data volume for fine-tuning is much smaller than that for pre-training, but the quality requirements are higher.

Understand with an Analogy

  • Pre-training is like reading ten thousand books — from elementary school to university, broadly learning all kinds of knowledge.

  • Fine-tuning is like vocational training — after graduation, you go to work at a company and learn how to communicate with colleagues, how to write emails, and how to complete work tasks.

Pretraining gives the model knowledge; fine-tuning makes the model usable.

This is also why, among large models, some are smart but hard to use, while others are average but easy to use—the difference often lies in fine-tuning.


Why Does AI Hallucinate?

Hallucination is one of AI's most famous problems: it will seriously fabricate facts, names, papers, and data that do not exist.

Understanding from the Perspective of Probabilistic Prediction

To understand hallucination, first go back to AI's most fundamental way of working:

It is not verifying facts, but predicting the most likely next word.

You ask: Who invented the telephone? — It knowsAlexander Graham BellThis sequence is the most likely to appear.

But you ask: Who invented the example tutorial? — There may be no such information in the training data, but it will fabricate a plausible-sounding name based on statistical patterns.

It doesn't know what it doesn't know—it only knows that at this position, these words have the highest probability of appearing.

Common Causes of Hallucination

CauseExplanationExample
Insufficient training dataThis topic appears rarely in the training data, so the model hasn't learned it wellAsking a very niche professional question
Knowledge cutoffThings that happened after the training data cutoff are unknown to the modelAsking "What happened in 2025" (assuming the model was trained in 2024)
Confusing sourcesMixing information from multiple sources together and getting it wrong"Zhang San wrote Paper A" (actually Li Si wrote it, but Zhang San and Li Si often appear together)
OverfittingMemorized noise during training and treated it as factFabricating citations that don't exist

How to Reduce the Impact of Hallucination

As a user, you can do this:

  • First, verify key information. When it involves medicine, law, investment, news facts, etc., be sure to check.

  • Second, ask the model to provide sources. Ask it "What is the source of this information?" "Can you give me a citation?"

  • Third, use retrieval-augmented generation (RAG). Feed your own documents to the model and have it answer based on these documents, rather than making things up.

  • Fourth, use multiple models for cross-validation. Ask several models the same question; if they all say the same thing, it is more credible.

Remember: no matter how real what AI says sounds, it may be fabricated. For high-risk decisions, always remain skeptical.

Other extensions