How AI Works
You might ask: I just want to use AI, why should I understand how it works?
The reason is simple: if you know where AI's capabilities come from, you can use it better.
You will know when to trust it and when to question it.
-
You will know what it is good at and what it is not good at.
-
You will know where hallucinations come from and how to reduce their impact.
Intuitive Understanding of Neural Networks
The core of modern AI isneural networks, a name that comes from the neuron structure of the human brain.
Analogy to Brain Neurons
The human brain has about 86 billion neurons that are interconnected and transmit signals. Each neuron receives input from other neurons, processes it, and then outputs to other neurons. The learning process is the process of adjusting the strength of these connections.
Artificial neural networks borrow this idea but make it greatly simplified.
Three Basic Layers
A typical neural network is divided into three layers: input layer, hidden layer, and output layer.

We use the example of recognizing whether an image is a cat or a dog:

| Layer | Function | In this example |
|---|---|---|
| Input layer | Receives raw data | Each pixel and color information of the image |
| Hidden layer | Extract features layer by layer | Edges → Textures → Organs such as ears and eyes |
| Output layer | Give the final result | "85% probability it's a cat", "15% probability it's a dog" |
Each layer has many neurons. Each neuron receives the output from the previous layer, does a simple calculation, and passes it to the next layer.
-
The first layer might recognize "there is a vertical line here" or "there is a circle here".
-
The second layer combines these: a vertical line plus a circle might be an ear.
-
The third layer continues combining: two pointed ears, whiskers, cat eyes — this is very likely a cat.
The amazing part is that these features are not designed by humans; the model learns them from data on its own.
The Simplest Neuron
Let's use a few lines of Python code to show what a neuron does:
Example
# Basic computation logic of an artificial neuron
# No complex math, just weighted sum + activation
# ============================================
def simple_neuron(inputs: list, weights: list, bias: float) -> float:
"""
The simplest neuron
inputs: input values (from neurons in the previous layer)
weights: weights (importance of each input, learned during training)
bias: bias (threshold, learned during training)
"""
# Step 1: Weighted sum
# Multiply each input by its corresponding weight, then sum them all up
weighted_sum = 0.0
for input_value, weight in zip(inputs, weights):
weighted_sum += input_value * weight
# Add bias
weighted_sum += bias
# Step 2: Activation function (makes the output non-linear)
# Here we use the simplest ReLU: negative numbers become 0, positive numbers stay unchanged
output = max(0.0, weighted_sum)
return output
# Simulation: a neuron that determines "is this a cat's ear"
# Input: [pointiness, height position, has fur]
inputs = [0.8, 0.9, 0.7] # These three features are all quite obvious
# Weights: learned after training (in the example example, we assume these values are already well-learned)
weights = [0.5, 0.4, 0.3]
# Bias: threshold
bias = -0.6
result = simple_neuron(inputs, weights, bias)
print(f"Neuron output: {result:.3f}")
print(f"Judgment: {'likely a cat ear' if result > 0 else 'not likely'}")
# Output: Neuron output: 0.660
# Output: Judgment: likely a cat ear
What this neuron does is very simple: take a weighted sum of the inputs, pass it through an activation function, and output the result.
But when thousands of such neurons are connected together, with each layer learning different features, the whole system produces astonishing intelligence.
Remember this intuition: a neural network = many simple computing units connected together, learning by adjusting connection weights.
Training vs. Inference: Two Different Stages
The AI lifecycle is divided into two completely distinct stages: training and inference.
Understanding the difference between these two stages can help you understand many things—such as why training is so expensive, while inference is relatively cheap.
Training: Let AI Learn
Training is a process: show the model large amounts of data, let it continuously adjust parameters, and make predictions more and more accurately.
For example, training a model to recognize cats and dogs:
-
1. Prepare millions of labeled images (this one is a cat, that one is a dog)
-
2. Let the model guess "what is this"; at first it will guess wrong a lot
-
3. Tell it "you guessed wrong, it should be a cat", and let it adjust the weights in the network
-
4. Repeat millions of times until the model predicts more and more accurately
The training phase requires enormous computational power and data.A large model may need thousands of GPUs trained for several months, costing millions of dollars.
Inference: Let AI Use
Inference is a process: use the trained model, give new input, get output.
You send ChatGPT a message, it replies to you—this is inference.
You use your phone to take a photo to identify a plant—this is also inference.
The characteristics of inference:
No need to adjust parameters, only use the trained weights for computation.
-
Usually only one GPU or even a phone chip is needed.
-
The cost is much lower than training.
Comparison of the Two
| Dimension | Training | Inference |
|---|---|---|
| Goal | Learn knowledge, adjust weights | Apply learned knowledge, provide answers |
| Data volume | Needs massive amounts of data | A single input is enough |
| Compute requirements | Extremely high (thousands of GPUs) | Relatively low (single GPU or phone) |
| Cost | Extremely high (millions of dollars) | Relatively low (a few cents per time) |
| Frequency | A few times or dozens of times | Millions of times per second |
| Who does it | Companies like OpenAI, Anthropic | Ordinary users or applications |
As an analogy: training is like "studying hard for ten years", inference is like "solving problems in an exam".
Studying requires a lot of time and energy, but once you learn it, solving problems becomes fast.
When you use ChatGPT, you are doing "inference"—the model does not "learn" or "become smarter" from your conversations. Its knowledge ends at the moment training is completed.
Introduction to the Transformer Architecture
In 2017, Google published a paper titled "Attention Is All You Need", proposing the Transformer architecture.
This paper changed the entire AI field. Today's large language models are almost all based on Transformer.
Why is Transformer So Important?
Before Transformer, sequence data (such as sentences) was processed using RNN or LSTM.
Their problem is: they can only process one word at a time, making it difficult to capture long-distance dependencies.
For example: "I left my wallet at a café in Beijing. When I went back to look for it the next day, ____ was still there." — fill in "it" in the blank. You know "it" refers to "wallet" because you remember the earlier content.
When old models processed "still there", they might have already forgotten the existence of "wallet".
Transformer's breakthrough lies in:It can see the entire sentence at once, and through the attention mechanism, know which words to focus on.
Encoder and Decoder
A complete Transformer is divided into two parts:
| Component | Function | Typical application scenarios |
|---|---|---|
| Encoder | Understands input and converts text into vector representations | Text classification, sentiment analysis, semantic search |
| Decoder | Generates output text based on understanding | Writing, translation, dialogue generation |
Some models use only the Encoder (such as BERT), some use only the Decoder (such as GPT), and some use both (such as T5).
The GPT series, Claude, and Llama all use "Decoder-only" architectures — their strength is generating fluent text.
What is the Attention Mechanism?
Attention is the core of Transformer and the key to its ability to handle long texts.
Analogy: Human Attention While Reading
When you read a sentence, you don't expend the same effort on every word.
For example: "It has a long tail and triangular ears" — when you see "it", you focus on "tail" and "ears" to determine what animal it is.
What the attention mechanism does is exactly this:Based on the current position being processed, it dynamically decides which words in the input to focus on.

Animation demo:
Each step does only one thing: predict the next word based on the previous context
How is Attention Computed?
In simple terms, each word generates three vectors:
-
Query:What information am I looking for?
-
Key:What information do I contain?
-
Value:What is my actual content?
Compute the similarity between the Query at each position and the Keys at all positions to obtain attention weights, then use these weights to compute a weighted sum of the Values, yielding the output for that position.
You don't need to remember the details, just remember this intuition:Attention = assigning a weight to each word in the sentence, indicating its importance to the current position.
Why is Attention So Important?
Attention brings several key advantages:
-
First, it can capture long-distance dependencies. No matter how far apart two words are, as long as they are related, attention can connect them.
-
Second, it supports parallel computation. Unlike RNNs, which process word by word, it can process the entire sentence simultaneously.
-
Third, a degree of interpretability. You can look at the attention weights and know what the model is focusing on.
For example, when translating "It has a long tail", you can see that when the model translates "it", the attention is mainly focused on "tail".
Pre-training and Fine-tuning
Today's large models typically use two-stage training: pre-training + fine-tuning.
Pre-training: Learning General Knowledge on Massive Data
Pre-training works like this: take large amounts of text from the internet (Wikipedia, books, web pages, code, etc.) and have the model do one thing — predict the next word.
You input "The weather today is really" and let it guess what the next word is.
-
You input "1 + 1 = " and let it guess what the next word is.
This process has no specific task; it simply lets the model broadly learn language, knowledge, logic, code, and so on.
Pre-training is the most expensive part: lots of data, massive computational power, and long duration.
Fine-tuning: Optimizing for Specific Tasks
The model after pre-training is very knowledgeable, but not necessarily obedient.
If you ask it a question, it might keep talking endlessly instead of answering directly.
If you ask it to write code, it might write halfway and then start writing something else.
Fine-tuning uses high-quality dialogue data to further train the model, teaching it to:
-
Understand instructions and output as required
-
Maintain conversational style and refuse harmful requests
-
Output safer and more useful content
-
The data volume for fine-tuning is much smaller than that for pre-training, but the quality requirements are higher.
Understand with an Analogy
-
Pre-training is like reading ten thousand books — from elementary school to university, broadly learning all kinds of knowledge.
-
Fine-tuning is like vocational training — after graduation, you go to work at a company and learn how to communicate with colleagues, how to write emails, and how to complete work tasks.
Pretraining gives the model knowledge; fine-tuning makes the model usable.
This is also why, among large models, some are smart but hard to use, while others are average but easy to use—the difference often lies in fine-tuning.
Why Does AI Hallucinate?
Hallucination is one of AI's most famous problems: it will seriously fabricate facts, names, papers, and data that do not exist.
Understanding from the Perspective of Probabilistic Prediction
To understand hallucination, first go back to AI's most fundamental way of working:
It is not verifying facts, but predicting the most likely next word.
You ask: Who invented the telephone? — It knowsAlexander Graham BellThis sequence is the most likely to appear.
But you ask: Who invented the example tutorial? — There may be no such information in the training data, but it will fabricate a plausible-sounding name based on statistical patterns.
It doesn't know what it doesn't know—it only knows that at this position, these words have the highest probability of appearing.
Common Causes of Hallucination
| Cause | Explanation | Example |
|---|---|---|
| Insufficient training data | This topic appears rarely in the training data, so the model hasn't learned it well | Asking a very niche professional question |
| Knowledge cutoff | Things that happened after the training data cutoff are unknown to the model | Asking "What happened in 2025" (assuming the model was trained in 2024) |
| Confusing sources | Mixing information from multiple sources together and getting it wrong | "Zhang San wrote Paper A" (actually Li Si wrote it, but Zhang San and Li Si often appear together) |
| Overfitting | Memorized noise during training and treated it as fact | Fabricating citations that don't exist |
How to Reduce the Impact of Hallucination
As a user, you can do this:
-
First, verify key information. When it involves medicine, law, investment, news facts, etc., be sure to check.
-
Second, ask the model to provide sources. Ask it "What is the source of this information?" "Can you give me a citation?"
-
Third, use retrieval-augmented generation (RAG). Feed your own documents to the model and have it answer based on these documents, rather than making things up.
-
Fourth, use multiple models for cross-validation. Ask several models the same question; if they all say the same thing, it is more credible.
Other extensionsRemember: no matter how real what AI says sounds, it may be fabricated. For high-risk decisions, always remain skeptical.