Large Language Model Basics (LLM)
Large Language Model (LLM) is the brain of an AI Agent; understanding it is the foundation for building intelligent agents.
The reason why a large language model can converse with you, write articles, and program is essentially that it, based on the text (prompt) you provide, word by word,guessproduces the most reasonable continuation.
In simple terms, a large language model is a deep learning model trained on massive text data, capable of understanding and generating human language. By analyzing massive text on the internet, the model learns the statistical patterns of language, and when it receives input, it generates the most reasonable continuation according to the learned patterns.

We can imagine a large language model as aextremely hardworking student with an extraordinary memory:
Learning phase (training): It has read almost all publicly available text on the internet—books, articles, web pages, code, etc. (the data volume can reach trillions of words). In this process, it is not memorizing, but learning an extremely complex set of language rules.
Application phase (inference): When you ask it a question or give it an instruction, it applies the learned rules to generate, word by word, the most logical and contextually appropriate answer.
Its "largeness" is mainly reflected in two aspects:
- Large parameter scale: The model contains tens of billions or even trillions of adjustable parameters internally, which record the learned language knowledge.
- Large training data: The volume of text data used for training is enormous, covering the essence of publicly available internet information.
The figure below shows the process of an LLM generating text word by word—it predicts one word at a time, then incorporates this new word into the input, continues to predict the next word, and repeats this cycle until a complete answer is generated:

Animation demo:
Each step does only one thing: based on the preceding text, predict the next word.
Although LLMs are powerful, they also have clear limitations:
| Capability | Description | Limitation |
|---|---|---|
| Knowledge cutoff | Training data has a cutoff date | Cannot know new information after training |
| Mathematical calculations | Can do simple calculations | Complex calculations are prone to errors |
| Real-time information | Requires external tools for assistance | Cannot access real-time data by itself |
| Factual accuracy | May generate incorrect information | Requires fact-checking |
| Long text processing | Context length has limitations | Extremely long texts may lose information |
| Logical consistency | May be contradictory | Requires careful design and validation |
Important reminder:LLMs are not omniscient; they are essentially statistical pattern-matching systems. Only by understanding their limitations can we better leverage their capabilities.
Core Working Principle: A Brief Analysis of the Transformer Architecture
The remarkable capabilities of LLMs are inseparable from their underlying core technology—Transformer architectureThere's no need to delve into complex mathematical principles, but you can understand its core ideas.
Imagine you want to write an article about the solar system:
- Read through materialsYou would first look at many relevant books and web pages.
- Identify key pointsYou would notice that words like Sun, planets, orbits, and gravity appear frequently and are interrelated.
- Organize languageBased on the point you want to make (e.g., introducing Mars), you would selectively use previously seen information about Mars's size, color, location, etc., and organize it into coherent sentences.
The way Transformer works is similar; its core process is divided into three stages:
- Input processingYour words are split into tokens (words or characters) and converted into numbers (vectors) that computers can understand.
- Context understanding (core):Self-Attention mechanismstarts working. It allows the model, when processing each word in a sentence, to weigh the importance ofall other wordsin the sentence. This process isparalleland extremely fast.
- Generation and iteration: The model calculates a probability distribution based on its understanding of all words, and predicts the most likely next word. After selecting and outputting this word, it uses it as new input and repeats the entire process until a complete answer is generated.

Self-attention is the most critical innovation of the Transformer. Taking the sentence "Apple's phone has a very large battery" as an example, when the model processesitthis word, the self-attention mechanism helps the model determineitandAppleandphoneare highly correlated. The figure below shows the attention weight distribution in this process:
It is precisely this ability toprocess in parallelanddeeply understand global contextthat allows Transformer-based LLMs to far surpass previous technologies (such as RNN) on language tasks.
How to Interact with LLM: Introduction to Prompt Engineering
PromptIt is the input you give to the LLM. It tells the model what you want, just like giving instructions to an assistant—the clearer the instructions, the better the results. The quality of the prompt directly determines the quality of the answer.
A good prompt usually consists of the following four parts:
Basic Principles
Be clear and specific: Avoid vague expressions. Don't sayWrite something about dogs, but instead sayIn vivid and lively language, write a short introduction of about 100 words about the personality traits of Golden Retrievers for children aged 6-8.。
Provide contextTell the model your identity, background, and goals. For example:You are an experienced Python programming mentor. Please explain what list comprehensions are to a beginner who has just learned basic syntax, and provide a simple example.
Specify format: If you need output in a specific format, please state it clearly, for example:Please summarize the following points into three bullet pointsorPlease output in JSON format。
Step-by-step thinking (Chain-of-Thought): For complex problems, you can guide the model to reason step by step in the prompt, for example:Please analyze this problem step by step: first list the known conditions, then derive the intermediate steps, and finally give the conclusion.This approach can significantly improve the accuracy of complex reasoning tasks.
Common Application Scenarios of LLM
| Scenario category | Specific example | Description |
|---|---|---|
| Content creation and editing | Write emails, reports, blogs; continue stories; polish copy; translate texts in different styles | Quickly generate drafts, provide inspiration and multiple ways of expression |
| Information retrieval and summarization | Quickly read long documents and extract key points; Q&A based on knowledge bases | Understands questions better than traditional search, can summarize and integrate |
| Programming assistance | Explain code, generate code snippets, debug errors, refactor code, write test cases | Acts as a 24/7 programming partner, greatly improving development efficiency |
| Conversation and customer service | Intelligent chatbots, personalized tutors, role-playing | Provides anthropomorphic, contextually coherent interactive experiences |
| Logical reasoning and analysis | Solve math problems, perform basic logical reasoning, analyze data trends, make plans | Demonstrates surprising reasoning abilities within limited domains |
API Calls and Parameter Settings
To build an AI Agent, you need to learn how to call LLMs via API.
In this chapter, we use the OpenAI API as an example to introduce basic calling methods.
openai is a powerful Python library for interacting with a range of OpenAI models and services. For details, refer to:Python OpenAI。
Open-source address:https://github.com/openai/openai-python
Basic API Calls
1. Install the required libraries
pip install openai
Then you need to register an account on the OpenAI official website and generate an API Key on the API keys page.
Example
from openai import OpenAI
client = OpenAI(
# This is the default and can be omitted
api_key=The API key you applied for,
)
response = client.responses.create(
model="gpt-4o",
instructions="You are a coding assistant that talks like a pirate.",
input="How do I check if a Python object is an instance of a class?",
)
print(response.output_text)
Accessing OpenAI domestically is still a bit troublesome. Many domestic services also support OpenAI, such as DeepSeek and Alibaba's Qwen.
DeepSeek
The DeepSeek API is fully compatible with the OpenAI API format. You only need to modify a few configurations to directly use the OpenAI SDK or compatible tools to access the DeepSeek API.
Core configuration parameters:
| Parameter | Value/Description |
|---|---|
| base_url | Required, fixed value:https://api.deepseek.com(Can also fill inhttps://api.deepseek.com/v1, only for OpenAI compatibility; v1 is unrelated to model version) |
| api_key | Required. You need to apply for an exclusive API Key on the DeepSeek official website first (Application address:https://platform.deepseek.com/) |
| model | Required,
|
Example
from openai import OpenAI
# Initialize client (core configuration: replace with your API Key)
client = OpenAI(
api_key=os.environ.get('DEEPSEEK_API_KEY'), # Recommended to configure via environment variable, or hardcode directly (not recommended)
base_url="https://api.deepseek.com" # DeepSeek fixed domain
)
# Call the chat API
try:
response = client.chat.completions.create(
model="deepseek-v4-flash", # Specify the model, options: deepseek-v4-flash / deepseek-v4-pro
messages=[
{"role": "system", "content": "You are a helpful assistant"}, # System role definition
{"role": "user", "content": "Hello"}, # User question
],
stream=False # Non-streaming output (returns complete result at once)
)
# Print reply content
print("Reply result:", response.choices[0].message.content)
except Exception as e:
print("Call failed:", str(e))
Alibaba Bailian
Alibaba Cloud Bailian's Tongyi Qianwen model supports the OpenAI-compatible interface. You only need to adjust the API Key, BASE_URL, and model name to migrate your existing OpenAI code to Alibaba Cloud Bailian service.
We need to activate the Alibaba Cloud Bailian model service and obtain an API-KEY.
We can first use the Alibaba Cloud primary account to access the Bailian model service platform:https://bailian.console.aliyun.com/, then click Login in the upper right corner. After logging in, click the gear ⚙️ icon in the upper right corner, select API key, then copy the API key. If you don't have one, you can also create an API key:


Activating Alibaba Cloud Bailian does not incur fees. Only model calls (after exceeding the free quota), model deployment, and model fine-tuning will incur corresponding charges.
Now to use the API, you are billed by token. Fortunately, it's not expensive. We can first purchase the cheapest package:Alibaba Cloud Bailian Large Model Service Platform。
You can also directly use the Coding Plan package from Bailian and Ark:https://www.example.com/claude-code/coding-plan.html。
Usage
Next, we use the OpenAI SDK to access the Qwen model on Bailian.
Non-streaming Call Example
Example
import os
def get_response():
client = OpenAI(
api_key="sk-xxx", # Please use the Alibaba Cloud Bailian API Key
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1", # Fill in the base_url of the DashScope SDK
)
completion = client.chat.completions.create(
model="qwen-plus", # Here, qwen-plus is used as an example; you can change the model name as needed. Model list: https://help.aliyun.com/zh/model-studio/getting-started/models
messages=[{'role': 'system', 'content': 'You are a helpful assistant.'},
{'role': 'user', 'content': 'Who are you?'}]
)
# JSON data
#print(completion.model_dump_json())
print(completion.choices[0].message.content)
if __name__ == '__main__':
get_response()
Running the code produces the following result:
I am Tongyi Qianwen, a large-scale language model independently developed by the Tongyi Lab under Alibaba Group. I can help you answer questions and create text, such as writing stories, official documents, emails, scripts, logical reasoning, programming, and more. I can also express opinions and play games. If you have any questions or need help, feel free to let me know at any time!
Streaming Call Example
Example
def get_response():
client = OpenAI(
api_key="sk-xxx",
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
)
completion = client.chat.completions.create(
model="qwen-plus",
messages=[
{'role': 'system', 'content': 'You are a helpful assistant.'},
{'role': 'user', 'content': 'Who are you?'}
],
stream=True,
stream_options={"include_usage": True}
)
for chunk in completion:
# The chunk may not contain choices or delta
if hasattr(chunk, "choices") and len(chunk.choices) > 0:
choice = chunk.choices[0]
if hasattr(choice, "delta") and hasattr(choice.delta, "content"):
print(choice.delta.content, end='', flush=True)
if __name__ == '__main__':
get_response()
Running the code produces the following result:
I am Tongyi Qianwen, a large-scale language model independently developed by the Tongyi Lab under Alibaba Group. I can help you answer questions and create text, such as writing stories, official documents, emails, scripts, logical reasoning, programming, and more. I can also express opinions and play games. If you have any questions or need help, feel free to let me know at any time!
Mainstream Large Language Models
The following is a compilation of official websites and API documentation addresses for mainstream large language models:
