AI API Development
It is convenient to use AI chat tools in a web interface, but to truly integrate AI capabilities into your own products, workflows, and automation scripts, you need an API.
An API (Application Programming Interface) is a bridge for communication between applications. Through an API, your code can directly send requests to an AI model and receive replies, without needing to manually open a webpage and copy-paste.
Imagine these scenarios:
Your e-commerce website automatically generates descriptive copy for each product.
-
Your note-taking app automatically summarizes long text entered by users.
-
Your customer service system automatically categorizes and replies to user inquiries.
-
All these features can be implemented through AI APIs.
The goal of this module is to take you from "only being able to chat with AI" to "being able to make AI work in your code."
API Basic Concepts
Before writing code, first understand a few core concepts.
What is an API
An API is a set of agreed-upon communication protocols. We send requests in a specific format, and the other party returns results in a specific format.
Use ordering food as an analogy:
-
You walk into a restaurant, get the menu (API documentation), and know what you can order and how to order (request format).
-
You say to the waiter: "One Kung Pao Chicken, less spicy" (sending an API request).
-
The kitchen prepares the dish and brings it to you (returning the API response).
In this process, you don't need to enter the kitchen, and you don't need to know how the dish is made—this is the value of an API:Encapsulate complex details, expose only a simple interface。
REST API Fundamentals
REST (Representational State Transfer) is currently the most popular API design style.
Simply put, it uses the HTTP protocol for communication:
| HTTP Method | Purpose | Example in AI APIs |
|---|---|---|
| POST | Create a resource | Send chat requests, generate images |
| GET | Retrieve a resource | Query model lists, get billing information |
| DELETE | Delete a resource | Delete conversation history |
The vast majority of AI APIs use POST requests, because we need to "send content to AI for processing."
JSON Format Basics
JSON (JavaScript Object Notation) is the most commonly used data format in API communication.
JSON's syntax is very simple; it's a combination of key-value pairs.
Examples
"model": "gpt-4",
"messages": [
{
"role": "user",
"content": "Hello, please introduce EXAMPLE"
}
],
"temperature": 0.7,
"max_tokens": 1000
}
This is a typical AI API request body.
Several core rules:
-
For objects, use
{ }to wrap, for arrays use[ ]to wrap. -
Strings use double quotes
", cannot use single quotes. -
Key-value pairs use a colon
:to separate, keys must be strings. -
Multiple key-value pairs use a comma
,to separate.
JSON looks a lot like Python dictionaries, but there are syntactic differences; pay attention when writing code.
Development Environment Setup
To do a good job, one must first sharpen one's tools.
Python Installation and Configuration
Python is the most commonly used language for calling AI APIs, with a mature ecosystem and rich libraries.
First, check whether Python is installed on your computer:
# 检查 Python 版本(Windows、macOS、Linux 通用) python --version # 或者 python3 --version
It is recommended to use Python 3.9 or higher.
If not installed, go topython.orgDownload the latest stable version.
pip Package Management
pip is Python's package manager, used to install third-party libraries.
# 检查 pip 版本 pip --version # 或者 pip3 --version # 升级 pip 到最新版本 pip install --upgrade pip # 配置国内镜像源(可选,提升下载速度) pip config set global.index-url https://pypi.tuna.tsinghua.edu.cn/simple
Code Editor Selection
For writing AI API code, we recommend one of the following editors:
| Editor | Features | Target Users |
|---|---|---|
| VS Code | Free, rich plugin ecosystem, lightweight | Most developers |
| Cursor | Built-in AI assistance, can directly generate code | For those who want AI to help write code |
| PyCharm | Most feature-rich, best Python support | Professional Python developers |
The example code in this tutorial can run in any editor.
OpenAI API Basics
The OpenAI API is the most classic and well-established AI API, and many other vendors are compatible with its format.
Registration and Getting API Key
The API Key is your pass, proving your identity and recording your usage. Steps to get one:
-
1. Visitplatform.openai.comRegister an account
-
2. Go to the API Keys page
-
3. Click "Create new secret key"
-
4. Copy and save this Key (it is only shown once!)
Important:Never commit your API Key to a public code repository. Once leaked, others may use your quota and incur charges.
First API Call
First, install the official OpenAI SDK:
pip install openai
Now write your first API call program:
Examples
# First OpenAI API call example
from openai import OpenAI
# Initialize the client
# Note: It is recommended to read the API Key from environment variables; do not hardcode it in code
# For demonstration purposes, it is written directly here. In real projects, use environment variables.
client = OpenAI(
api_key="your-api-key-here" # Replace with your API Key
)
# Send a chat request
response = client.chat.completions.create(
model="gpt-3.5-turbo", # Select the model
messages=[
{
"role": "user", # Role: user represents the user
"content": "Please introduce EXAMPLE's Rookie Tutorial in one sentence." # User input content
}
]
)
# Print the full response (see what was returned)
print("Full response:")
print(response)
print("\n" + "="*50 + "\n")
# Extract the AI's reply
ai_reply = response.choices[0].message.content
print("AI reply:")
print(ai_reply)
If you don't have an OpenAI api_key, you can also use domestic alternatives. Many domestic models are compatible with OpenAI, such as DeepSeek. You can go tohttps://platform.deepseek.com/api_keysapply for an api_key and then replace yours with the following codeapi_keyandmodel, DeepSeek supports the models deepseek-v4-flash and deepseek-v4-pro (thinking mode), and you can use the API to call large models:
Examples
# Initialize the client
# Note: It is recommended to read the API Key from an environment variable, do not hardcode it in the code
# For demonstration purposes, it is written directly here. In actual projects, please use environment variables
client = OpenAI(
api_key="sk-xxxx", # Set the api_key
base_url="https://api.deepseek.com") # Set the request address
# Send a chat request
response = client.chat.completions.create(
model="deepseek-v4-flash", # Select the model
messages=[
{
"role": "user", # Role: user represents the user
"content": "Please introduce EXAMPLE's Rookie Tutorial in one sentence." # User input content
}
]
)
# Print the full response (see what was returned)
print("Full response:")
print(response)
print("\n" + "="*50 + "\n")
# Extract the AI's reply
ai_reply = response.choices[0].message.content
print("AI reply:")
print(ai_reply)
Run this program:
python first_api_call.py
If everything works normally, you will see output similar to this:
完整响应: ChatCompletion( id='chatcmpl-xxx', choices=[Choice(finish_reason='stop', index=0, message=ChatMessage(content='Example是一个专注于提供编程技术教程的学习平台...', role='assistant'))], model='gpt-3.5-turbo', usage=CompletionUsage(completion_tokens=30, prompt_tokens=18, total_tokens=48) ) ================================================== AI 回复: Example是一个专注于提供编程技术教程的学习平台,涵盖多种编程语言和技术领域。
Congratulations! You have successfully called an AI API.
Request Parameters Explained
The OpenAI API has many parameters that can be adjusted. Here are the most commonly used ones:
| Parameter Name | Type | Required | Description | Default Value |
|---|---|---|---|---|
| model | string | Yes | The model ID to use, such as gpt-3.5-turbo, gpt-4 | None |
| messages | array | Yes | List of conversation history messages | None |
| temperature | number | no | Sampling temperature, between 0-2. The higher it is, the more random; the lower it is, the more deterministic. | 1.0 |
| max_tokens | integer | no | Maximum number of tokens to generate | Unlimited |
| top_p | number | no | Nucleus sampling parameter, between 0-1 | 1.0 |
| stop | string/array | no | Stop sequence; stop generation when these characters are encountered | null |
Let's focus on temperature, which is the most commonly used adjustment parameter:
Examples
# Demonstrate the effect of the temperature parameter on output
from openai import OpenAI
client = OpenAI(api_key="your-api-key-here")
def ask_ai(temperature_value: float, question: str) -> str:
"""Specify the temperature for the question"""
response = client.chat.completions.create(
model="gpt-3.5-turbo",
messages=[{"role": "user", "content": question}],
temperature=temperature_value,
max_tokens=200
)
return response.choices[0].message.content
# Ask the same question three times with different temperature
question = "Give EXAMPLE a slogan"
print("--- temperature=0 (most deterministic, results are similar each time) ---")
print(ask_ai(0.0, question))
print("\n--- temperature=0.7 (balanced, recommended for daily use) ---")
print(ask_ai(0.7, question))
print("\n--- temperature=1.8 (most random, results vary greatly each time) ---")
print(ask_ai(1.8, question))
A rule of thumb:
-
temperature=0: factual answers, code generation, scenarios requiring precise results
-
temperature=0.7: everyday conversation, general tasks
-
temperature>1.0: creative writing, brainstorming
Anthropic Claude API
Claude is Anthropic's large language model, known for its safety and long-text processing capabilities.
Similarities and Differences with OpenAI API
Both can do text generation, but there are some differences in design philosophy:
| Comparison item | OpenAI API | Claude API |
|---|---|---|
| Message structure | system, user, assistant alternate | system is set separately, user and assistant alternate |
| Context length | 16k-128k varies | Claude 3 series supports 200k |
| Parameter design | temperature, top_p, etc. | temperature, top_p, etc., similar concepts |
Claude Messages API Structure
First install the Claude SDK:
pip install anthropic
Basic calls to the Claude API:
Examples
# Basic Claude API call example
import anthropic
# Initialize the client
client = anthropic.Anthropic(
api_key="your-api-key-here" # Replace with your Claude API Key
)
# Send message request
response = client.messages.create(
model="claude-3-sonnet-20240229", # Select the Claude model
max_tokens=1024, # Claude requires max_tokens to be set
messages=[
{
"role": "user",
"content": "Please introduce EXAMPLE Rookie Tutorial in three sentences"
}
]
)
# Print response
print("Full response:")
print(response)
print("\n" + "="*50 + "\n")
# Extract reply
print("Claude's reply:")
print(response.content[0].text)
Similarly, we can replace it with a domestically compatible large model, such as DeepSeek, specifyingapi_key 、 base_url、modelThese three parameters:
Examples
# Initialize client
client = anthropic.Anthropic(
api_key="sk-xxx", # Replace with your API Key
base_url="https://api.deepseek.com/anthropic" # Base URL for Claude API
)
# Send message request
response = client.messages.create(
model="deepseek-v4-flash", # Choose Claude model
max_tokens=1024, # Claude requires max_tokens to be set
messages=[
{
"role": "user",
"content": "Please introduce the EXAMPLE tutorial in three sentences."
}
]
)
# Print response
print("Full response:")
print(response)
print("\n" + "="*50 + "\n")
# Extract reply
print("Claude reply:")
print(response.content[0].thinking)
System Prompt Settings
In the Claude API, System Prompt is an independent parameter, not placed in the messages array:
Examples
# Claude System Prompt usage example
import anthropic
client = anthropic.Anthropic(api_key="your-api-key-here")
response = client.messages.create(
model="claude-3-sonnet-20240229",
max_tokens=1024,
# System Prompt is set here
system="You are a professional technical documentation editor. Keep answers concise and accurate, and use lists where appropriate.",
messages=[
{
"role": "user",
"content": "What core knowledge points should be mastered when learning Python?"
}
]
)
print(response.content[0].text)
Multi-turn Conversation Management
Single-turn conversations are simple, but truly useful applications need to "remember context."
Ways to Maintain Conversation History
The approach is simple:Store all previous conversations and send them to the AI with each request。
Examples
# Multi-turn conversation management example
from openai import OpenAI
client = OpenAI(api_key="your-api-key-here")
class ChatBot:
"""A simple multi-turn conversation bot"""
def __init__(self, system_prompt: str = "You are a helpful assistant"):
self.messages = [] # Use a list to store conversation history
# First add the system prompt
self.messages.append({
"role": "system",
"content": system_prompt
})
def chat(self, user_input: str) -> str:
"""Send a message and get a reply"""
# Add user message to history
self.messages.append({
"role": "user",
"content": user_input
})
# Call the API
response = client.chat.completions.create(
model="gpt-3.5-turbo",
messages=self.messages
)
# Get AI reply
ai_reply = response.choices[0].message.content
# Also add the AI reply to history
self.messages.append({
"role": "assistant",
"content": ai_reply
})
return ai_reply
def get_history(self) -> list:
"""Get conversation history"""
return self.messages
# Usage example
if __name__ == "__main__":
bot = ChatBot(system_prompt="You are a programming assistant specializing in answering questions about EXAMPLE and Python")
print("=== Multi-turn conversation demo ===")
# First round
r1 = bot.chat("What is Python?")
print(User: What is Python?)
print("AI:", r1)
print()
# Second round (AI can remember the previous topic)
r2 = bot.chat(What are its features?)
print(User: What are its features?)
print("AI:", r2)
print()
# Third round
r3 = bot.chat(Are there Python tutorials on EXAMPLE?)
print(User: Are there Python tutorials on EXAMPLE?)
print("AI:", r3)
Key point: The conversation history is an ordered list, strictly alternating between user-assistant-user-assistant. Each request must send the full history so the AI can understand the context. History consumes tokens, so it needs to be cleaned up regularly.
Handling Context Window Overflow
Each model has a context window limit, e.g., gpt-3.5-turbo is 16k, gpt-4 is 8k or 32k.
When the conversation history is too long, an error is triggered.
Several common handling strategies:
Examples
# Example of context window management strategies
from openai import OpenAI
client = OpenAI(api_key="your-api-key-here")
class ChatBotWithContextLimit:
"""Chatbot with context limit"""
def __init__(self, system_prompt: str, max_messages: int = 10):
self.system_prompt = system_prompt
self.max_messages = max_messages # How many messages to keep at most
self.messages = []
def _trim_messages(self):
"""Trim message list, keep the most recent messages"""
# Always keep the system prompt, then only keep the most recent max_messages
if len(self.messages) > self.max_messages:
# Keep the most recent N messages
self.messages = self.messages[-self.max_messages:]
def chat(self, user_input: str) -> str:
# Add user message
self.messages.append({
"role": "user",
"content": user_input
})
# Trim history
self._trim_messages()
# Construct full request (system prompt + history)
full_messages = [
{"role": "system", "content": self.system_prompt}
] + self.messages
try:
response = client.chat.completions.create(
model="gpt-3.5-turbo",
messages=full_messages
)
ai_reply = response.choices[0].message.content
self.messages.append({
"role": "assistant",
"content": ai_reply
})
return ai_reply
except Exception as e:
# If it still errors, more aggressive trimming may be needed
return f"Error: {str(e)}"
# Strategy 2: Use AI to summarize conversation history
def summarize_conversation(messages: list) -> str:
"""Use AI to summarize the conversation history, replacing the original history"""
history_text = "\n".join([
f"{m['role']}: {m['content']}"
for m in messages
])
prompt = f"""Please summarize the following conversation into a concise summary, preserving key information:
{history_text}
Summary: """
response = client.chat.completions.create(
model="gpt-3.5-turbo",
messages=[{"role": "user", "content": prompt}],
temperature=0
)
return response.choices[0].message.content
# Use summarization strategy
print("--- Conversation history summary example ---")
sample_history = [
{"role": "user", "content": "What's my name?"},
{"role": "assistant", "content": "Sorry, I don't know your name. You can tell me."},
{"role": "user", "content": "My name is Xiao Ming, I'm learning Python at EXAMPLE."},
{"role": "assistant", "content": "Hello Xiao Ming! Nice to meet you. EXAMPLE is a great learning platform."},
{"role": "user", "content": "I want to learn about lists, can you teach me?"},
{"role": "assistant", "content": "Of course! Python lists are..."},
]
summary = summarize_conversation(sample_history)
print("Conversation summary:", summary)
Summarize the methods of context management:
| Strategy | Advantages | Disadvantages | Applicable scenarios |
|---|---|---|---|
| Keep only the most recent N entries | Simple, fast | May lose important information | Short conversations, casual chat scenarios |
| Use AI to summarize history | Preserves semantic information | Requires additional API calls | Long conversations, many important pieces of information |
| Sliding window | Balances simplicity and completeness | Slightly more complex to implement | General purpose for common scenarios |
Streaming Output
When chatting with AI on a web page, you see text appear word by word — that is streaming output.
What is Streaming Output
-
Normal mode: AI returns the complete reply at once after generating it; you may have to wait several seconds or even tens of seconds.
-
Streaming mode: AI sends each token as soon as it is generated, providing a better user experience.
Technically, streaming output uses the SSE (Server-Sent Events) protocol, where the server continuously pushes data.
Implementing the Typewriter Effect
Examples
# Streaming output demo (typewriter effect)
from openai import OpenAI
client = OpenAI(api_key="your-api-key-here")
print("=== Streaming Output Demo ===")
print("Q: Please write an introduction to EXAMPLE\n")
print("A:", end="", flush=True)
# Key parameter: stream=True
response = client.chat.completions.create(
model="gpt-3.5-turbo",
messages=[
{"role": "user", "content": "Please write an introduction to the EXAMPLE tutorial"}
],
stream=True # Enable streaming output
)
collected_reply = []
for chunk in response:
# Extract the content of the current chunk
if chunk.choices[0].delta.content:
content = chunk.choices[0].delta.content
print(content, end="", flush=True) # Print in real time without newlines
collected_reply.append(content)
print() # Output a newline at the end
print("\n" + "="*50)
print("Complete reply:", "".join(collected_reply))
Run this program, and you will see text appear word by word.
Claude's streaming output usage is similar:
Examples
# Claude streaming output example
import anthropic
client = anthropic.Anthropic(api_key="your-api-key-here")
print("=== Claude Streaming Output ===")
print("Q: How to learn programming?\n")
print("A:", end="", flush=True)
with client.messages.stream(
model="claude-3-sonnet-20240229",
max_tokens=1024,
messages=[{"role": "user", "content": "How to learn programming systematically?"}]
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
print()
User experience advice: As long as the scenario does not require "must wait for the complete result", prioritize streaming output. Letting users see progress can significantly improve the experience.
API Cost Calculation and Optimization
AI APIs are billed based on usage; the more you use, the more you spend.
Token Billing Principles
AI API is not billed by word count or number of conversations, but byToken,Tokenis the basic unit of AI text processing.
One token is approximately equal to:
English: 0.75 words (or 4 letters)
-
Chinese: 1-2 Chinese characters
For example:
-
"Hello, World" — approximately 3-4 tokens
-
"EXAMPLE Rookie Tutorial is great" — approximately 6-8 tokens
Billing method:
Prompt tokens(你发给 AI 的) + Completion tokens(AI 返回给你的)= 总 tokens
Different models have different prices, for example gpt-3.5-turbo is $0.0015/1k prompt tokens, $0.002/1k completion tokens.
How to Control API Costs
Examples
# API cost tracking and control example
from openai import OpenAI
client = OpenAI(api_key="your-api-key-here")
# Model price list (example, actual prices please check official documentation)
# Unit: USD / 1k tokens
MODEL_PRICES = {
"gpt-3.5-turbo": {
"prompt": 0.0015,
"completion": 0.002
},
"gpt-4": {
"prompt": 0.03,
"completion": 0.06
}
}
def chat_with_cost_tracking(model: str, messages: list):
"""Chat with cost statistics"""
response = client.chat.completions.create(
model=model,
messages=messages
)
# Get token usage from the response
usage = response.usage
prompt_tokens = usage.prompt_tokens
completion_tokens = usage.completion_tokens
total_tokens = usage.total_tokens
# Calculate the cost
price = MODEL_PRICES.get(model, {"prompt": 0, "completion": 0})
prompt_cost = (prompt_tokens / 1000) * price["prompt"]
completion_cost = (completion_tokens / 1000) * price["completion"]
total_cost = prompt_cost + completion_cost
return {
"reply": response.choices[0].message.content,
"usage": {
"prompt_tokens": prompt_tokens,
"completion_tokens": completion_tokens,
"total_tokens": total_tokens
},
"cost": {
"prompt_cost": prompt_cost,
"completion_cost": completion_cost,
"total_cost": total_cost
}
}
# Usage example
if __name__ == "__main__":
result = chat_with_cost_tracking(
model="gpt-3.5-turbo",
messages=[{"role": "user", "content": "Introduce Python"}]
)
print("AI reply:", result["reply"])
print()
print("Usage statistics:")
print(f" Prompt tokens: {result['usage']['prompt_tokens']}")
print(f" Completion tokens: {result['usage']['completion_tokens']}")
print(f" Total tokens: {result['usage']['total_tokens']}")
print()
print("Cost statistics (USD):")
print(f" Prompt cost: ${result['cost']['prompt_cost']:.6f}")
print(f" Completion cost: ${result['cost']['completion_cost']:.6f}")
print(f" Total cost: ${result['cost']['total_cost']:.6f}")
Some practical tips for cost optimization:
| Tip | How to do it | Savings ratio |
|---|---|---|
| Choose the right model | Use gpt-3.5-turbo for simple tasks, not gpt-4 | 90%+ |
| Limit max_tokens | Set a reasonable upper limit to avoid the AI saying too much | Depends on the situation |
| Trim context | Keep only necessary conversation history | 50%+ |
| Keep prompts concise | Don't make the system prompt too long | Depends on the situation |
| Result caching | Return cached results directly for identical questions | Depends on the situation |
Caching Strategies
For repeated questions, caching results can save a lot of money:
Examples
# Simple response cache implementation
import hashlib
import json
from openai import OpenAI
client = OpenAI(api_key="your-api-key-here")
class SimpleCache:
"""Simple in-memory cache"""
def __init__(self):
self.cache = {}
def _make_key(self, model: str, messages: list) -> str:
"""Generate cache key based on request content"""
# Serialize key information and compute a hash
key_data = {
"model": model,
"messages": messages
}
key_str = json.dumps(key_data, sort_keys=True)
return hashlib.md5(key_str.encode()).hexdigest()
def get(self, model: str, messages: list):
"""Get cache"""
key = self._make_key(model, messages)
return self.cache.get(key)
def set(self, model: str, messages: list, value: str):
"""Set cache"""
key = self._make_key(model, messages)
self.cache[key] = value
# Chat function with cache
cache = SimpleCache()
def chat_with_cache(model: str, messages: list):
# Check cache first
cached = cache.get(model, messages)
if cached:
print("(Cache hit, return directly)")
return cached
# Cache miss, call API
response = client.chat.completions.create(
model=model,
messages=messages
)
reply = response.choices[0].message.content
# Store in cache
cache.set(model, messages, reply)
return reply
# Test
if __name__ == "__main__":
messages = [{"role": "user", "content": "What kind of website is EXAMPLE?"}]
print("First call...")
r1 = chat_with_cache("gpt-3.5-turbo", messages)
print(r1)
print()
print("Second call (same question)...")
r2 = chat_with_cache("gpt-3.5-turbo", messages)
print(r2)
Error Handling and Retry Mechanisms
Network requests can always fail, and APIs are no exception. Robust code must handle various exception scenarios.
Common Error Types
| Error type | HTTP status code | Cause | Handling method |
|---|---|---|---|
| Authentication failure | 401 | API Key incorrect or expired | Check API Key |
| Insufficient quota | 402 | Account has no money | Recharge or switch account |
| Limit exceeded | 429 | Request too fast or quota exhausted | Wait a while and try again |
| Server error | 5xx | AI server-side issue | Retry |
Exponential Backoff Retry
When encountering transient errors (such as 429, 5xx), the most common strategy is exponential backoff retry:
-
First failure → wait 1 second → retry
-
Second failure → wait 2 seconds → retry
-
Third failure → wait 4 seconds → retry
-
Fourth failure → wait 8 seconds → retry
Each wait time doubles, giving the service a chance to recover.
Examples
# Error handling and exponential backoff retry
import time
from openai import OpenAI, APIError, RateLimitError, APIConnectionError
client = OpenAI(api_key="your-api-key-here")
def chat_with_retry(
model: str,
messages: list,
max_retries: int = 3,
initial_wait: float = 1.0
):
"""Chat function with retry mechanism"""
retries = 0
while retries <= max_retries:
try:
response = client.chat.completions.create(
model=model,
messages=messages
)
return {
"success": True,
"reply": response.choices[0].message.content
}
except RateLimitError as e:
# Rate limit error: wait a while and try again
retries += 1
if retries > max_retries:
return {
"success": False,
"error": f"Retry attempts exhausted: {str(e)}"
}
wait_time = initial_wait * (2 ** (retries - 1))
print(f"Rate limit triggered, waiting {wait_time} seconds before retrying...")
time.sleep(wait_time)
except APIConnectionError as e:
# Network connection error: retry
retries += 1
if retries > max_retries:
return {
"success": False,
"error": f"Network connection failed: {str(e)}"
}
wait_time = initial_wait * (2 ** (retries - 1))
print(f"Network error, waiting {wait_time} seconds before retrying...")
time.sleep(wait_time)
except APIError as e:
# API error: retry depending on the situation
if e.status_code and 500 <= e.status_code < 600:
retries += 1
if retries > max_retries:
return {
"success": False,
"error": f"Server error: {str(e)}"
}
wait_time = initial_wait * (2 ** (retries - 1))
print(f"Server error, waiting {wait_time} seconds before retrying...")
time.sleep(wait_time)
else:
# Non-5xx errors, do not retry
return {
"success": False,
"error": fAPI error: {str(e)}
}
except Exception as e:
# Other errors, do not retry
return {
"success": False,
"error": fUnknown error: {str(e)}
}
# Usage example
if __name__ == "__main__":
result = chat_with_retry(
model="gpt-3.5-turbo",
messages=[{"role": "user", "content": "Hello, EXAMPLE"}]
)
if result["success"]:
print(Success:, result["reply"])
else:
print(Failure:, result["error"])
You can also use an existing library to simplify retry logic, such as tenacity, install it with the following command:
pip install tenacity
Examples
# Use tenacity library to simplify retry logic
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
from openai import OpenAI, RateLimitError, APIError
client = OpenAI(api_key="your-api-key-here")
@retry(
stop=stop_after_attempt(3), # Retry up to 3 times
wait=wait_exponential(multiplier=1, min=1, max=10), # Exponential backoff: 1, 2, 4... up to 10 seconds
retry=retry_if_exception_type((RateLimitError, APIError)) # Retry only on specific exceptions
)
def chat_with_tenacity(model: str, messages: list):
"""Use tenacity decorator to implement retry"""
response = client.chat.completions.create(
model=model,
messages=messages
)
return response.choices[0].message.content
# Usage
try:
reply = chat_with_tenacity(
"gpt-3.5-turbo",
[{"role": "user", "content": "Introduce Python"}]
)
print(reply)
except Exception as e:
print(fFinal failure: {e})
Hands-on Project: Command-line AI Chat Tool
Now let's integrate what we've learned earlier and build a complete command line chat tool.
Examples
# EXAMPLE AI Chatbot - Complete Command Line Tool
import sys
import json
import os
from datetime import datetime
from openai import OpenAI
class ExampleChatBot:
"""A fully functional command line chatbot"""
def __init__(self, config_file: str = "config.json"):
# Load configuration
self.config = self._load_config(config_file)
self.client = OpenAI(api_key=self.config["api_key"])
# Conversation state
self.messages = []
self.total_tokens = 0
self.total_cost = 0.0
# Initialize system prompt
self._init_system_prompt()
def _load_config(self, config_file: str) -> dict:
"""Load configuration file"""
default_config = {
"api_key": "your-api-key-here",
"model": "gpt-3.5-turbo",
"system_prompt": "You are a helpful AI assistant powered by EXAMPLE.",
"max_history": 20,
"stream": True,
"temperature": 0.7
}
if os.path.exists(config_file):
with open(config_file, "r", encoding="utf-8") as f:
user_config = json.load(f)
default_config.update(user_config)
return default_config
def _save_config(self, config_file: str = "config.json"):
"""Save configuration file"""
with open(config_file, "w", encoding="utf-8") as f:
json.dump(self.config, f, indent=2, ensure_ascii=False)
def _init_system_prompt(self):
"""Initialize system prompt"""
self.messages = [
{
"role": "system",
"content": self.config["system_prompt"]
}
]
def _trim_history(self):
"""Trim conversation history"""
# Keep system prompt + the most recent N conversations
max_messages = self.config["max_history"]
if len(self.messages) > max_messages + 1:
self.messages = [self.messages[0]] + self.messages[-max_messages:]
def chat(self, user_input: str) -> str:
"""Send message and get reply"""
# Add user message
self.messages.append({
"role": "user",
"content": user_input
})
# Trim history
self._trim_history()
if self.config["stream"]:
return self._chat_stream()
else:
return self._chat_normal()
def _chat_normal(self) -> str:
"""Normal mode (non-streaming)"""
response = self.client.chat.completions.create(
model=self.config["model"],
messages=self.messages,
temperature=self.config["temperature"]
)
ai_reply = response.choices[0].message.content
# Count usage
if hasattr(response, "usage"):
self.total_tokens += response.usage.total_tokens
# Roughly estimate cost (gpt-3.5-turbo price)
self.total_cost += (response.usage.total_tokens / 1000) * 0.002
self.messages.append({
"role": "assistant",
"content": ai_reply
})
return ai_reply
def _chat_stream(self) -> str:
"""Streaming mode"""
response = self.client.chat.completions.create(
model=self.config["model"],
messages=self.messages,
temperature=self.config["temperature"],
stream=True
)
collected_chunks = []
print("AI:", end="", flush=True)
for chunk in response:
if chunk.choices[0].delta.content:
content = chunk.choices[0].delta.content
print(content, end="", flush=True)
collected_chunks.append(content)
print()
ai_reply = "".join(collected_chunks)
self.messages.append({
"role": "assistant",
"content": ai_reply
})
# Token counting in streaming mode requires additional handling, simplified here
# In real projects, you can call the usage API or estimate
self.total_tokens += len(ai_reply) // 2 # Rough estimate
return ai_reply
def save_conversation(self, filename: str = None):
"""Save conversation history"""
if not filename:
timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
filename = f"conversation_{timestamp}.json"
data = {
"model": self.config["model"],
"total_tokens": self.total_tokens,
"estimated_cost": self.total_cost,
"messages": self.messages,
"saved_at": datetime.now().isoformat()
}
with open(filename, "w", encoding="utf-8") as f:
json.dump(data, f, indent=2, ensure_ascii=False)
return filename
def clear_history(self):
"""Clear conversation history"""
self._init_system_prompt()
self.total_tokens = 0
self.total_cost = 0.0
def show_stats(self):
"""Display statistics"""
print("\n" + "="*40)
print(Conversation Statistics)
print("="*40)
print(fModel: {self.config['model']})
print(fMessage count: {len(self.messages) - 1})
print(fTotal tokens: {self.total_tokens})
print(fEstimated cost: ${self.total_cost:.4f})
print("="*40 + "\n")
def print_help():
"""Print help information"""
print("\n" + "="*40)
print(EXAMPLE AI Chat Tool - Command List)
print("="*40)
print(/help - Show help)
print(/clear - Clear conversation history)
print(/stats - Show statistics)
print(/save - Save conversation history)
print(/model - Switch model)
print(/temp - Adjust temperature)
print(/system - Modify system prompt)
print(/quit - Exit program)
print("="*40 + "\n")
def main():
print("="*40)
print(EXAMPLE AI Chat Tool)
print("="*40)
print(Enter /help to view commands, or type text directly to start chatting\n")
bot = ExampleChatBot()
while True:
try:
user_input = input(You:).strip()
if not user_input:
continue
# Handle commands
if user_input.startswith("/"):
cmd = user_input.lower()
if cmd in ["/quit", "/exit", "/q"]:
print(Goodbye!)
break
elif cmd == "/help":
print_help()
elif cmd == "/clear":
bot.clear_history()
print(Conversation history cleared)
elif cmd == "/stats":
bot.show_stats()
elif cmd == "/save":
filename = bot.save_conversation()
print(fConversation saved to: {filename})
elif cmd.startswith("/model "):
model = user_input[7:].strip()
bot.config["model"] = model
bot._save_config()
print(fModel switched to: {model})
elif cmd.startswith("/temp "):
try:
temp = float(user_input[6:].strip())
if 0 <= temp <= 2:
bot.config["temperature"] = temp
bot._save_config()
print(fTemperature set to: {temp})
else:
print(Temperature must be between 0-2)
except ValueError:
print(Please enter a valid number)
elif cmd.startswith("/system "):
system_prompt = user_input[8:].strip()
bot.config["system_prompt"] = system_prompt
bot.clear_history()
bot._save_config()
print(System prompt updated, conversation reset)
else:
print(Unknown command, enter /help for help)
else:
# Regular chat
if not bot.config["stream"]:
print("AI:", end="")
bot.chat(user_input)
print()
except KeyboardInterrupt:
print("\nEnter /quit to exit, or continue chatting)
except Exception as e:
print(fError: {e})
if __name__ == "__main__":
main()
Create a configuration file:
{
"api_key": "your-api-key-here",
"model": "gpt-3.5-turbo",
"system_prompt": "你是一个乐于助人的 AI 助手,由 EXAMPLE 提供技术支持。",
"max_history": 20,
"stream": true,
"temperature": 0.7
}
Now you can run this chat tool:
python example_chatbot.py
You will see:
======================================== EXAMPLE AI 聊天工具 ======================================== 输入 /help 查看命令,直接输入文字开始聊天 你:你好 AI:你好!很高兴见到你。我是由 EXAMPLE 提供技术支持的 AI 助手,有什么可以帮你的吗? 你:/stats ======================================== 对话统计 ======================================== 模型:gpt-3.5-turbo 消息数:2 总 tokens:50 估算费用:$0.0001 ======================================== 你:/quit 再见!Other extensions