LangChain Chat Model Common Parameters
In the previous section, we learnedinit_chat_model()the basic usage of
This section will detail the common control parameters used when calling models, helping you better control model behavior.
temperature - Controlling Creativity and Determinism
temperatureIt is the most commonly used parameter, with a value range of 0 to 2. It controls the randomness of the model output.
Example
# Comparison of different temperature values for the same question
question = "Introduce Python Tutorial Python in one sentence"
# temperature=0: Output is very deterministic, almost identical every time
model_low = init_chat_model("deepseek:deepseek-v4-flash", temperature=0)
resp1 = model_low.invoke(question)
resp2 = model_low.invoke(question)
print(f"temperature=0 1st time: {resp1.content}")
print(f"temperature=0 2nd time: {resp2.content}")
print(f"Both results are identical: {resp1.content == resp2.content}")
print()
# temperature=1.5: Output is diverse, may differ each time
model_high = init_chat_model("deepseek:deepseek-v4-flash", temperature=1.5)
resp1 = model_high.invoke(question)
resp2 = model_high.invoke(question)
print(f"temperature=1.5 1st time: {resp1.content}")
print(f"temperature=1.5 2nd time: {resp2.content}")
Output:
temperature=0 第1次: Example是一个为编程初学者提供免费教程的学习平台。 temperature=0 第2次: Example是一个为编程初学者提供免费教程的学习平台。 两次结果相同: True temperature=1.5 第1次: Example是一个简单易懂的编程与网络技术入门学习平台。 temperature=1.5 第2次: Example是深受编程新手喜爱的中文免费在线技术学习平台。
| temperature value | Effect | Use Cases |
|---|---|---|
| 0 ~ 0.3 | Stable and deterministic output, results are nearly identical each time | Data extraction, classification, code generation, translation |
| 0.5 ~ 0.7 | Moderate creativity, natural output that stays on topic | Daily conversation, content summarization |
| 0.8 ~ 1.2 | Diverse output with more room for creativity | Creative writing, brainstorming |
| 1.3 ~ 2.0 | Very random output, unexpected content may appear | Exploratory generation (not recommended for production) |
temperature=0 does not mean "exactly identical". Due to floating-point precision differences in the model's internal calculations, there may still be tiny differences in extreme cases. If absolute determinism is needed, some models provide a seed parameter.
max_tokens - Controlling Output Length and Cost
max_tokensLimits the maximum number of tokens in the model output. One token is approximately equivalent to 0.75 English words or 0.5 Chinese characters.
Example
model = init_chat_model("deepseek:deepseek-v4-flash", temperature=0)
# max_tokens=30: Limit output to within 30 tokens
response_short = model.invoke(
"Describe the Example Tutorial EXAMPLE platform in detail",
max_tokens=30
)
print(f"Limited to 30 tokens ({len(response_short.content)} characters):")
print(response_short.content)
print()
# max_tokens=200: Allow longer output
response_long = model.invoke(
"Describe the Example Tutorial EXAMPLE platform in detail",
max_tokens=200
)
print(f"Limited to 200 tokens ({len(response_long.content)} characters):")
print(response_long.content)
Output:
限制 30 tokens (42 字符): Example是一个面向编程初学者的在线学习平台,提供各种编程语言的教程和实例。 限制 200 tokens (306 字符): Example是一个面向编程初学者的在线学习平台,提供了丰富的编程教程和参考资料。 该平台涵盖了 HTML、CSS、JavaScript、Python、Java、C、C++、PHP、SQL 等多种编程语言和技术。 Example的特点包括: 1. 内容系统化:从基础到进阶,帮助学习者逐步掌握编程知识 2. 实例丰富:每个知识点都配有可运行的代码示例 3. ...
max_tokens is a hard limit on output length. If set too low, the model's response may be abruptly cut off mid-sentence. It is generally recommended to set it between 100-2000 depending on the scenario. For "one-sentence answer" scenarios, 30-100 is sufficient; for "detailed explanation" scenarios, 500-2000 is recommended.
timeout and max_retries - Network Reliability
In production environments, network requests may fail for various reasons. These two parameters help you control request behavior.
Example
# Recommended configuration for production environment
model = init_chat_model(
"deepseek:deepseek-v4-flash",
# Wait at most 30 seconds for a single request
timeout=30,
# Retry up to 3 times after failure (4 total request attempts)
max_retries=3,
)
# Simulate a normal call
try:
response = model.invoke("What is Python Tutorial Python?")
print(f"Call successful: {response.content[:50]}...")
except Exception as e:
print(f"Call failed: {e}")
| Parameter | Description | Recommended Value |
|---|---|---|
| timeout | Maximum waiting time for a single request (seconds). None means no limit | 30~60 (too short may cause timeouts; too long degrades user experience) |
| max_retries | Number of retries after failure. 0 means no retry | 2~3 (sufficient for handling occasional network issues) |
Time cost analysis for setting timeout and max_retries:
假设 timeout=30, max_retries=3: 单次成功: ~2s 第1次失败+重试成功: ~2s + ~2s = ~4s 全部失败: 4次 × 30s = 最多 120s 才抛出异常
base_url - Custom API Address
The base_url parameter is very useful when you need to access models through a proxy, relay service, or private deployment.
Example
# Scenario 1: Accessing OpenAI through a proxy
model = init_chat_model(
"deepseek:deepseek-v4-flash",
base_url="https://your-proxy-domain.com/v1", # Proxy address
)
# Scenario 2: Using a third-party service compatible with the OpenAI API
# Many domestic models provide OpenAI-compatible interfaces
model = init_chat_model(
"deepseek:deepseek-v4-flash", # Set provider to openai
base_url="https://api.third-party.com/v1", # But it actually points to a third party
api_key="your-third-party-key", # Third-party API Key
)
# Scenario 3: Connecting to local models (e.g., vLLM, Ollama)
model = init_chat_model(
"openai:qwen2.5", # Local model name
base_url="http://localhost:8000/v1", # Local service address
api_key="not-needed", # Local usually does not require a Key
)
base_url changes the API endpoint address, but the provider parameter determines the behavior mode. For example, provider="openai" will use OpenAI's message format, even if base_url points to a different service. Ensure the target service is compatible with the provider format you specify.
Other Common Parameters
top_p - Nucleus Sampling
Example
# top_p is another way to control randomness
# The model only samples from words whose cumulative probability reaches top_p
# top_p=0.1 means only selecting from the top 10% highest-probability words
model = init_chat_model(
"deepseek:deepseek-v4-flash",
top_p=0.9, # Only consider words in the top 90% of cumulative probability
)
response = model.invoke("Introduce Python Tutorial Python")
print(response.content[:100])
It is generally recommended to adjust only one of temperature or top_p, not both at the same time. If both are set, the model's behavior may be difficult to predict.
stop - Stop Sequence
Example
model = init_chat_model("deepseek:deepseek-v4-flash")
# The stop parameter specifies stop sequences; the model immediately stops generating when it encounters these words
response = model.invoke(
"List five programming learning websites, one per line",
stop=["\n"] # Stop when a newline is encountered, only return the first one
)
print(f"Limit stop=['\n']: {response.content}"\\n']: {response.content}")
Output:
限制 stop=['\n']: 1. Example
seed - Reproducibility (Supported by Some Models)
Example
# Some models support the seed parameter to obtain deterministic output
# Same seed + same input = same output
model = init_chat_model("deepseek:deepseek-v4-flash", seed=42, temperature=0)
resp1 = model.invoke("Introduce Python Tutorial")
resp2 = model.invoke("Introduce Python Tutorial")
print(f"seed=42, results identical: {resp1.content == resp2.content}")
Parameter Quick Reference Table
| Parameter | Type | Default Value | When to Use |
|---|---|---|---|
| temperature | float | Varies by model | Set to 0~0.3 when the task requires stability, and 0.7~1.0 when creativity is needed |
| max_tokens | int | Model upper limit | When output length needs to be controlled |
| timeout | int/float | None | Recommended to always set in production environments |
| max_retries | int | Varies by model | Recommended 2~3 when the network is unstable |
| base_url | str | Official address | When using a proxy, relay, or local service |
| top_p | float | 1.0 | When nucleus sampling control is needed (as an alternative to temperature) |
| stop | list[str] | None | When precise control over the output ending is needed |