LangChain Chat Model Common Parameters

In the previous section, we learnedinit_chat_model()the basic usage of

This section will detail the common control parameters used when calling models, helping you better control model behavior.


temperature - Controlling Creativity and Determinism

temperatureIt is the most commonly used parameter, with a value range of 0 to 2. It controls the randomness of the model output.

Example

from langchain.chat_models import init_chat_model

# Comparison of different temperature values for the same question
question = "Introduce Python Tutorial Python in one sentence"

# temperature=0: Output is very deterministic, almost identical every time
model_low = init_chat_model("deepseek:deepseek-v4-flash", temperature=0)
resp1 = model_low.invoke(question)
resp2 = model_low.invoke(question)
print(f"temperature=0 1st time: {resp1.content}")
print(f"temperature=0 2nd time: {resp2.content}")
print(f"Both results are identical: {resp1.content == resp2.content}")
print()

# temperature=1.5: Output is diverse, may differ each time
model_high = init_chat_model("deepseek:deepseek-v4-flash", temperature=1.5)
resp1 = model_high.invoke(question)
resp2 = model_high.invoke(question)
print(f"temperature=1.5 1st time: {resp1.content}")
print(f"temperature=1.5 2nd time: {resp2.content}")

Output:

temperature=0 第1次: Example是一个为编程初学者提供免费教程的学习平台。
temperature=0 第2次: Example是一个为编程初学者提供免费教程的学习平台。
两次结果相同: True

temperature=1.5 第1次: Example是一个简单易懂的编程与网络技术入门学习平台。
temperature=1.5 第2次: Example是深受编程新手喜爱的中文免费在线技术学习平台。
temperature valueEffectUse Cases
0 ~ 0.3Stable and deterministic output, results are nearly identical each timeData extraction, classification, code generation, translation
0.5 ~ 0.7Moderate creativity, natural output that stays on topicDaily conversation, content summarization
0.8 ~ 1.2Diverse output with more room for creativityCreative writing, brainstorming
1.3 ~ 2.0Very random output, unexpected content may appearExploratory generation (not recommended for production)

temperature=0 does not mean "exactly identical". Due to floating-point precision differences in the model's internal calculations, there may still be tiny differences in extreme cases. If absolute determinism is needed, some models provide a seed parameter.


max_tokens - Controlling Output Length and Cost

max_tokensLimits the maximum number of tokens in the model output. One token is approximately equivalent to 0.75 English words or 0.5 Chinese characters.

Example

from langchain.chat_models import init_chat_model

model = init_chat_model("deepseek:deepseek-v4-flash", temperature=0)

# max_tokens=30: Limit output to within 30 tokens
response_short = model.invoke(
    "Describe the Example Tutorial EXAMPLE platform in detail",
    max_tokens=30
)
print(f"Limited to 30 tokens ({len(response_short.content)} characters):")
print(response_short.content)
print()

# max_tokens=200: Allow longer output
response_long = model.invoke(
    "Describe the Example Tutorial EXAMPLE platform in detail",
    max_tokens=200
)
print(f"Limited to 200 tokens ({len(response_long.content)} characters):")
print(response_long.content)

Output:

限制 30 tokens (42 字符):
Example是一个面向编程初学者的在线学习平台,提供各种编程语言的教程和实例。

限制 200 tokens (306 字符):
Example是一个面向编程初学者的在线学习平台,提供了丰富的编程教程和参考资料。
该平台涵盖了 HTML、CSS、JavaScript、Python、Java、C、C++、PHP、SQL 等多种编程语言和技术。
Example的特点包括:
1. 内容系统化:从基础到进阶,帮助学习者逐步掌握编程知识
2. 实例丰富:每个知识点都配有可运行的代码示例
3. ...

max_tokens is a hard limit on output length. If set too low, the model's response may be abruptly cut off mid-sentence. It is generally recommended to set it between 100-2000 depending on the scenario. For "one-sentence answer" scenarios, 30-100 is sufficient; for "detailed explanation" scenarios, 500-2000 is recommended.


timeout and max_retries - Network Reliability

In production environments, network requests may fail for various reasons. These two parameters help you control request behavior.

Example

from langchain.chat_models import init_chat_model

# Recommended configuration for production environment
model = init_chat_model(
    "deepseek:deepseek-v4-flash",

    # Wait at most 30 seconds for a single request
    timeout=30,

    # Retry up to 3 times after failure (4 total request attempts)
    max_retries=3,
)

# Simulate a normal call
try:
    response = model.invoke("What is Python Tutorial Python?")
    print(f"Call successful: {response.content[:50]}...")
except Exception as e:
    print(f"Call failed: {e}")
ParameterDescriptionRecommended Value
timeoutMaximum waiting time for a single request (seconds). None means no limit30~60 (too short may cause timeouts; too long degrades user experience)
max_retriesNumber of retries after failure. 0 means no retry2~3 (sufficient for handling occasional network issues)

Time cost analysis for setting timeout and max_retries:

假设 timeout=30, max_retries=3:
单次成功: ~2s
第1次失败+重试成功: ~2s + ~2s = ~4s
全部失败: 4次 × 30s = 最多 120s 才抛出异常

base_url - Custom API Address

The base_url parameter is very useful when you need to access models through a proxy, relay service, or private deployment.

Example

from langchain.chat_models import init_chat_model

# Scenario 1: Accessing OpenAI through a proxy
model = init_chat_model(
    "deepseek:deepseek-v4-flash",
    base_url="https://your-proxy-domain.com/v1",  # Proxy address
)

# Scenario 2: Using a third-party service compatible with the OpenAI API
# Many domestic models provide OpenAI-compatible interfaces
model = init_chat_model(
    "deepseek:deepseek-v4-flash",           # Set provider to openai
    base_url="https://api.third-party.com/v1",  # But it actually points to a third party
    api_key="your-third-party-key",  # Third-party API Key
)

# Scenario 3: Connecting to local models (e.g., vLLM, Ollama)
model = init_chat_model(
    "openai:qwen2.5",               # Local model name
    base_url="http://localhost:8000/v1",  # Local service address
    api_key="not-needed",           # Local usually does not require a Key
)

base_url changes the API endpoint address, but the provider parameter determines the behavior mode. For example, provider="openai" will use OpenAI's message format, even if base_url points to a different service. Ensure the target service is compatible with the provider format you specify.


Other Common Parameters

top_p - Nucleus Sampling

Example

from langchain.chat_models import init_chat_model

# top_p is another way to control randomness
# The model only samples from words whose cumulative probability reaches top_p
# top_p=0.1 means only selecting from the top 10% highest-probability words
model = init_chat_model(
    "deepseek:deepseek-v4-flash",
    top_p=0.9,       # Only consider words in the top 90% of cumulative probability
)

response = model.invoke("Introduce Python Tutorial Python")
print(response.content[:100])

It is generally recommended to adjust only one of temperature or top_p, not both at the same time. If both are set, the model's behavior may be difficult to predict.

stop - Stop Sequence

Example

from langchain.chat_models import init_chat_model

model = init_chat_model("deepseek:deepseek-v4-flash")

# The stop parameter specifies stop sequences; the model immediately stops generating when it encounters these words
response = model.invoke(
    "List five programming learning websites, one per line",
    stop=["\n"]  # Stop when a newline is encountered, only return the first one
)
print(f"Limit stop=['\n']: {response.content}"\\n']: {response.content}")

Output:

限制 stop=['\n']: 1. Example

seed - Reproducibility (Supported by Some Models)

Example

from langchain.chat_models import init_chat_model

# Some models support the seed parameter to obtain deterministic output
# Same seed + same input = same output
model = init_chat_model("deepseek:deepseek-v4-flash", seed=42, temperature=0)
resp1 = model.invoke("Introduce Python Tutorial")
resp2 = model.invoke("Introduce Python Tutorial")
print(f"seed=42, results identical: {resp1.content == resp2.content}")

Parameter Quick Reference Table

ParameterTypeDefault ValueWhen to Use
temperaturefloatVaries by modelSet to 0~0.3 when the task requires stability, and 0.7~1.0 when creativity is needed
max_tokensintModel upper limitWhen output length needs to be controlled
timeoutint/floatNoneRecommended to always set in production environments
max_retriesintVaries by modelRecommended 2~3 when the network is unstable
base_urlstrOfficial addressWhen using a proxy, relay, or local service
top_pfloat1.0When nucleus sampling control is needed (as an alternative to temperature)
stoplist[str]NoneWhen precise control over the output ending is needed
Other Extensions