Hermes Agent sub-agent delegation and batch processing
This chapter introduces Hermes' two parallel processing capabilities: sub-agent delegation, which allows a main agent to distribute tasks to multiple sub-agents for parallel execution, and a batch processing system that can run agents in parallel on hundreds of prompts and generate structured training data.
Sub-agent delegation and parallelism
Hermes supports delegating tasks from the main Agent to sub-Agents, enabling parallel workflows.
Sub-agents viadelegationToolset implementation:
Example
# Ensure the delegation toolset is enabled
toolsets:
- delegation
Main agent usesspawn_agentorspawn_parallelTool-derived Agent.
The following pseudocode illustrates the mechanism of parallel delegation:
Example
# spawn_parallel dispatches multiple sub-agents simultaneously, each executing independently.
spawn_parallel([
{"prompt": "Analyze Q1 sales data and generate trend report"},
{"prompt": "Analyze Q2 sales data and generate trend report"},
{"prompt": "Analyze Q3 sales data and generate trend report"}
])
# Three sub-Agents work in parallel, with results aggregated and returned to the main Agent
Background task
Start background tasks directly in the session:
Example
/background check all servers and report downtime
Kanban Multi-Agent Collaboration
The Kanban toolset provides more structured multi-Agent collaboration — an orchestrator Agent distributes tasks to the board, and multiple worker Agents claim and execute them:
Example
hermes tools enable kanban
# Add role descriptions to Profile so the orchestrator knows each Agent's capabilities.
hermes profile create researcher \
--description "Skilled at reading source code and external documentation, producing research reports"
hermes profile create coder \
--description "Skilled at implementing features, fixing bugs, and writing tests"
The Kanban toolset must be explicitly enabled — even all or * will not automatically turn it on. This is intentional design, because multi-Agent collaboration consumes more tokens.
Batch Processing and RL Training Data Generation
Batch processing lets you run agents in parallel on hundreds or thousands of prompts, generating structured trajectory data—mainly used for model fine-tuning training data generation and evaluation.
Quick start
Example
cat data/prompts.jsonl
# {"prompt": "Write a Python function to find the longest palindromic substring"}
# {"prompt": "Create user authentication REST API endpoints with Flask"}
# {"prompt": "DebugThisError:TypeError: cannot unpack non-iterable"}
# Run Batch Processing
python batch_runner.py \
--dataset_file=data/prompts.jsonl \
--batch_size=10 \
--run_name=my_first_run \
--model=anthropic/claude-sonnet-4-6 \
--num_workers=4
# Resume interrupted task from checkpoint
python batch_runner.py \
--dataset_file=data/prompts.jsonl \
--batch_size=10 \
--run_name=my_first_run \
--resume
Dataset format
Each line is a JSON object, must containpromptField, optionally contains additional context:
Example
{"prompt": "DebugThisError", "cwd": "/home/user/project", "docker_image": "python:3.12"}
Main parameters
| Parameter | Default value | Description |
|---|---|---|
| --dataset_file | Required | JSONL dataset path |
| --batch_size | Required | Number of prompts per batch |
| --run_name | Required | Run name (used for output directory and checkpoint resumption) |
| --model | claude-sonnet-4.6 | Models used |
| --num_workers | 4 | Number of parallel worker processes |
| --max_turns | 10 | Maximum tool call rounds per prompt |
| --distribution | "default" | Toolset distribution (randomly sampling different tool combinations) |
| --reasoning_effort | — | Reasoning effort: none/minimal/low/medium/high/xhigh |
| --resume | false | Resume from checkpoint |
Output structure
Batch processing generates the following directory structure:
data/my_first_run/ ├── trajectories.jsonl # 最终合并输出(所有批次) ├── batch_0.jsonl # 各批次结果 ├── batch_1.jsonl ├── checkpoint.json # 断点续传记录 └── statistics.json # 工具使用统计
Each trajectory is output in ShareGPT format, including full conversation history, tool invocation records, and reasoning coverage statistics:
Example
"prompt_index": 42,
"conversations": [
{"from": "human", "value": "Write a function to find palindromes..."},
{"from": "gpt", "value": "I'll create this function...",}
"tool_calls": [{"name": "terminal", "arguments": {...}}]},
{"from": "tool", "value": "...execution result..."},
{"from": "gpt", "value": "This is the completed function..."}
],
"completed": true,
"toolsets_used": ["terminal", "file"],
"tool_stats": {
"terminal": {"count": 2, "success": 2, "failure": 0}
}
}
Quality filtering
After the batch run ends, two filters are automatically applied:
| filter item | Rules | Purpose |
|---|---|---|
| Zero inference filtering | Samples without any trace of reasoning were discarded. | Ensure the output contains the thinking process to improve training data quality |
| Hallucinated tool name filtering | Entries containing tool calls that are not in the valid tool list are filtered out | Exclude invalid tool calls produced by model hallucination |
Reasoning traces refer to <REASONING_SCRATCHPAD> markers or native thinking tokens. The filtered data is more suitable for model fine-tuning.
Integration with the Atropos RL environment
Batch processing can directly generate Atropos-compatible training data:
Example
python batch_runner.py \
--dataset_file=data/coding_tasks.jsonl \
--run_name=atropos_run \
--model=nous-hermes-3.1 \
--num_workers=8 \
--reasoning_effort=high
# The output trajectories.jsonl can be directly used as Atropos training input