Hermes Agent Toolkit

Tools are the only channel through which Hermes Agent interacts with the external world—reading and writing files, executing commands, searching the web, generating images; all capabilities are exposed through tools.

Hermes has 70+ built-in tools, organized into 28 tool collections. This chapter will analyze the core tools, parameters, and best practices of each tool collection one by one.


Tool system architecture

Before diving into specific tools, let's first understand how tools are organized.

Hermes uses a three-layer structure to manage tools:

工具(Tool)          ← 单个可调用函数,如 read_file、web_search
  └── 工具集(Toolset)   ← 相关工具的逻辑分组,如 file、web
        └── 平台配置        ← 哪个平台加载哪些工具集

Key Design: Disabling a tool collection makes its tools completely disappear from the system prompt—not just uncallable, but the Agent is entirely unaware of their existence. This both saves Tokens and prevents the Agent from misusing tools it shouldn't use.

Ways to manage tool sets:

# TUI 交互式管理(推荐)
hermes tools

# 命令行管理
hermes tools list                # 列出所有工具及启用状态
hermes tools enable browser      # 启用指定工具集
hermes tools disable audio       # 禁用指定工具集

Configure in config.yaml:

# 文件路径:~/.hermes/config.yaml
# 按平台指定启用的工具集
platform_toolsets:
  cli:
    - hermes-cli         # CLI 核心工具集
    - file
    - terminal
    - web
    - memory
    - skills

Desktop version, in the left-side Skills & Tools menu:

Tools can also be managed on the dashboard, in the left-side skills menu:


File and system tool sets

This type of tool collection enables the Agent to manipulate the local file system, execute commands, and manage processes.

file — file operations

The most fundamental tool collection; the Agent relies on it for nearly all tasks to read and write files.

ToolsFeaturesKey Parameters
read_fileRead file contentfile_path (required), offset (starting line), limit (maximum number of lines)
write_fileCreate or overwrite filefile_path (required), content (required)
edit_filePrecisely replace text in filesfile_path、old_string、new_string、replace_all
list_dirList directory contentspath (required), recursive (whether to recurse)
search_filesSearch by file name patternpattern (required, supports glob), path (search root directory)
search_contentSearch text in file contentpattern (required, supports regex), path, file_types

How edit_file works: It uses exact substring matching to locate the modification point. The Agent provides old_string (text that actually exists in the file), the tool finds it and replaces it with new_string. If old_string matches multiple locations, set replace_all: true to replace all occurrences.

edit_file is the Agent's most commonly used file modification method—it is safer than write_file because if old_string is not found, the operation fails with an error instead of accidentally overwriting the file.

terminal — Shell command execution

Allows the Agent to execute Shell commands in a specified backend environment.

ToolsFeaturesKey Parameters
run_commandExecute a single Shell commandcommand (required), workdir (working directory), timeout (timeout in milliseconds)
run_scriptExecute multi-line Shell scriptsscript (required, multi-line script content), workdir, timeout

Commands execute in the configured terminal backend—local, docker (container), ssh (remote), etc.

Execution results include stdout, stderr, and exit code. The Agent can use this information to determine whether the command succeeded.

code — sandbox code execution

Executes code snippets in an isolated sandbox, supporting Python, JavaScript, and Bash.

ToolsFeaturesKey Parameters
execute_codeRun code in sandboxlanguage (python/js/bash), code (required), timeout

Difference from terminal: the code tool runs in an independent sandbox and cannot access the file system (unless explicitly mounted). It is suitable for running code that does not require file interaction, such as algorithm validation and data processing.

computer — desktop GUI automation

Controls the mouse, keyboard, and screenshots of the desktop environment.

ToolsFeatures
screenshotCapture the screen or a specified region
clickClick the mouse at specified coordinates
typeSimulate keyboard text input
scrollScroll mouse wheel
moveMove the mouse to specified coordinates

The computer tool collection is mainly used for GUI application testing and automated operations. If your use case doesn't involve desktop applications, you can disable it to save Tokens.


Network and content tool sets

These tool sets enable the Agent to search the internet, extract web page content, and automate browser operations.

web — Web search and content extraction

The main channel through which the Agent obtains external information.

ToolsFeaturesKey Parameters
web_searchSearch internet contentquery (required), num_results (number of results)
web_extractExtract the main content of a specified URLurl (required), extract_mode (auto/text/markdown)

web_search returns a list of search results (title + URL + summary). The Agent typically searches first, then calls web_extract on pages of interest to get detailed content.

A search API must be configured to use it. Multiple providers are supported:

Example

# File path: ~/.hermes/config.yaml
# Configure web search backend
web
:
  search_provider
: firecrawl    # or brave, serpapi, tavily
  # Set the corresponding API Key in .env
  # FIRECRAWL_API_KEY=fc-...
  # Or BRAVE_API_KEY=BS...

browser — Browser automation

Full browser control capabilities—navigation, clicking, form filling, screenshots—operating web pages like a human.

ToolsFeaturesKey Parameters
navigateNavigate to a specified URLurl (required)
clickClick page elementselector (CSS selector or text description)
fill_formFill form fieldsfields (mapping from field names to values)
screenshotCapture current pagefull_page (whether to take a full-page screenshot)
extractExtract structured data from a pageselector、format(json/text)
execute_jsExecutes JavaScript in the pagecode (required)

The browser tool collection requires Playwright or a cloud browser backend.

The browser tool set consumes additional resources when launching browser instances. If you only need to extract web page text content, prefer web_extract over browser extract—the former is a lightweight HTTP request, while the latter requires launching a full browser.

vision — image analysis

Allows the Agent to "understand" image content.

ToolsFeaturesKey Parameters
analyze_imageAnalyze image content and generate a descriptionimage_path or image_url (required), question (optional, specific question about the image)
describe_imageGenerate a detailed textual description of an imageimage_path or image_url (required)

The vision tool sends images to a vision-capable LLM (such as claude-sonnet-4-6, gpt-4o), and the model recognizes and describes the image content.

audio — audio processing

Speech-to-text and text-to-speech.

ToolsFeaturesKey Parameters
transcribeConvert audio files to textaudio_path (required), language (optional)
ttsConvert text to speech filestext (required), voice (voice selection), provider

Agent capability tool sets

This type of tool collection manages the Agent's own state—memory, skills, and sub-Agent collaboration.

memory — memory management

Read and write persistent memory and search cross-session history.

ToolsFeaturesKey Parameters
memoryManage persistent memory (add, delete, modify)action(add/replace/remove)、target(memory/user)、content
session_searchFull-text search historical sessionsquery (required, supports FTS5 syntax)

The memory tool has no read operation—persistent memory is already injected into the system prompt at session startup, so the Agent doesn't need to read it separately.

skills — skill management

Loads and manages reusable workflow skills.

ToolsFeaturesKey Parameters
skills_listList all installed skills and their descriptionsNone (automatically called at session startup)
skill_viewLoad the full content of a skillname (required), path (optional, load references sub-files)
skill_manageCreate, modify, or delete skill filesaction、name、content

delegation — sub-agent delegation

Spawns sub-Agents to execute tasks in parallel.

ToolsFeaturesKey Parameters
spawn_agentCreates a single sub-Agent to execute a taskprompt (required), model (optional)
spawn_parallelCreate multiple sub-Agents in paralleltasks (required, task list)

Sub-Agents have independent context windows and do not pollute the main Agent's conversation history. After execution, they return a summary of results.

kanban — Multi-agent kanban board

A structured multi-Agent collaboration system that requires explicit enabling (all/* does not automatically turn it on).

ToolsFeatures
task_createCreate a new task in the kanban board
task_updateUpdate task status (todo/in_progress/done)
task_listList all tasks of the tenant
task_assignAssigns tasks to designated worker Agents

Kanban is the core mechanism for multi-Profile collaboration—one orchestrator Profile creates tasks, and multiple worker Profiles claim and execute them. This requires the gateways of multiple Profiles to run simultaneously.


Media creation toolset

This type of tool collection typically requires a Nous Portal subscription or the corresponding third-party API Key.

image_gen — AI image generation

Generates images from text descriptions, supporting 9 models.

ToolsFeaturesKey Parameters
generate_imageGenerate an image based on a Promptprompt(Required)、model(flux/gpt-image/ideogram etc.)、size、style

tts — text-to-speech

TTS functionality independent of the audio tool collection, offering a richer selection of providers.

ToolsFeatures
text_to_speechConvert text to speech files
list_voicesList available voice timbres

Custom toolset

You can combine existing tool collections into project-specific custom tool collections:

Example

# File path: ~/.hermes/config.yaml
# Define project-specific toolset combination
custom_toolsets
:
 # Tools needed for data science projects
  data-science
:
   - file
    - terminal
    - code
    - web

  # Pure writing projects only need file operations
  writing
:
   - file
    - web
    - memory

  # Ops projects need terminal + files, not a browser
  devops
:
   - file
    - terminal
    - memory

When using, through--toolsetsParameter specification:

Example

# Use custom toolset at startup
hermes --toolsets data-science chat

# Or set as default in Profile configuration
coder config set toolsets data-science

Tool call flow

Understanding the complete flow of a tool call helps to understand the Agent's behavior and troubleshoot issues.

A complete tool call goes through the following steps:

  1. The LLM decides to call a tool while generating its reply, outputting the tool name and parameters
  2. tool_request Middleware execution (can modify parameters)
  3. Approval check—if the tool is a dangerous operation, decides whether to show a confirmation prompt based on approvals.mode
  4. Actual tool execution
  5. tool_execution Middleware executes (can modify results)
  6. post_tool_call Hook records telemetry data
  7. The execution result is returned to the LLM, and the LLM continues generating a reply

If your Middleware modifies tool parameters, the modified parameters are the input for approval and actual execution. If you short-circuit execution in the Middleware (return a custom result), the actual tool will not run.


Tool usage statistics

Knowing which tools are frequently used helps optimize tool collection configuration and Token consumption.

In a session, view the tool call statistics for the current session:

Example

# View tool usage statistics for this session
/usage

# Output example:
# Token usage: Input 12,340 | Output 5,678 | Total 18,018
# Tool calls: read_file (15 times) | edit_file (8 times)
# terminal (3 times) | web_search (2 times)

In batch processing scenarios, the statistics.json file records more detailed statistics:

Example

// File path: data/my_first_run/statistics.json
{
  "total_tool_calls": 1523,
  "tools": {
    "read_file": {"count": 520, "success": 510, "failure": 10},
    "edit_file": {"count": 380, "success": 375, "failure": 5},
    "terminal": {"count": 290, "success": 270, "failure": 20},
    "web_search": {"count": 180, "success": 178, "failure": 2}
  },
  "avg_tool_calls_per_turn": 2.4,
  "most_common_tool_sequence": [
    "read_file", "edit_file", "terminal"
  ]
}
other extensions