Hermes Agent Memory System

Hermes's core differentiating capability is not tool calling, but rather itsReally remembers you。

Most AI tools start from zero with every conversation—you spend an hour today teaching it your project structure, naming conventions, and deployment process, and tomorrow when you start a new session, it all resets.

This chapter provides an in-depth analysis of how Hermes solves this problem through a three-tier memory structure, and how the Agent autonomously manages memory, enabling cross-session search and recall.


Three-tier memory architecture

Hermes uses a three-layer memory architecture of mutually independent, functionally complementary layers to store and recall information.

Core design philosophy of the three-tier architecture:High-frequency important information remains in context, low-frequency historical information is recalled on demand, and deep patterns are automatically inferred by AI.—striking a balance between memory capability and token cost.

┌─────────────────────────────────────────────────┐
│       第一层:持久记忆(Persistent Memory)        │
│  MEMORY.md(环境事实)+ USER.md(用户画像)          │
│  → 每次会话启动时自动注入系统提示词,始终可见          │
├─────────────────────────────────────────────────┤
│       第二层:情节日志(Session History)            │
│  SQLite + FTS5 全文检索,存储所有历史对话             │
│  → Agent 主动搜索时按需加载,不占用固定 Token         │
├─────────────────────────────────────────────────┤
│       第三层:用户建模(Honcho,可选)               │
│  辩证推理引擎,提取跨会话模式、偏好与目标              │
│  → 比你亲口说出来的更深层地理解你                     │
└─────────────────────────────────────────────────┘
Memory HierarchyStorage locationCapacityLoading MethodPurpose
Persistent MemoryMarkdown files under ~/.hermes/memories/~1,300 tokens totalAutomatically injected into the system prompt at session startupKey facts, always visible
Story LogSQLite database + JSONL transcript filesUnlimited (all historical conversations)Agent searches on demand, FTS5 full-text searchFind specific historical discussions
Honcho user modelingHoncho server-side APIIsolated by agentBackground inference injection per round / every N roundsAutomatically infers deep preferences and goals

The three layers are not substitutes but complements. Persistent memory answers "I always need to know these," session history answers "Did we discuss this last week," and Honcho answers "What do you really want."


Persistent memory file

Persistent memory is the first and most fundamental layer of the Hermes memory system.

It is stored in two~/.hermes/memories/Composed of Markdown files under, automatically injected into the system prompt at each session start.

Two core files

The two files each have their own role and do not overlap.

FilePurposeCharacter Limitapproximately equal to
MEMORY.mdThe Agent's personal notes—environment facts, project conventions, things learned on the job2,200 characters~800 tokens
USER.mdUser profile—your preferences, communication style, work habits, skill level1,375 characters~500 tokens

The character limit is intentional—it ensures memory content stays concise and focused. If a new entry exceeds the limit, the memory tool returns an error prompting the Agent to free up space before writing, rather than silently discarding old entries.

How memory is injected into the system prompt

At each session start, the contents of both files are injected into the system prompt in a fixed format, as follows:

══════════════════════════════════════════════
MEMORY(Agent 个人笔记)[67% — 1,474/2,200 字符]
══════════════════════════════════════════════
用户的项目是位于 ~/code/myapi 的 Rust Web 服务,使用 Axum + SQLx
§
本机运行 Ubuntu 22.04,已安装 Docker 和 Podman
§
用户偏好简洁回复,不喜欢冗长解释

Here are a few key design details:

Freeze Snapshot Mode: Memory content in the system prompt is captured once at the start of the session and is not updated mid-session.

This is a deliberate design—to keep the LLM's prefix cache (Prompt Cache) valid and reduce token cost.

When the Agent modifies memory during a session, changes are written to disk immediately, but must wait forLaunch on next sessionOnly then will it appear in the context.

Usage percentage visible:[67% — 1,474/2,200]Keeps the Agent aware of remaining capacity at all times, so it can proactively reorganize when approaching the limit.

§ as item separator: each memory entry is separated by §, and entries themselves can be multi-line text.

How the Agent manages memory

The Agent uses built-inmemorytools to manage memory, supporting three operations.

OperationDescriptionParameter
addAdd a new memory entrytarget(memory/user)、content
replaceReplace an existing entry (located by substring matching)target、old_text、content
removeDelete an existing entry (located by substring matching)target、old_text

Note:No read operation—memory content is already injected into the system prompt at session start, so the Agent sees it directly in context without needing additional reads.

replaceandremoveUsageSubstring MatchingTo locate an entry, you don't need to provide the full entry text:

Example

# Scenario: Current memory has "User prefers dark theme for all editors"
# Now needs to be updated to a more precise version
# Use substring matching to locate — just provide a uniquely identifiable snippet

memory(
    action="replace",
    target="memory",
    old_text="Dark Theme",  # A unique substring is enough to locate the original entry.
    content="User VS Code uses light theme, terminal uses dark theme"  # New content after replacement
)

If the substring matches multiple entries, the tool returns an error requesting a more precise match string.


What to remember, what not to remember

The Agent automatically determines and saves valuable information—you usually don't need to ask explicitly.

But understanding the criteria helps you understand the Agent's behavior and manually guide it when needed.

Should be proactively saved to memory (environment & work)

The following types of information trigger the Agent to write to MEMORY.md:

TypeExample
Environmental FactsThis machine runs Debian 12, PostgreSQL 16, with Docker replaced by Podman
Project Specification~/code/api Usage Go 1.22,sqlc query,chi Routing;Testinguse make test
Tool ExperienceStaging server SSH port is 2222, not 22; key is at ~/.ssh/staging_ed25519
Correction RecordDon't run Docker commands with sudo; the user is already in the docker group
Complete work2026-01-15 Completed migration of database from MySQL to PostgreSQL

Should be proactively saved to user (about you)

The following types of information trigger the Agent to write to USER.md:

TypeExample
Identity informationName, job title, timezone
Communication preferencesPrefers concise replies, no need to explain obvious steps
disliked itemDon't add excessive comments in code
Work habitsDon't check messages before 3 PM every day
Skill levelFamiliar with Rust and Python, Go is newly learned

Content that should not be saved

The following content does not deserve precious memory space:

TypeExampleReason
Overly broadUser has a projectNo informational value, cannot be reused
Can re-check anytimePython 3.12 supports nested f-stringsJust search; not worth occupying space
Raw dataLarge log segments, code blocks, data tablesExceeds character limit; structured memory is more useful
Temporary contentTemporary file paths for this debugging sessionNot usable next time
Content already in the contextInformation already present in project SOUL.md, AGENTS.mdRepeated injection wastes tokens

Good memory vs. bad memory

The same information, well-written, can be worth ten entries; poorly written, it's worthless.

RatingExampleReview
OKuserof macOS 14 Sonoma,Homebrew,Docker Desktop,Shell Yes带 oh-my-zsh of zsh,EditorYes VS Code + Vim KeybitsHigh information density; related facts are bundled together
OKProject ~/code/api Usage Go 1.22,sqlc 做 DB query,chi route.Test:make test。CI:GitHub ActionsSpecific actionable rules
OKstaging Server(10.0.1.50)SSH port 2222 is not 22,dense钥 ~/.ssh/staging_ed25519Contextual lessons learned, avoiding pitfalls
PoorUsers have projectsToo vague, cannot guide any action
PoorOn January 5, 2026, the user asked me to look at their project at ~/code/api, and I found it uses Go 1.22...Too verbose; the core information is buried in the narrative

Memory capacity management

Memory files have a strict character limit; the Agent needs to proactively organize when approaching the limit.

What happens when approaching the limit

When a new entry would cause the character limit to be exceeded, the tool returns an error rather than silently truncating:

{
  "success": false,
  "error": "Memory 已用 2,100/2,200 字符。新条目(250 字符)会超出上限。
            请先整理:用 replace 合并重叠条目,或 remove 过时条目,
            再重试——在同一轮对话中完成。",
  "current_entries": [
    "用户的 macOS 14 Sonoma,Homebrew,Docker Desktop",
    "项目 ~/code/api 使用 Go 1.22,sqlc 做 DB 查询"
  ],
  "usage": "2,100/2,200"
}

Upon receiving this error, the Agent's correct approach is:

  1. Read all current entries (already included in the error response)
  2. Find entries that can be deleted or merged
  3. usereplaceMerge several related entries into a more refined version
  4. Then retry adding the new entry

Best practice: when the system prompt shows memory usage exceeding 80%, organize proactively—don't wait for an error. Prevention always saves more tokens than remediation.

Capacity reference

StorageUpper limitTypical entry count
MEMORY.md2,200 characters8–15 items
USER.md1,375 characters5–10 items

Controlling memory write permissions

By default, the Agent can write to memory freely—includingBackground auto-triggered self-improvement review(Runs in the background after a conversation round ends, distilling memory and skills.)

If you want to review content before anything is written to memory, you can enable write approval.

Enable write approval

Just set it in the configuration file:

Example

# File path: ~/.hermes/config.yaml
# Enable memory write approval — all write operations require manual approval
memory
:
  write_approval
: true   # Default false (free writing)

Behavior changes after enabling:

ScenariosBehavior
When the Agent wants to write memory during CLI interactionInline prompt in the terminal asking for your approval
Messaging platform / background automatic reviewWrites are staged, waiting for your manual approval

Approval management commands

Use the following slash commands to manage pending memory writes:

Example

# View pending memory writes (auto-triggered tag [auto])
/memory pending

# Approve writes for the specified ID, or approve all.
/memory approve <id>
/memory approve all

# Reject writes for the specified ID, or reject all.
/memory reject <id>
/memory reject all

# Enable or disable approval at runtime (setting persists)
/memory approval on
/memory approval off

Use cases: if the Agent saved an incorrect assumption about you, or you want full control over memory content, enabling write_approval is the most direct approach.

About the background self-improvement review

After each conversation round, Hermes runs a background self-improvement review (background review) in the background, analyzing the conversation and distilling memories or skills worth keeping.

By default, a line of notification is displayed in the chatMemory updated。

You can adjust the level of notification detail:

Example

# File path: ~/.hermes/config.yaml
# Controls the notification detail level of background review
display
:
  memory_notifications
: on      # off | on (default) | verbose
ValueEffect
offNo notification shown, but the review still runs and still writes
on (default)Short prompts, such as "Memory updated"
verboseIncludes a change preview, such as Memory + User prefers concise replies

If your main model is expensive, you can run background reviews on a cheaper model:

Example

# File path: ~/.hermes/config.yaml
# Background Review Uses an Independent Model to Reduce Costs
auxiliary
:
  background_review
:
    provider
: openrouter
    model
: google/gemini-3-flash-preview   # Default auto = main model

Episode directory and cross-session search

The second memory layer—Session History—stores complete records of all historical conversations, supporting full-text search and on-demand recall.

All conversations are automatically saved

Every session—whether from CLI, Telegram, Discord, or any other platform—is automatically saved to two locations:

Storage locationFormattingcontent
~/.hermes/state.dbSQLite database + FTS5 full-text indexStructured metadata, complete message history (including tool calls and results), token usage, timestamps
~/.hermes/sessions/JSONL transcript filesRaw conversation transcripts, including tool invocation records

session_search tool

Built-in Agentsession_searchA tool that can perform full-text search across all historical conversations.

The entire process requires no LLM involvement in the search phase; it's done directly by the SQLite FTS5 engine:

Workflow:

  1. FTS5 retrieves matching messages sorted by relevance
  2. Grouped by session, taking the most relevant sessions (3 by default)
  3. Loads each session's conversation, capturing approximately 100,000 characters around the match point
  4. A lightweight summarization model generates focused summaries
  5. Returns each session's summary and context to the Agent

Supported search syntax (FTS5 standard):

SyntaxExampleDescription
Keyword retrievaldocker deploymentMatches messages containing any keyword
Phrase matching"exact phrase"Exact match for the entire phrase
Boolean ORdocker OR kubernetesMatch any keyword
Boolean NOTpython NOT javaExclude specific keywords
Prefix wildcarddeploy*Matches words starting with "deploy"

Trigger timing: the Agent automatically calls session_search when you mention "what we discussed last time..." or "that previous proposal...", instead of making you repeat the explanation.

session_search vs persistent memory

The two are complementary, not a replacement relationship:

Comparison DimensionPersistent memory (MEMORY.md / USER.md)Event log (Session Search)
Capacity~1,300 tokens totalUnlimited (all historical conversations)
SpeedInstant (already in the system prompt)~20ms (FTS5 query)
Token costFixed consumption per sessionOn demand; no query, no cost
PurposeKey facts, always visibleFind specific historical discussions
Management styleAgent proactively organizesFully automated, no intervention needed

MemoryUsed for "I always need to know these"—environment facts, user preferences, project conventions.

Story LogUsed for "Did we talk about this last week?"—finding specific historical discussions.


Session Management

Each session in the session history can be named, resumed, exported, and cleaned up.

Naming and recovery

Hermes automatically generates a short session title (3–7 words) after the first conversation round, running asynchronously in the background without affecting response speed.

You can also name it manually:

Example

# Manually set the title in the session
/title refactor authentication module

After naming, sessions can be resumed directly by title:

Example

# Restore session by title
hermes -c Refactor authentication module

# Restore the most recent session
hermes --continue

# Restore by session ID
hermes -r <session_id>

Auto-Lineage: when a session undergoes context compaction, Hermes automatically creates and numbers continuation sessions:

"我的项目" → "我的项目 #2" → "我的项目 #3"

When resuming by title, the latest continuation session is automatically selected.

Common session management commands

Example

# List the most recent 20 sessions
hermes sessions list

# Filter by source platform
hermes sessions list --source telegram

# Show more entries
hermes sessions list --limit 50

# Rename specified session
hermes sessions rename <id> "New title"

# Export all sessions to file
hermes sessions export backup.jsonl

# Delete specified session
hermes sessions delete <id>

# Clean up sessions older than 30 days
hermes sessions prune --older-than 30

# View statistics
hermes sessions stats

hermes sessions listOutput example:

标题                 预览                               最近活跃    ID
────────────────────────────────────────────────────────────────────
重构认证模块          帮我重构认证模块的代码                2 小时前   20260305_09...
我的项目 #3          检查一下测试失败的原因                昨天        20260304_14...
—                   Python 装饰器怎么用                  3 天前      20260303_10...

Automatic session cleanup (optional)

By default, session history is retained permanently (because it provides data for session_search).

If automatic cleanup is needed, enable it in the configuration:

Example

# File path: ~/.hermes/config.yaml
# Enable automatic conversation cleanup (default off)
sessions
:
  auto_prune
: true        # default false
  retention_days
: 90      # Keep the last 90 days
  vacuum_after_prune
: true
  min_interval_hours
: 24

Honcho user modeling

Honcho is an AI-native memory backend developed by Plastic Labs, integrated into Hermes as a plugin.

It doesn't replace MEMORY.md and USER.md, but rather builds on top of themAdd dialectical reasoning layer—by analyzing your conversation patterns, it automatically infers preferences, goals, and habits you haven't explicitly expressed.

Unlike the manual curation of MEMORY.md, Honcho is fully automatic: after each conversation turn (frequency configurable), it reasons in the background, accumulating a deeper understanding of you.

Honcho vs. built-in memory

SkillsBuilt-in memoryHoncho
Cross-session persistenceLocal filesServer-side API
User profileManually curated (USER.md)Automatic dialectical reasoning
Session summary—Session-level context injection
Multi-agent isolation—Modeled separately per Agent
Semantic searchFTS5 full textBased on reasoning conclusions

Two-layer context injection

Each conversation turn, Honcho assembles two layers of context injected into the system prompt:

Base Layer (Base Context): session summary + user characterization + user Peer Card + AI self-characterization. Refreshed according to contextCadence (every N turns). This is the "who is this user" layer.

Dialectic layer (Dialectic Supplement): Generated in real time by Honcho's LLM engine, focusing on "what is most relevant right now." Refreshed according to dialecticCadence.

Two startup modes for dialectical reasoning:

ModeConditionsQuery method
cold startNo historical dataGeneric query—"Who is this user? What are their preferences, goals, and working style?"
hot startExisting historical dataSession-focused query — "Based on what has been discussed in this session, what is the most relevant context about this user?"

Three independent adjustment knobs

ParameterControlled contentDefault value
contextCadenceBase layer refresh interval (calls API once every N turns)1
dialecticCadenceDialectical layer refresh interval (calls LLM once every N turns)2 (Recommended 1–5)
dialecticDepthReasoning depth per dialectic (1–3 passes)1

The three knobs are completely independent — you can refresh the base layer at high frequency and run dialectical reasoning at low frequency, or vice versa:

Example

{
  "contextCadence": 1,
  "dialecticCadence": 5,
  "dialecticDepth": 2
}

The above configuration means: refresh the base layer every turn, and perform a 2-pass deep dialectical reasoning every 5 turns.

Multi-pass dialectical reasoning

When dialecticDepth is set to greater than 1, each dialectical reasoning run executes multiple rounds of LLM calls:

PassTaskDescription
Pass 0Cold/warm start reasoningEstablish initial assessment
Pass 1Self-auditIdentifies blind spots in the initial assessment, synthesizing evidence from recent sessions
Pass 2ReconciliationChecks contradictions from the previous two turns, producing a final synthesized conclusion

Depth 3 doesn't necessarily mean 3 LLM calls—if the previous turn already produced high-quality output, the next turn exits early to save tokens.

Quick start with Honcho

Example

# Interactive setup, choose honcho provider
hermes memory setup

Or manual configuration:

Example

# File path: ~/.hermes/config.yaml
# Enable Honcho as the memory provider.
memory
:
  provider
: honcho

Example

# Write API Key to environment variable file
echo 'HONCHO_API_KEY=your_key' >> ~/.hermes/.env

Get API Key at honcho.dev.

Tools provided by Honcho

After enabling Honcho, the Agent can use five exclusive tools:

ToolsPurpose
honcho_profileRead or update the user's Peer Card
honcho_searchSemantic search based on reasoning conclusions (raw fragments, no LLM synthesis)
honcho_contextRetrieve complete session context (summary, user characterization, recent messages)
honcho_reasoningCall Honcho LLM inference (reasoning_level can be specified)
honcho_concludeCreate or delete reasoning conclusions

Honcho CLI commands

Example

# View Honcho connection status, configuration and key settings
hermes honcho status

# View or set session policy
hermes honcho strategy

# View or Set Recall Mode (hybrid/context/tools)
hermes honcho mode

# View or set Token budget
hermes honcho tokens

# List known Honcho session mappings
hermes honcho sessions

Memory system configuration reference

The following is the complete reference for all memory-related configuration, which can be copied directly into your configuration file for use.

Example

# ============================================
# File path: ~/.hermes/config.yaml
# Hermes memory system complete configuration reference
# ============================================

# --- Persistent memory configuration ---
memory
:
  memory_enabled
: true          # Whether to enable persistent memory (MEMORY.md)
  user_profile_enabled
: true    # Whether to enable user profile (USER.md)
  memory_char_limit
: 2200       # MEMORY.md character limit (approximately 800 tokens)
  user_char_limit
: 1375         # USER.md character limit (approx. 500 tokens)
  write_approval
: false         # true = all memory writes require manual approval
  provider
: honcho              # Optional: Enable Honcho memory provider

# --- Session auto-cleanup (optional, disabled by default) ---
sessions
:
  auto_prune
: false             # Whether to automatically clean up old sessions
  retention_days
: 90            # Keep sessions from the last N days
  vacuum_after_prune
: true      # Whether to reclaim disk space after cleanup
  min_interval_hours
: 24        # Minimum interval between two cleanups

# --- Background review notifications ---
display
:
  memory_notifications
: on      # off | on | verbose

# --- Background review using cheaper model (optional) ---
auxiliary
:
  background_review
:
    provider
: openrouter        # Model Provider
    model
: google/gemini-3-flash-preview  # Model Name

Quick reference for common memory operations

Directly manipulate memory in conversation

You can instruct Hermes to operate on memory using natural language, without needing to memorize any commands:

Example

# Add Memory
Remember: this project uses Poetry for dependency management, not pip

# Delete or overwrite memory
Forget the previous records about MySQL, we have migrated to PostgreSQL

# Update information in memory
Update the staging server's IP to 10.0.2.100

# Query Memory
Tell me what you currently remember about me.

Slash Commands

Example

# View current memory contents
/memory

# View pending writes
/memory pending

# Approve all pending
/memory approve all

# Reject all pending
/memory reject all

# Enable Write Approval
/memory approval on

# Disable Write Approval
/memory approval off

Directly edit memory files

Memory files are plain Markdown, which can be opened and edited directly with an editor:

Example

# View current memory contents
cat ~/.hermes/memories/MEMORY.md
cat ~/.hermes/memories/USER.md

# Open in editor to modify
nano ~/.hermes/memories/MEMORY.md

Note: after directly editing files, changes take effect at the next session start and won't immediately appear in the current session's context.


Memory security and privacy

Security scan

Memory content undergoes a security scan before being written, detecting injection attacks and data theft patterns.

The system detects the following threat patterns:

Threat typeDetection content
Prompt injectionInjection patterns attempting to override the Agent's system instructions
Credential theftInstructions inducing the Agent to leak API keys or tokens
Backdoor installationInstallation instructions for SSH backdoors or remote access
Hidden charactersInvisible Unicode characters (zero-width spaces, directional overrides, etc.)

Content matching threat patterns will beDeny write。

Sensitive information in logs is masked

Sensitive information such as API Keys and Tokens is automatically redacted in all log files—even if it appears in tool outputs, it will never be recorded in plaintext:

Example

# File path: ~/.hermes/config.yaml
# Enabled by default, automatically desensitizes sensitive information
security
:
  redact_secrets
: true

Local data storage

All memory data (MEMORY.md, USER.md, state.db) is stored locally by default~/.hermes/Directory, not uploaded to the cloud

When choosing to use Honcho, conversation data is sent to Honcho's servers for processing—this is a deliberate trade-off: exchanging data uploads for deeper reasoning capabilities.

other extensions