Hermes Agent Memory System
Hermes's core differentiating capability is not tool calling, but rather itsReally remembers you。
Most AI tools start from zero with every conversation—you spend an hour today teaching it your project structure, naming conventions, and deployment process, and tomorrow when you start a new session, it all resets.
This chapter provides an in-depth analysis of how Hermes solves this problem through a three-tier memory structure, and how the Agent autonomously manages memory, enabling cross-session search and recall.
Three-tier memory architecture
Hermes uses a three-layer memory architecture of mutually independent, functionally complementary layers to store and recall information.
Core design philosophy of the three-tier architecture:High-frequency important information remains in context, low-frequency historical information is recalled on demand, and deep patterns are automatically inferred by AI.—striking a balance between memory capability and token cost.
┌─────────────────────────────────────────────────┐ │ 第一层:持久记忆(Persistent Memory) │ │ MEMORY.md(环境事实)+ USER.md(用户画像) │ │ → 每次会话启动时自动注入系统提示词,始终可见 │ ├─────────────────────────────────────────────────┤ │ 第二层:情节日志(Session History) │ │ SQLite + FTS5 全文检索,存储所有历史对话 │ │ → Agent 主动搜索时按需加载,不占用固定 Token │ ├─────────────────────────────────────────────────┤ │ 第三层:用户建模(Honcho,可选) │ │ 辩证推理引擎,提取跨会话模式、偏好与目标 │ │ → 比你亲口说出来的更深层地理解你 │ └─────────────────────────────────────────────────┘
| Memory Hierarchy | Storage location | Capacity | Loading Method | Purpose |
|---|---|---|---|---|
| Persistent Memory | Markdown files under ~/.hermes/memories/ | ~1,300 tokens total | Automatically injected into the system prompt at session startup | Key facts, always visible |
| Story Log | SQLite database + JSONL transcript files | Unlimited (all historical conversations) | Agent searches on demand, FTS5 full-text search | Find specific historical discussions |
| Honcho user modeling | Honcho server-side API | Isolated by agent | Background inference injection per round / every N rounds | Automatically infers deep preferences and goals |
The three layers are not substitutes but complements. Persistent memory answers "I always need to know these," session history answers "Did we discuss this last week," and Honcho answers "What do you really want."
Persistent memory file
Persistent memory is the first and most fundamental layer of the Hermes memory system.
It is stored in two~/.hermes/memories/Composed of Markdown files under, automatically injected into the system prompt at each session start.
Two core files
The two files each have their own role and do not overlap.
| File | Purpose | Character Limit | approximately equal to |
|---|---|---|---|
| MEMORY.md | The Agent's personal notes—environment facts, project conventions, things learned on the job | 2,200 characters | ~800 tokens |
| USER.md | User profile—your preferences, communication style, work habits, skill level | 1,375 characters | ~500 tokens |
The character limit is intentional—it ensures memory content stays concise and focused. If a new entry exceeds the limit, the memory tool returns an error prompting the Agent to free up space before writing, rather than silently discarding old entries.
How memory is injected into the system prompt
At each session start, the contents of both files are injected into the system prompt in a fixed format, as follows:
══════════════════════════════════════════════ MEMORY(Agent 个人笔记)[67% — 1,474/2,200 字符] ══════════════════════════════════════════════ 用户的项目是位于 ~/code/myapi 的 Rust Web 服务,使用 Axum + SQLx § 本机运行 Ubuntu 22.04,已安装 Docker 和 Podman § 用户偏好简洁回复,不喜欢冗长解释
Here are a few key design details:
Freeze Snapshot Mode: Memory content in the system prompt is captured once at the start of the session and is not updated mid-session.
This is a deliberate design—to keep the LLM's prefix cache (Prompt Cache) valid and reduce token cost.
When the Agent modifies memory during a session, changes are written to disk immediately, but must wait forLaunch on next sessionOnly then will it appear in the context.
Usage percentage visible:[67% — 1,474/2,200]Keeps the Agent aware of remaining capacity at all times, so it can proactively reorganize when approaching the limit.
§ as item separator: each memory entry is separated by §, and entries themselves can be multi-line text.
How the Agent manages memory
The Agent uses built-inmemorytools to manage memory, supporting three operations.
| Operation | Description | Parameter |
|---|---|---|
| add | Add a new memory entry | target(memory/user)、content |
| replace | Replace an existing entry (located by substring matching) | target、old_text、content |
| remove | Delete an existing entry (located by substring matching) | target、old_text |
Note:No read operation—memory content is already injected into the system prompt at session start, so the Agent sees it directly in context without needing additional reads.
replaceandremoveUsageSubstring MatchingTo locate an entry, you don't need to provide the full entry text:
Example
# Now needs to be updated to a more precise version
# Use substring matching to locate — just provide a uniquely identifiable snippet
memory(
action="replace",
target="memory",
old_text="Dark Theme", # A unique substring is enough to locate the original entry.
content="User VS Code uses light theme, terminal uses dark theme" # New content after replacement
)
If the substring matches multiple entries, the tool returns an error requesting a more precise match string.
What to remember, what not to remember
The Agent automatically determines and saves valuable information—you usually don't need to ask explicitly.
But understanding the criteria helps you understand the Agent's behavior and manually guide it when needed.
Should be proactively saved to memory (environment & work)
The following types of information trigger the Agent to write to MEMORY.md:
| Type | Example |
|---|---|
| Environmental Facts | This machine runs Debian 12, PostgreSQL 16, with Docker replaced by Podman |
| Project Specification | ~/code/api Usage Go 1.22,sqlc query,chi Routing;Testinguse make test |
| Tool Experience | Staging server SSH port is 2222, not 22; key is at ~/.ssh/staging_ed25519 |
| Correction Record | Don't run Docker commands with sudo; the user is already in the docker group |
| Complete work | 2026-01-15 Completed migration of database from MySQL to PostgreSQL |
Should be proactively saved to user (about you)
The following types of information trigger the Agent to write to USER.md:
| Type | Example |
|---|---|
| Identity information | Name, job title, timezone |
| Communication preferences | Prefers concise replies, no need to explain obvious steps |
| disliked item | Don't add excessive comments in code |
| Work habits | Don't check messages before 3 PM every day |
| Skill level | Familiar with Rust and Python, Go is newly learned |
Content that should not be saved
The following content does not deserve precious memory space:
| Type | Example | Reason |
|---|---|---|
| Overly broad | User has a project | No informational value, cannot be reused |
| Can re-check anytime | Python 3.12 supports nested f-strings | Just search; not worth occupying space |
| Raw data | Large log segments, code blocks, data tables | Exceeds character limit; structured memory is more useful |
| Temporary content | Temporary file paths for this debugging session | Not usable next time |
| Content already in the context | Information already present in project SOUL.md, AGENTS.md | Repeated injection wastes tokens |
Good memory vs. bad memory
The same information, well-written, can be worth ten entries; poorly written, it's worthless.
| Rating | Example | Review |
|---|---|---|
| OK | userof macOS 14 Sonoma,Homebrew,Docker Desktop,Shell Yes带 oh-my-zsh of zsh,EditorYes VS Code + Vim Keybits | High information density; related facts are bundled together |
| OK | Project ~/code/api Usage Go 1.22,sqlc 做 DB query,chi route.Test:make test。CI:GitHub Actions | Specific actionable rules |
| OK | staging Server(10.0.1.50)SSH port 2222 is not 22,dense钥 ~/.ssh/staging_ed25519 | Contextual lessons learned, avoiding pitfalls |
| Poor | Users have projects | Too vague, cannot guide any action |
| Poor | On January 5, 2026, the user asked me to look at their project at ~/code/api, and I found it uses Go 1.22... | Too verbose; the core information is buried in the narrative |
Memory capacity management
Memory files have a strict character limit; the Agent needs to proactively organize when approaching the limit.
What happens when approaching the limit
When a new entry would cause the character limit to be exceeded, the tool returns an error rather than silently truncating:
{
"success": false,
"error": "Memory 已用 2,100/2,200 字符。新条目(250 字符)会超出上限。
请先整理:用 replace 合并重叠条目,或 remove 过时条目,
再重试——在同一轮对话中完成。",
"current_entries": [
"用户的 macOS 14 Sonoma,Homebrew,Docker Desktop",
"项目 ~/code/api 使用 Go 1.22,sqlc 做 DB 查询"
],
"usage": "2,100/2,200"
}
Upon receiving this error, the Agent's correct approach is:
- Read all current entries (already included in the error response)
- Find entries that can be deleted or merged
- use
replaceMerge several related entries into a more refined version - Then retry adding the new entry
Best practice: when the system prompt shows memory usage exceeding 80%, organize proactively—don't wait for an error. Prevention always saves more tokens than remediation.
Capacity reference
| Storage | Upper limit | Typical entry count |
|---|---|---|
| MEMORY.md | 2,200 characters | 8–15 items |
| USER.md | 1,375 characters | 5–10 items |
Controlling memory write permissions
By default, the Agent can write to memory freely—includingBackground auto-triggered self-improvement review(Runs in the background after a conversation round ends, distilling memory and skills.)
If you want to review content before anything is written to memory, you can enable write approval.
Enable write approval
Just set it in the configuration file:
Example
# Enable memory write approval — all write operations require manual approval
memory:
write_approval: true # Default false (free writing)
Behavior changes after enabling:
| Scenarios | Behavior |
|---|---|
| When the Agent wants to write memory during CLI interaction | Inline prompt in the terminal asking for your approval |
| Messaging platform / background automatic review | Writes are staged, waiting for your manual approval |
Approval management commands
Use the following slash commands to manage pending memory writes:
Example
/memory pending
# Approve writes for the specified ID, or approve all.
/memory approve <id>
/memory approve all
# Reject writes for the specified ID, or reject all.
/memory reject <id>
/memory reject all
# Enable or disable approval at runtime (setting persists)
/memory approval on
/memory approval off
Use cases: if the Agent saved an incorrect assumption about you, or you want full control over memory content, enabling write_approval is the most direct approach.
About the background self-improvement review
After each conversation round, Hermes runs a background self-improvement review (background review) in the background, analyzing the conversation and distilling memories or skills worth keeping.
By default, a line of notification is displayed in the chatMemory updated。
You can adjust the level of notification detail:
Example
# Controls the notification detail level of background review
display:
memory_notifications: on # off | on (default) | verbose
| Value | Effect |
|---|---|
| off | No notification shown, but the review still runs and still writes |
| on (default) | Short prompts, such as "Memory updated" |
| verbose | Includes a change preview, such as Memory + User prefers concise replies |
If your main model is expensive, you can run background reviews on a cheaper model:
Example
# Background Review Uses an Independent Model to Reduce Costs
auxiliary:
background_review:
provider: openrouter
model: google/gemini-3-flash-preview # Default auto = main model
Episode directory and cross-session search
The second memory layer—Session History—stores complete records of all historical conversations, supporting full-text search and on-demand recall.
All conversations are automatically saved
Every session—whether from CLI, Telegram, Discord, or any other platform—is automatically saved to two locations:
| Storage location | Formatting | content |
|---|---|---|
| ~/.hermes/state.db | SQLite database + FTS5 full-text index | Structured metadata, complete message history (including tool calls and results), token usage, timestamps |
| ~/.hermes/sessions/ | JSONL transcript files | Raw conversation transcripts, including tool invocation records |
session_search tool
Built-in Agentsession_searchA tool that can perform full-text search across all historical conversations.
The entire process requires no LLM involvement in the search phase; it's done directly by the SQLite FTS5 engine:
Workflow:
- FTS5 retrieves matching messages sorted by relevance
- Grouped by session, taking the most relevant sessions (3 by default)
- Loads each session's conversation, capturing approximately 100,000 characters around the match point
- A lightweight summarization model generates focused summaries
- Returns each session's summary and context to the Agent
Supported search syntax (FTS5 standard):
| Syntax | Example | Description |
|---|---|---|
| Keyword retrieval | docker deployment | Matches messages containing any keyword |
| Phrase matching | "exact phrase" | Exact match for the entire phrase |
| Boolean OR | docker OR kubernetes | Match any keyword |
| Boolean NOT | python NOT java | Exclude specific keywords |
| Prefix wildcard | deploy* | Matches words starting with "deploy" |
Trigger timing: the Agent automatically calls session_search when you mention "what we discussed last time..." or "that previous proposal...", instead of making you repeat the explanation.
session_search vs persistent memory
The two are complementary, not a replacement relationship:
| Comparison Dimension | Persistent memory (MEMORY.md / USER.md) | Event log (Session Search) |
|---|---|---|
| Capacity | ~1,300 tokens total | Unlimited (all historical conversations) |
| Speed | Instant (already in the system prompt) | ~20ms (FTS5 query) |
| Token cost | Fixed consumption per session | On demand; no query, no cost |
| Purpose | Key facts, always visible | Find specific historical discussions |
| Management style | Agent proactively organizes | Fully automated, no intervention needed |
MemoryUsed for "I always need to know these"—environment facts, user preferences, project conventions.
Story LogUsed for "Did we talk about this last week?"—finding specific historical discussions.
Session Management
Each session in the session history can be named, resumed, exported, and cleaned up.
Naming and recovery
Hermes automatically generates a short session title (3–7 words) after the first conversation round, running asynchronously in the background without affecting response speed.
You can also name it manually:
Example
/title refactor authentication module
After naming, sessions can be resumed directly by title:
Example
hermes -c Refactor authentication module
# Restore the most recent session
hermes --continue
# Restore by session ID
hermes -r <session_id>
Auto-Lineage: when a session undergoes context compaction, Hermes automatically creates and numbers continuation sessions:
"我的项目" → "我的项目 #2" → "我的项目 #3"
When resuming by title, the latest continuation session is automatically selected.
Common session management commands
Example
hermes sessions list
# Filter by source platform
hermes sessions list --source telegram
# Show more entries
hermes sessions list --limit 50
# Rename specified session
hermes sessions rename <id> "New title"
# Export all sessions to file
hermes sessions export backup.jsonl
# Delete specified session
hermes sessions delete <id>
# Clean up sessions older than 30 days
hermes sessions prune --older-than 30
# View statistics
hermes sessions stats
hermes sessions listOutput example:
标题 预览 最近活跃 ID ──────────────────────────────────────────────────────────────────── 重构认证模块 帮我重构认证模块的代码 2 小时前 20260305_09... 我的项目 #3 检查一下测试失败的原因 昨天 20260304_14... — Python 装饰器怎么用 3 天前 20260303_10...
Automatic session cleanup (optional)
By default, session history is retained permanently (because it provides data for session_search).
If automatic cleanup is needed, enable it in the configuration:
Example
# Enable automatic conversation cleanup (default off)
sessions:
auto_prune: true # default false
retention_days: 90 # Keep the last 90 days
vacuum_after_prune: true
min_interval_hours: 24
Honcho user modeling
Honcho is an AI-native memory backend developed by Plastic Labs, integrated into Hermes as a plugin.
It doesn't replace MEMORY.md and USER.md, but rather builds on top of themAdd dialectical reasoning layer—by analyzing your conversation patterns, it automatically infers preferences, goals, and habits you haven't explicitly expressed.
Unlike the manual curation of MEMORY.md, Honcho is fully automatic: after each conversation turn (frequency configurable), it reasons in the background, accumulating a deeper understanding of you.
Honcho vs. built-in memory
| Skills | Built-in memory | Honcho |
|---|---|---|
| Cross-session persistence | Local files | Server-side API |
| User profile | Manually curated (USER.md) | Automatic dialectical reasoning |
| Session summary | — | Session-level context injection |
| Multi-agent isolation | — | Modeled separately per Agent |
| Semantic search | FTS5 full text | Based on reasoning conclusions |
Two-layer context injection
Each conversation turn, Honcho assembles two layers of context injected into the system prompt:
Base Layer (Base Context): session summary + user characterization + user Peer Card + AI self-characterization. Refreshed according to contextCadence (every N turns). This is the "who is this user" layer.
Dialectic layer (Dialectic Supplement): Generated in real time by Honcho's LLM engine, focusing on "what is most relevant right now." Refreshed according to dialecticCadence.
Two startup modes for dialectical reasoning:
| Mode | Conditions | Query method |
|---|---|---|
| cold start | No historical data | Generic query—"Who is this user? What are their preferences, goals, and working style?" |
| hot start | Existing historical data | Session-focused query — "Based on what has been discussed in this session, what is the most relevant context about this user?" |
Three independent adjustment knobs
| Parameter | Controlled content | Default value |
|---|---|---|
| contextCadence | Base layer refresh interval (calls API once every N turns) | 1 |
| dialecticCadence | Dialectical layer refresh interval (calls LLM once every N turns) | 2 (Recommended 1–5) |
| dialecticDepth | Reasoning depth per dialectic (1–3 passes) | 1 |
The three knobs are completely independent — you can refresh the base layer at high frequency and run dialectical reasoning at low frequency, or vice versa:
Example
"contextCadence": 1,
"dialecticCadence": 5,
"dialecticDepth": 2
}
The above configuration means: refresh the base layer every turn, and perform a 2-pass deep dialectical reasoning every 5 turns.
Multi-pass dialectical reasoning
When dialecticDepth is set to greater than 1, each dialectical reasoning run executes multiple rounds of LLM calls:
| Pass | Task | Description |
|---|---|---|
| Pass 0 | Cold/warm start reasoning | Establish initial assessment |
| Pass 1 | Self-audit | Identifies blind spots in the initial assessment, synthesizing evidence from recent sessions |
| Pass 2 | Reconciliation | Checks contradictions from the previous two turns, producing a final synthesized conclusion |
Depth 3 doesn't necessarily mean 3 LLM calls—if the previous turn already produced high-quality output, the next turn exits early to save tokens.
Quick start with Honcho
Example
hermes memory setup
Or manual configuration:
Example
# Enable Honcho as the memory provider.
memory:
provider: honcho
Example
echo 'HONCHO_API_KEY=your_key' >> ~/.hermes/.env
Get API Key at honcho.dev.
Tools provided by Honcho
After enabling Honcho, the Agent can use five exclusive tools:
| Tools | Purpose |
|---|---|
| honcho_profile | Read or update the user's Peer Card |
| honcho_search | Semantic search based on reasoning conclusions (raw fragments, no LLM synthesis) |
| honcho_context | Retrieve complete session context (summary, user characterization, recent messages) |
| honcho_reasoning | Call Honcho LLM inference (reasoning_level can be specified) |
| honcho_conclude | Create or delete reasoning conclusions |
Honcho CLI commands
Example
hermes honcho status
# View or set session policy
hermes honcho strategy
# View or Set Recall Mode (hybrid/context/tools)
hermes honcho mode
# View or set Token budget
hermes honcho tokens
# List known Honcho session mappings
hermes honcho sessions
Memory system configuration reference
The following is the complete reference for all memory-related configuration, which can be copied directly into your configuration file for use.
Example
# File path: ~/.hermes/config.yaml
# Hermes memory system complete configuration reference
# ============================================
# --- Persistent memory configuration ---
memory:
memory_enabled: true # Whether to enable persistent memory (MEMORY.md)
user_profile_enabled: true # Whether to enable user profile (USER.md)
memory_char_limit: 2200 # MEMORY.md character limit (approximately 800 tokens)
user_char_limit: 1375 # USER.md character limit (approx. 500 tokens)
write_approval: false # true = all memory writes require manual approval
provider: honcho # Optional: Enable Honcho memory provider
# --- Session auto-cleanup (optional, disabled by default) ---
sessions:
auto_prune: false # Whether to automatically clean up old sessions
retention_days: 90 # Keep sessions from the last N days
vacuum_after_prune: true # Whether to reclaim disk space after cleanup
min_interval_hours: 24 # Minimum interval between two cleanups
# --- Background review notifications ---
display:
memory_notifications: on # off | on | verbose
# --- Background review using cheaper model (optional) ---
auxiliary:
background_review:
provider: openrouter # Model Provider
model: google/gemini-3-flash-preview # Model Name
Quick reference for common memory operations
Directly manipulate memory in conversation
You can instruct Hermes to operate on memory using natural language, without needing to memorize any commands:
Example
Remember: this project uses Poetry for dependency management, not pip
# Delete or overwrite memory
Forget the previous records about MySQL, we have migrated to PostgreSQL
# Update information in memory
Update the staging server's IP to 10.0.2.100
# Query Memory
Tell me what you currently remember about me.
Slash Commands
Example
/memory
# View pending writes
/memory pending
# Approve all pending
/memory approve all
# Reject all pending
/memory reject all
# Enable Write Approval
/memory approval on
# Disable Write Approval
/memory approval off
Directly edit memory files
Memory files are plain Markdown, which can be opened and edited directly with an editor:
Example
cat ~/.hermes/memories/MEMORY.md
cat ~/.hermes/memories/USER.md
# Open in editor to modify
nano ~/.hermes/memories/MEMORY.md
Note: after directly editing files, changes take effect at the next session start and won't immediately appear in the current session's context.
Memory security and privacy
Security scan
Memory content undergoes a security scan before being written, detecting injection attacks and data theft patterns.
The system detects the following threat patterns:
| Threat type | Detection content |
|---|---|
| Prompt injection | Injection patterns attempting to override the Agent's system instructions |
| Credential theft | Instructions inducing the Agent to leak API keys or tokens |
| Backdoor installation | Installation instructions for SSH backdoors or remote access |
| Hidden characters | Invisible Unicode characters (zero-width spaces, directional overrides, etc.) |
Content matching threat patterns will beDeny write。
Sensitive information in logs is masked
Sensitive information such as API Keys and Tokens is automatically redacted in all log files—even if it appears in tool outputs, it will never be recorded in plaintext:
Example
# Enabled by default, automatically desensitizes sensitive information
security:
redact_secrets: true
Local data storage
All memory data (MEMORY.md, USER.md, state.db) is stored locally by default~/.hermes/Directory, not uploaded to the cloud
When choosing to use Honcho, conversation data is sent to Honcho's servers for processing—this is a deliberate trade-off: exchanging data uploads for deeper reasoning capabilities.
other extensions