DeeSeek Harness Postmortems and Engineering Culture
The dsh repository turns every bug that "shouldn't have happened but did" into a postmortem, and leaves guards at the test, documentation, and rule levels.
Understand this postmortem culture and testing discipline, and you can put it to work in the example project: not just use dsh, but think like a maintainer.
In one sentence: a bug's value is not in the one-line fix, but in why the process let it through, and what guard was added so similar issues fail loudly next time.
Retrospective Culture: Four Questions
A postmortem records: a bug appeared where it shouldn't have — in front of real users, in a merged PR, or in a released version.
What's worth noting isWhy did our process let it through?, rather than just that one-line fix.
It is a retrospective record of failure, answering four questions:
| Four Questions | What to answer |
|---|---|
| What's broken? | A short paragraph that lets busy readers absorb the key points in thirty seconds. |
| What is the mechanism? | State the root cause in plain words, without blaming individuals. |
| Why every safety net failed to catch it | Identify gaps in tests, tooling, and conventions — not a one-off typo. |
| What new protections were added | Tests, AGENTS.md rules, ADRs — make similar bugs fail loudly next time. |
Not every bug deserves a postmortem.
Only write one when the bug meets all three conditions:Hidden(the mechanism is non-obvious; even a careful engineer would have to work hard to re-derive it),Systematic(the reason it escaped was a gap in tests, tooling, or conventions),The cost of rediscovery is high(it consumed real debugging time, and will again next time).
Four Real Cases
The repository currently has four postmortems, each corresponding to a class of engineering mistakes that tends to recur.
I'll go through each one with the focus on "why it escaped" and "what guard was added."
Review 0001
ACP server crashes on connection: export default drops the plugin's inject.
What broke: the moment the editor (Zed) connected, the first
session/newJust reportcannot get property "agents" without inject。Mechanism: The plugin wrote one extra.
export default apply, Loader'sunwrapExportsIt retrieved the bare function, and the one on the namespace...injectLost it entirely.Why the safety net didn't catch it: 178 green unit tests + 100% line coverage were all there, but every test was manually
ctx.plugin(...)mounted, bypassing the real Loader's loading path.Guards added: removed the default export; added a keyless real-Loader smoke test; rule: "test the real entry path — line coverage is not behavior coverage."
Review 0002
The filesystem snapshot tool was permanently disabled by a literal !!js object.
What broke: seven filesystem scenarios invoked a tool that does not exist in the registry, returning...
UNKNOWN_TOOL。Mechanism: The author uses
disabled: !!js ...We wanted to conditionally enable the filesystem plugin, but Cordis only evaluates JS expressionsconfiginside the plugin; when directly readingdisabledthe config item, it saw a truthy object.Why the safety net didn't catch it: snapshot refresh treated "deterministic replay" as "correct behavior"—it proved the regression was stably reproduced, but did not prove that the filesystem tool was actually registered.
Guards added: switched to an explicit filesystem overlay; a static config guard rejects expression nodes in Loader config metadata; the snapshot framework rejects structured
UNKNOWN_TOOLResult.
Review 0003
The web agent accepted a replacement server rather than the GUI hosting its session.
What broke: the agent modified the GUI source code, but didn't know which URL the current session corresponded to or which process was hosting it.
Mechanism: it treated the HTTP 200 returned by the bare Vite server as success (actually a blank screen), then went to accept a replacement
dsh webserver on another port, never probing port 3081.Why the safety net didn't catch it: the Web composition did not provide the model with identity information about the current GUI, canonical URL, or running mode; the first regression test also used "process timeout" to impersonate "fail fast", producing a false positive.
Add protection:StartdevicepublishSpecificationCycle回 URL andreal际生产/DevelopmentMode(Environment Variables + Hintword区segment);Independent Vite ServicesModeInConfigurationPhaseRejectStart;DividelayertruerealpathTesting覆盖 CLI、Hintword、runtime事realand浏览device HMR.
Review 0004
Landlock partial-enforcement notices caused child-process failures to be misclassified.
What broke: on kernels with an older Landlock ABI, ripgrep exits normally with code 1 when there are no matches, but it was presented as a failure.
SANDBOX_UNAVAILABLESandbox failure.Mechanism: the launcher prints a harmless
landlock-run: partial enforcement (older Landlock ABI)notice; the harness uses a case-insensitivelandlock-run:substring combined with any non-zero exit to misclassify it as a runner failure.Why the safety net didn't catch it: the sandbox result type can only express a set of substrings; it cannot express "Landlock failure must exit with code 125 + one line of fatal diagnostic"; the test matrix never constructs the combination of "notification followed by non-zero subprocess exit".
Add protection:
RunnerFailureRuleCarry informational lines with allowed exit codes, per-line fatal signatures, and precise exclusions; filesystem search switches to...ctx.subprocessRun the bundled ripgrep, no longer routing it through sandboxed bash.
Looking at the four cases together, a common lesson emerges:Tests must go through the real entry path。
Manual mounting, mocking everything, and treating snapshot refresh as acceptance — all of these make "unit tests all green, product broken" possible.
Four test layers
The repository's testing strategy is layered; each layer covers blind spots the layer before it cannot catch.
| Tier | Command | What to capture |
|---|---|---|
| Unit Testing | pnpm run test | Vitest runs in-package tests, prioritizing edge cases, error paths, event ordering, and concurrency races. |
| Coverage gate | pnpm run test:coverage | 100% coverage per file; uncovered lines are usually dead code that should be deleted. |
| Real API e2e | pnpm run test:e2e | Calls real provider APIs with keys; automatically skips when keys are missing, keeping keyless CI green. |
| Snapshot | pnpm run test:snapshot / test:web | Keyless expected outputs cover external behavior; browser snapshots are replayed and compared using Chromium. |
The repository is DeepSeek's own, so there is one special principle:Inference is cheap here, don't skimp on real API tests.。
Keyless tests can only prove the underlying pathway; only running with keys can prove that the agent works correctly with a real model.
The highest-value ones are smoke tests: boot a real example, send a single prompt, and inspect the outside world.
They catch the "unit tests all green, product broken" class of problems that mocks cannot discover.
Line coverage is a necessary condition, never a sufficient one.
It proves a line was executed, not that the feature works as delivered.
In case 0001, 100% coverage still let two integration bugs slip through — the perfect footnote.
Another principle isVerify the external world, rather than self-report。
E2E assertions should rerun the command or re-read the file externally; keyword probing of the agent's own output would let a cheating agent pass.
Assert that unmodified files remain byte-for-byte identical.
Documentation discipline: one fact, one home.
The repository's documentation is not "done once written"; a mechanism prevents "docs and source drift."
verify-type-equivIt is a gate: it uses a TypeScript parser to extract type-declaration symbols and their attached JSDoc from the source, and asserts that documentation code blocks match both.
When you change a documented type declaration or its JSDoc, the gate fails until you update the pasted content.
This guarantees that type definitions in the docs always match the source — no "docs copied from an older version."
Another discipline isOne fact, one home: each fact is maintained in exactly one file; all other files reference it.
For example, the "source of truth" for the tool schema is inadding-a-tool.md, other pages reference rather than copy.
Chinese and English documents are maintained as bilingual pairs; when updating, first run...pnpm run gen-doc-graphsUpdate English, then update Chinese and verify the pairing.
This applies to your own projects too: keep the "single source of truth" in one place; don't scatter the same config across three documents.
Hands-on example: write a thirty-second executive summary.
Every postmortem opens with an executive summary so a busy reader can absorb the key points in thirty seconds.
Try writing a postmortem for an Agent failure you encountered in the example project, using the template below:
# 复盘 0005:example 演示环境的工具黑名单没有生效 ## 摘要 demo 环境里模型仍能调用被禁用的 fs_write,直接写入了工作区。 根因是黑名单插件用 `tools/pre-execute` 返回了 `next()`, 而真正负责拦截的守卫注册在插件卸载之后——加载顺序错了。 逃逸原因是加载顺序没有测试覆盖,`cordis.yml` 只测了「能起来」。 新增防护:给黑名单插件补一条「顺序敏感」的装配测试, 并在 AGENTS.md 记录「pre-execute 与 guard 的注册顺序」规则。
The formula for a thirty-second summary: what broke → root cause in plain words → why it escaped → a lesson you can carry forward long-term.
Series conclusion
At this point, all 28 installments of the DeepSeek Harness introductory tutorial conclude.
Look back along the road: from the first article, learning dsh's four guardrails, to installing, using the Web UI, and using the Python SDK; from writing your first plugin, to understanding plugin lifecycle, services and scopes, and the event system; from designing capabilities with the three-role pattern and integrating any LLM, to packaging and releasing. Finally, these last five articles complete the last piece of the puzzle: security and engineering culture.
What you truly take away is not just how to use the API, but a way of thinking:
The model is responsible for intelligence, the Harness is responsible for reliability.Constraints are not restrictions; they are the foundation that makes Agents predictable, auditable, and replayable.
Everything is a plugin.Put policy at extension points, not inside the loop, and the system can evolve without losing control.
The real entry path beats all mocks.Coverage, snapshots, and smoke tests each do their own job, together preventing "all green yet broken."
Failure is not scary. What's scary is not knowing why.Postmortem culture turns a single incident into a set of defenses.
other extensions