Pimp My IDE / manual service
Back to garage
October 3, 2026 | agents / documentation / durable context

Memory needs a service manual.

A searchable transcript can recover old words. A maintained document can tell the next agent what still governs the work. Those are different jobs.

Do not dump every useful sentence into one memory bin. Route rules, decisions, active work, and evidence into separate records. Give each record an owner and a reason to recheck it.

The argument against memory is really an argument against hidden retrieval.

Kevin Liao's October 3 essay attacks a common memory design. Session transcripts become small snippets, a similarity search picks a few, and the agent receives them without a clear view of what was omitted or allowed to go stale. His alternative stores specs, decisions, research, and indexes as readable Markdown that can be committed and changed.[1]

The article also introduces Operator Memory. Its public repository exists and describes a document loop with shared, project, and local Markdown. The repository is evidence that the implementation is inspectable. It is not evidence that document retrieval beats every memory system or that its automatic updates stay correct on a specific codebase.[2]

The useful split is not memory versus no memory. It is hidden recall versus inspectable project state.

One instruction file is an entry point, not a project history.

AGENTS.md defines a predictable place for setup commands, tests, conventions, and other instructions. Its documentation says nested files can carry closer instructions for subprojects, and the closest file takes precedence. That makes scope visible and keeps a root file from becoming a junk drawer.[3]

But rules are only one kind of durable context. A rule says how to work. A decision record says why one option won. An active task file says what remains. A proof record pins commands and results to a revision. Mixing those jobs creates conflicts that no filename convention can settle.

Long runs need a live work record.

Anthropic's current Opus 5.5 guide recommends a task file for long coding runs because older turns may be summarized as context fills. It also recommends defining the finish line and stop conditions, checking subagent evidence, and marking claims that could not be confirmed.[4]

That advice puts different clocks on different records. Project rules may last for months. A design decision lasts until its premise changes. A task list changes every few minutes. Test output belongs to one exact revision. The maintenance schedule should match the clock.

A document can be confidently stale.

Readable files fix observability, not truth. An agent can write a crisp explanation of code that changed five commits ago. A copied command can keep running while checking the wrong thing. An automatic updater can erase the disagreement that a reviewer needed to see.

Put three service fields beside durable context. Name the source that supports it. Name the person or process responsible for it. Name the event that forces a recheck, such as an API change, a dependency upgrade, a failed replay, or the end of the task. Keep transcripts available as evidence when needed, but do not make the transcript the policy.

Use the smallest document that has one maintenance rule.

Start with four homes. Put commands and conventions in scoped instruction files. Put tradeoffs and rejected alternatives in decision records. Put current goals and unchecked work in a task file. Put revision-bound outputs in proof records. Link between them instead of copying paragraphs.

The result is boring in the best way. A new session can find what governs the work. A reviewer can see who owns each claim. A stale record has a named trigger. Search remains useful, but it searches a workshop with labeled drawers instead of a pile of yesterday's conversations.

Interactive makeover / documentation service bay

Route the note before it hardens into policy.

This replaces one undifferentiated memory box. Write a note, choose the record that should own it, and select the maintenance fields your review will require. The generated card stays explicit about missing content and proof.

Rules
Decisions
Work
Evidence

Intake desk

The note is local to this page. Nothing is stored or sent.

Owning record
Maintenance fields

Service card

The lit drawer follows the native radio control. Selecting all fields prepares the card structure. It does not verify the note.

Route: rules0 of 3 fields selected
This tool drafts a routing card. It does not inspect a repository, create a document, identify a real owner, check freshness, run a command, or attach evidence.

Sources read

Source log and evidence boundary
  1. Kevin Liao, "Agents Don't Need Memory. They Need Documentation.", published and read October 3, 2026. It supplies the critique of snippet retrieval and the case for inspectable Markdown. It is an author's argument and introduces the author's own project.
  2. aerovato/operator-memory, inspected October 3, 2026. The public repository showed 86 commits, 35 tags, source directories, a BSD-3-Clause license, and documented adapters. This pass inspected the repository page and README but did not install or run the package.
  3. AGENTS.md documentation, read October 3, 2026. It defines the file as agent-focused project guidance, documents nested scope and closest-file precedence, and describes the format as plain Markdown.
  4. Addy Osmani, "Getting the most out of Opus 5.5 in Claude and Claude Code", published September 22 and read October 3, 2026. It supports the guidance on finish lines, stop conditions, task files, subagent evidence, and unconfirmed claims. Product comparisons on that page are vendor claims.
  5. Hacker News item 49945933, resolved through the official API and read October 3, 2026. It was the discovery route. Comments were not used as technical evidence.

Evidence boundary. The four-drawer model and service fields are Pimp My IDE's editorial recommendation. This page did not benchmark transcript search against document retrieval. It did not install Operator Memory or test how any named coding agent loads, updates, or prioritizes project files.