A reviewer that forgets is a linter · Brandon Miller
·6 min read· agents· ai· review· tooling

A reviewer that forgets is a linter

Every review agent on Quartra keeps its own memory directory: indexed, one file per lesson, written at the moment it learned the thing. Four entries in, it had already caught a test assertion that always passes.

Quartra's iOS and Android trees each carry their own agent harness under .claude/: agents, commands, skills, hooks, and a project.json that every template renders from. Most of that is scaffolding I'd expect any project to grow. The piece that has earned its keep, and the piece I'd port first to anything else, is .claude/agents/memory/.

Every agent gets a directory. Inside it: MEMORY.md, an index of one-line hooks, and one file per lesson. At the start of a task the agent reads the index and opens only the entries whose hook touches what it's about to do.

The format

---
name: samsung-muxer-truncates-last-sample
description: MediaMuxer on Samsung API 33 drops the final sample unless the track is flushed before stop()
type: gotcha
date: 2026-08-31
---

MediaMuxer.stop() on Samsung devices running API 33 finalizes the moov atom without
the last queued sample, producing a file that plays but is ~40ms short. Reproduced
on SM-S911B, not on Pixel 7 or the API 33 emulator.

**Why it matters:** the truncation is silent — the output is a valid MP4, so
validation that only checks "does it open" passes.

**How to apply:** validate output duration against the expected segment duration
after every remux, and fall back to re-encode when it is short.

type is one of gotcha, decision, baseline, feedback, reference. Dates are absolute, never "last week". Related entries link with [[wikilinks]], and a link to an entry that doesn't exist yet is fine. It marks something worth writing.

Then exactly one line goes into MEMORY.md.

The index is the load-bearing part. Reading a memory corpus costs tokens on every single task; reading an index of hooks costs almost nothing. The agent pays for recall by the hook, not by the corpus. Which is why the protocol is blunt about hooks having to be specific:

"Samsung MediaMuxer truncates the last sample" beats "muxer bug".

What's actually in there

Four entries under swift-reviewer, four under swiftui-reviewer, one under kotlin-reviewer. Every other agent's directory is still empty, which is the correct state: the harness ships the structure and nothing in it. Two entries are worth reading in full.

The test assertion that always passes. SourceImport.copy writes to Library/Caches/imports/<UUID>/<original filename>. The picked filename only ever appears on the file inside the folder, never on the folder itself, which is a bare UUID. So a cleanup assertion shaped like this:

contentsOfDirectory(at: SourceImport.directory)
    .contains { $0.lastPathComponent.contains("Song Name") }

is vacuous. It passes whether or not the failure path cleaned anything up. It was sitting in BackgroundJobTests, in a test named whatNothingCanDecodeFailsWithAMessageWorthReading.

That is not a bug you find by reading the diff. The diff shows a test that reads correctly and asserts something plausible. You find it by knowing how SourceImport lays out its directories. That is precisely the kind of thing a reviewer either carries in its head or doesn't. The entry ends with a rule aimed squarely at the next review: when looking at any SourceImport cleanup test, check what level of the tree the assertion actually inspects. Plus a second-order note, because swift-testing runs @Test functions in parallel against one shared real directory, so a naive before/after set diff will flake.

The disabled state that isn't. DesignSystem.swift has no isEnabled handling anywhere, and the Library and Stems buttons use .buttonStyle(.plain) with explicit foreground colors on their content. So .disabled(true) produces no dimming at all. The control looks fully live and silently eats taps.

What makes that a good memory rather than just a good bug report is the last line: VoiceOver still announces "dimmed". The bug is sighted-user-only. An accessibility audit passes it. It turned up on the import gating, where the picker swapped its label and showed a spinner (fine) while the Recent rows and the sample row were gated with .disabled alone and looked completely untouched for the entire import.

The rule that went in: when reviewing any new .disabled(...), require a visible state change on that same control (a label swap, a spinner, an opacity change), not just the modifier.

The half that decides whether this survives

Anyone can create a memory directory. What determines whether it's useful in six months is the list of things that are not allowed in it, and that list is longer than the list of things that are:

  • Anything derivable by reading the code. "The export path lives in ExportRepository" is not a memory; it is a grep.
  • What git already records. "Fixed the null check in ViewModel". The commit says that.
  • Restatements of CLAUDE.md. Already loaded.
  • Task narration. "I implemented the feature and it worked."
  • Duplicates. Check the index first and update the existing file. Two memories that disagree are worse than one that is stale.
  • Speculation. If you didn't verify it, don't write it as fact. Say what you observed and what you inferred, separately.

And the filter that does most of the work:

Could you have written this memory without doing the task? If yes, do not save it.

Almost everything an agent wants to write down at the end of a task fails that test. Which is the point: the default behaviour of a model asked to record lessons is to produce a tidy summary of what it just did, and a directory full of those is worse than an empty one, because now the index costs something to read and returns nothing.

Memory rots, and the protocol treats that as a reading rule

Not as a maintenance chore someone will get to:

Verify before relying. A memory records what was true when it was written. If it names a file, function, dependency version, or flag, confirm that still exists before acting on it. A stale memory that sends an agent to a deleted file costs more than no memory at all.

Maintenance is deletion, not annotation. When an entry turns out to be wrong, the file and its index line go. It does not get a correction appended to the bottom. When two cover the same ground, they merge. The index stays short enough to read in full at the start of every task, and that constraint is what forces the pruning.

Why this is the piece I'd port first

A reviewer that forgets is a linter. It knows a fixed set of rules and applies them uniformly, which is useful and also the ceiling on how useful it can be.

Most of what makes a senior reviewer valuable on a codebase they've worked in for a year isn't general knowledge about Swift or Kotlin. It's local, specific and almost entirely undocumented: this API lies to you, this test is vacuous, this modifier does nothing here, this helper looks shared but has one caller. Nobody writes those into a CONTRIBUTING.md, because you learn them at 11pm while fixing something else and the knowledge never gets promoted anywhere.

An agent will write them down, if you give it somewhere to put them and a strict rule about what's worth putting there.

One honest caveat: the entries cluster almost entirely in the two reviewer agents. That tracks (review is where you find the surprising thing), but it also means the format is still unproven for the build, QA and product agents, whose directories have been empty since day one. It's possible the protocol only really fits agents whose job is to look at code someone else wrote.