Skip to content
Lazy Engineer's Tech Blog
Go back

Of the 4 most list-heavy files, the subagent missed 2

한국어로 읽기 →

What AI Verification Misses · Part 2 of 3

The Obsidian plugin I’m building takes source material (raw files), has an LLM read it, and splits it into wiki notes, one per concept. That means bugs keep turning up where something gets dropped or duplicated on the way from source to notes. To catch them, I’ve been handing a full audit to Claude Code subagents (separate AI agents you can split work across): each one compares a raw file, one at a time, against the notes that were generated from it. For each file, a subagent reports CLEAN if nothing is missing, or lists what is.

In Part 1 (Korean only), I found a kind of failure that the plugin’s own audit step (an extra LLM check before two notes get merged) couldn’t see by design. This time the hole wasn’t in the plugin. It was in the method I use every time to check those results: the subagent full audit.

I asked which issues had gone quiet, and didn’t get a real answer

Looking at the plugin’s GitHub issues, I asked Claude whether any of them hadn’t reproduced lately. Claude picked a few and answered based on the last full audit. So I asked back: “Why didn’t you look at 136, 143, 95, 97, and 50? Shouldn’t you check those while auditing every note?” It had just reported auditing all 56 raw files in three subagent batches, and then hadn’t actually matched those results against the issues one by one.

Two of the files the subagents called CLEAN were wrong

While matching the files the subagents had reported as CLEAN against the issue list, I opened a few files with a high density of (bulleted) lists myself. In “EDGE AI London 2026 행사 및 시장 인사이트” (an event and market recap), 14 partner companies were nowhere in the derived notes — exactly the same pattern as an issue that was already open. “AI is removing the middle class of software engineering” was the same story: a checklist item that had already been missed once earlier in the same work was missing again. I picked the 4 files with the highest list density and checked them, and 2 were wrong. Half.

The AI audit judged lists as a whole instead of counting items — a prompt A/B test

The instruction I’d given the subagents was: “Read the whole raw file, read all the derived notes, and compare.” That catches a short quote or a single number well. But faced with a list of a dozen-plus items, it sometimes judged “it’s all there” from the overall impression instead of checking the items one by one. Giving the same file and the same instruction again reproduced it, so this wasn’t bad luck; it was a structural gap in the instruction itself.

To confirm it, I ran an A/B test on the same file with prompt A (the original “read and compare”) and prompt B (“First turn every bullet and list item in the source into a numbered checklist, then judge each item individually as present or absent. Don’t lump several items into one sentence and call them ‘covered’.”). A finished in 71 seconds and missed the same item again. B took 304 seconds (4.3x as long) but correctly flagged that item as MISSING. I also checked that both misses sat in the middle-to-late part of the batch, not at the very end — so it wasn’t “it got tired and skimmed.” The cause was the list density inside the file.

A prompt that forces a checklist caught the missing list items (spot-checking by bullet count)

Using prompt B for whole batches wasn’t realistic (4.3x slower, and the reports get too long to handle). Instead I added one more step: after a batch audit, rank the raw files by bullet count, and spot-check the top ones myself — including the ones the subagents reported as CLEAN — rather than handing them back to a subagent. Only the files where a real omission turns up go to prompt B for a precise re-check. I wrote this procedure down in one document, so the next time I run this audit I don’t have to rebuild the prompt from memory.

What stays with me

Part 9 of another series (Korean only) and Part 1 of this one were both about measuring how far the plugin’s audit step could really be trusted. This time I had to put the method I used to check that audit step — handing a full comparison to subagents — under measurement itself. “The subagent said CLEAN” and “the subagent actually checked every item” turned out to be different statements, and the difference only showed up, quietly, in files with dense lists. Handing audits to a tool still makes sense. I just relearned that before taking its “I checked everything” at face value, I need to count at least a few by eye myself.


Share this post:

What AI Verification Misses

  1. 1. 몇 주 뒤, 그 감사 장치가 못 보는 경우를 찾았다 (Korean only)
  2. 2. Of the 4 most list-heavy files, the subagent missed 2
  3. 3. I verified on a different pair and never rechecked the original case