Series: What AI Verification Misses
Putting the checks that verify AI-generated output under measurement to see what they actually miss: a merge audit that mistook a parent concept and its subtype for the same thing, a subagent-driven full audit that skimmed list-heavy files instead of checking item by item, re-verifying only against other cases while never rechecking the original false positive, an ingest step where sibling notes copied paragraphs verbatim and nothing caught it at first, and even the embedding filter built to catch that duplication flipping its verdicts depending on model size (small/large).
-
몇 주 뒤, 그 감사 장치가 못 보는 경우를 찾았다
(Korean only)예전에 32번 중 32번을 다 맞혔던 병합 감사 장치가, 실전에서 상위 개념 노트와 그 하위 유형 노트를 동일 대상으로 오판했다. 원인은 판단력이 아니라 설계 자체 — 그 감사는 처음부터 이 종류의 오판을 잡도록 만들어진 게 아니었다.
-
Of the 4 most list-heavy files, the subagent missed 2
A full audit I delegated to Claude Code subagents reported 2 of 56 files as CLEAN when they weren't. An A/B test on the prompt showed that on files dense with bullet lists, 'I read it all' meant a skim rather than an item-by-item check, and a prompt that forces a checklist fixed it.
-
I verified on a different pair and never rechecked the original case
I fixed a false positive in my Obsidian plugin's LLM merge-judgment prompt and even closed the GitHub issues, but never retested the case that caused the false positive in the first place. Rerun days later under the real full-vault conditions, the judgment flipped, and actually performing the merge suggested the original call was likely wrong all along.