Skip to content
Lazy Engineer's Tech Blog
Go back

I verified on a different pair and never rechecked the original case

한국어로 읽기 →

What AI Verification Misses · Part 3 of 3

The Obsidian plugin I’m building has a feature that scans the whole wiki, finds two notes covering the same concept, and suggests merging them. Before a suggestion is carried out, an LLM is asked one more time: “Are these two really the same concept?” This post is about what happened while I was fixing that judgment prompt.

In Part 3 of another series (Korean only), I reopened a decision a few hours after closing it. This time I reopened it much later: an issue I’d closed two days earlier with a “confirmed fixed” note, and only when I looked at it again did I realize that what I’d confirmed back then was the answer to a different question.

Fixing one false positive, I missed a true positive on the other side — the merge-judgment prompt

Two notes, each about “the impact of AI on the labor market,” came up as a merge suggestion. One cited BLS data and interviews with labor economists; the other cited the Stanford AI Index and a McKinsey survey. They came from unrelated articles. Looking at the merge-judgment prompt, the cause was clear: it only said “overlapping title/scope means the same concept,” and had no criterion at all for checking “whether the sources point to the same specific event.” Any two notes explaining the same general topic with different evidence got judged “the same concept.” I added that criterion to fix it.

But when I reran the scan with the revised prompt, a completely different pair broke. Two notes on Hyundai Motor’s autonomous-driving strategy — a pair that was effectively about the same event, overlapping on everything from the January 2026 appointment of an executive from NVIDIA, to the same full switch to Alpamayo, the same SEooC approach, the same toolchain, and the same mid-2028 production target — were now wrongly filtered out as “different concepts” just because “the sources are different.” I’d set a criterion to catch the false positive, and it pulled too hard, so this time it missed a true positive.

I verified with the opposite case and closed the issues trusting only that

I relaxed the criterion again, adding a balance: “The fact that the sources differ is not, by itself, evidence of a different concept. When different outlets each report the same event, it’s the same concept even with two sources.” Naturally, I re-verified with the Hyundai pair I’d just missed. The pair that had been missed six times in a row was now caught correctly. A related issue looked resolved as a side effect too, and the next day both issues were closed with the judgment “effectively already resolved, no code changes needed, just close it.”

That verification wasn’t wrong. It was true that the Hyundai pair now merged correctly, and that check came from comparing the note bodies directly. The problem was elsewhere: the case that started this investigation in the first place, the AI labor-market pair, was never retested even once while the criterion changed twice. “Relaxing the criterion recovered the true positive” and “the original false positive still doesn’t get through even after relaxing it” are different questions, and I’d closed the second using only the answer to the first.

Days later, I reran the same pair under full-vault conditions

Days later, while revisiting that decision, I did something I hadn’t done back then. Instead of putting just the two notes of the AI labor-market pair into the prompt on their own, I ran it the way it runs in practice, with the full list of notes in the vault included. The results split. The judgment that came out correctly as “different concepts” when the pair was tested alone flipped under real conditions (one of the titles already had a _dup number appended to its filename because the titles collided, and embedding similarity was over 0.9). The LLM suggested merging the two notes again, citing “they share the same category and title” — the exact reason the prompt explicitly said not to use.

That meant the rule simply didn’t hold under these conditions. But once I’d confirmed that, a different question was left: should this merge really be blocked at all?

When I actually ran the merge, there was no reason to block it (topic vs. individual event rule)

Instead of predicting in words, I actually ran the merge. The result was a tidy document that preserved the facts from both sources as-is, with nothing made up. A follow-up scan didn’t suggest splitting it again either. Traceability back to each individual source also stayed intact through the notes that had been created separately for each source. On the surface, it was a result nobody would find odd.

So the premise that started this investigation — “if different sources merely cover the same general topic, they shouldn’t be merged” — was most likely wrong for this case to begin with. I removed this case from where I’d hard-coded it into the rule as an example (a worked example). I also reopened the two issues I’d already closed under that rule, leaving a comment saying the decision was under review.

What stays with me

Both issues had been closed as “not reproducible, verification complete.” None of those verifications lied — each was exactly right about the question it answered. But somewhere in between, the original question had quietly turned into a different one. Fixing a criterion and checking that it recovers a case it missed on the other side is not the same verification as checking that the criterion still judges correctly the case that made you create it. While I changed the judgment twice, I thought I was verifying the whole time, but I was actually checking while quietly swapping out what I was verifying.


Share this post:

What AI Verification Misses

  1. 1. 몇 주 뒤, 그 감사 장치가 못 보는 경우를 찾았다 (Korean only)
  2. 2. Of the 4 most list-heavy files, the subagent missed 2
  3. 3. I verified on a different pair and never rechecked the original case