Skip to content
Lazy Engineer's Tech Blog

Series: What AI Verification Misses

Putting the checks that verify AI-generated output under measurement to see what they actually miss: a merge audit that mistook a parent concept and its subtype for the same thing, a subagent-driven full audit that skimmed list-heavy files instead of checking item by item, re-verifying only against other cases while never rechecking the original false positive, an ingest step where sibling notes copied paragraphs verbatim and nothing caught it at first, and even the embedding filter built to catch that duplication flipping its verdicts depending on model size (small/large).