Posts
All the articles I've posted.
-
I verified on a different pair and never rechecked the original case
I fixed a false positive in my Obsidian plugin's LLM merge-judgment prompt and even closed the GitHub issues, but never retested the case that caused the false positive in the first place. Rerun days later under the real full-vault conditions, the judgment flipped, and actually performing the merge suggested the original call was likely wrong all along.
-
Of the 4 most list-heavy files, the subagent missed 2
A full audit I delegated to Claude Code subagents reported 2 of 56 files as CLEAN when they weren't. An A/B test on the prompt showed that on files dense with bullet lists, 'I read it all' meant a skim rather than an item-by-item check, and a prompt that forces a checklist fixed it.
-
Wrong, but 90% confident
I tried Laya, the open-source reproduction of TypeSafe's Jev, on my Obsidian plugin's duplicate-note decisions (Jev itself only through its docs, since my account isn't open yet). I measured CPU speed, 78% accuracy, and what fine-tuning would take, and the real problem turned out to be broken calibration: confidence around 0.9 on wrong answers.