Skip to content
Lazy Engineer's Tech Blog
Go back

Wrong, but 90% confident

한국어로 읽기 →

I’m building an Obsidian plugin on the side (not public yet). You feed it source material, and an LLM reads it and organizes it into one note per concept. Feed it enough material over time and the same note title sometimes gets created twice. When that happens, the plugin sends both notes to the OpenAI API for a second look. If the same concept just happened to be created twice, it merges them; if the titles match but the content is different, it renames one so they’re distinguishable.

I’d been thinking it was wasteful to spend a full LLM call on every little “merge or not” decision, and right then I heard about models built to do only this kind of “low-level judgment”: TypeSafe’s Jev, and its open-source rival Laya. I wanted to see whether one of them could actually take over the decision I’m making today.

What a “System One” model is (ModernBERT, 421M)

Jev and Laya aren’t LLMs. They don’t generate text. You give them a situation (the state) and a set of typed questions, and they hand back answers with probabilities. They answer with just three primitives: choice (one of the options), score (an ordinal score), and noul (the probability that something is true). The point is that there’s nothing to parse and no free-form language to hallucinate in. Jev is proprietary and came out in early access on 2026-09-15; Laya is an open-source project that reproduces it, built on ModernBERT-large, 421M parameters, Apache 2.0.

I had the wrong source, twice — GGUF and a GitHub mirror

While trying to actually run Laya, I went down the wrong path twice. First, I saw a “GGUF” file on Hugging Face and assumed I could load it straight into Ollama or LM Studio. But it wasn’t the standard GGUF that llama.cpp uses; it was a format produced by a separate compiler called ggmlc. The real Ollama app won’t open it.

The second mistake was picking the wrong original repository. The first GitHub repo I found (he-jev/laya) looked like the real thing: its README matched the actual model information exactly. But the “Links” section at the very bottom of that README pointed to a different address as the GitHub original, not to itself. When I checked, the real original was NandhaKishorM/laya with 27,891 stars, and the one I’d found first was a mirror with 7 stars. Accurate content didn’t mean it was the original.

Speed — the advertised 33 ms was a GPU number

I tested on my home PC with no GPU (12 cores, 23 GB of available RAM). I installed it with pip install laya and pulled the English checkpoint (convaiinnovations/laya). Cold load took 23 seconds (including the model download), and memory after loading was 2.7 GB. Latency per call averaged 1.3 seconds.

That’s more than 30x off from the “33 ms, 39.5 ms” figures the README leads with, so I went back and checked: those benchmarks were measured on a Tesla T4 GPU. My test was CPU-only, and a gap like that is normal when you run a 421M-parameter transformer on a CPU. The “microsecond-scale gating” marketing line was about machines with a GPU; on a CPU-only machine, I had to redo the math in seconds.

Accuracy — scored against real past cases (18 of 23, 78%)

As it happened, the note vault I’d been using to test the plugin still had old duplicate notes sitting in it. They were genuine duplicates, created when the same original article ended up in two different source files. I opened both the original and the duplicate of each of these 17 pairs myself to confirm they really were the same concept (the right answer was “same” for every one), then asked Laya: “These two notes share a title. Are they the same concept or different concepts?” On top of that, since no real cases of “same title, completely different concept” were left in the vault, I built 6 of them by hand from pairs of unrelated notes.

The result: 18 of 23 correct (78%). It got 14 of the 17 real duplicates and 4 of the 6 synthetic “different concept” cases.

The real problem wasn’t accuracy, it was confidence (temperature)

Going by the number alone, it looked like “not there yet, but maybe worth trying.” But when I went through the answers one by one, there was a worse problem. Plenty of the correct answers had confidence like 0.01, 0.011, or 0.016, which is basically a coin flip. And, decisively, there were cases where it called two notes with completely different content “the same,” got it wrong, and showed high confidence doing it: 0.802, 0.896.

A warning had been there since the moment the model loaded: this checkpoint has a corrupted temperature value, so its confidence can’t be trusted. The model was telling me itself. The whole value of a System One model like Jev or Laya is that you can trust the probability and gate automatically on a threshold, and with this checkpoint that premise had already collapsed.

Just fine-tune it? The bottleneck wasn’t the GPU (RLCD, Kaggle 2×T4)

According to the official docs (docs/finetune.md), the procedure goes like this: train with RLCD on Kaggle’s free 2×T4 GPUs, hold out up to 400 examples from the training data, and re-fit the confidence temperature on them. The GPU is free, so that wasn’t the problem.

The real bottleneck was data. What training needs isn’t correct labels but “the state + the questions + the probability distribution a teacher model assigned to each option.” Even the demo in the docs uses 1,200 cases (6,000 questions), and the real benchmark is on the order of 30k cases. I have only 23 real cases. Doing meaningful fine-tuning would mean steadily collecting every duplicate decision the plugin makes, along with its probabilities. Not a one- or two-day job.

Would Jev be different? It can’t be fine-tuned at all (no LoRA)

Seeing how much work fine-tuning Laya would take, I thought “then why not just use Jev?” and looked into it. TypeSafe’s answer pointed the opposite way. Jev can’t be fine-tuned on customer data, and LoRA isn’t available either. Every account shares the same weights, and customization happens only at the prompt level: filling state with your domain data and spelling out your decision criteria in instructions/criteria.

So the two go in opposite directions. Laya’s strength is that you can improve it yourself, but to make it usable you have to go through all the data, GPU, and calibration work I just described. Jev skips that work entirely, but its accuracy ceiling is fixed at whatever the vendor’s model can do.

Conclusion

I decided not to use Laya any further for now. Duplicate-note decisions will keep going through the OpenAI API as they do today. That approach is already proven on both accuracy and how far its confidence can be trusted, and there was no reason to swap it for an uncertain path. Jev hasn’t been opened up to my account yet, so I couldn’t even try it this time; I’ll take another look once it is.

What stays with me

The biggest lesson from this wasn’t the accuracy number itself. It was that you can’t trust a “confidence” value just because the marketing copy says so. If I’d only looked at 78%, I might have shrugged and said “could be usable.” Opening the answers one by one showed the wrong ones actually carried higher confidence. The speed benchmark was the same story: I nearly took 33 ms measured on a GPU as what to expect on a CPU. It was one more day of relearning that you always have to check what conditions a number came from.


Share this post: