Skip to main content
Research•8 min read

I tested Jev's confidence. Don't trust the number yet.

October 1, 2026 by Asif Waliuddin

evalscalibrationjevtypesafemodel-selection
I tested Jev's confidence. Don't trust the number yet.

TypeSafe launched Jev with a sharp promise. Typed decisions, not strings, and every answer carries a probability you can build on. The TypeSafe homepage says Jev "returns typed decisions with calibrated probabilities, so your software can account for uncertainty." The launch post puts it more plainly: "Calibrated: higher confidence means higher accuracy." I tested that on 273 held-out decisions with known answers. Jev was 66% confident on average and 53% right.

That is exactly the property I want. At 0.95, auto-accept. At 0.6, ask a human. That pattern only works if the number means what it says.

So I went looking for the calibration data. On the homepage and the launch post, as of 1 October 2026, I found speed and cost numbers but no Brier score, no ECE and no reliability curve. So I measured it.

What I ran

I wrote the bars down before a single call was made. The pre-registration fixes the questions, the metrics and the verdict words. I committed it before the first call, but that commit sits in our private repo, so you have to take its time on trust. That is why I published part 2's plan before running it.

Three models answered the same 340 questions:

  • Jev (jev-1.13.0), using the probabilities it returns.
  • Claude Opus 5.5, my frontier baseline, asked to state a probability per option (stated, not measured).
  • Qwen3-14B, running locally on one 4090, with probabilities read from its own token logprobs.

Every question has a right answer that was on record before the test, not produced by any model under test. 160 come from CUAD, the Atticus Project's contract dataset with expert annotations, released under CC BY 4.0: "which of 20 clause types is this?" and "does this passage contain clause X, yes or no?" The other 180 come from my company's own operations records: who wrote this team post (captured by the system), which program a work item was assigned to, and which lane was named owner of a commitment when it was created. Some of those assignments were made by my own agents, so read them as recorded decisions, not human labels. Names, mentions and IDs were stripped from those on purpose.

Unless noted, every number below is on the 273-item held-out test set.

Where Jev wins

Accuracy on contract clause typing: a tie with Opus. On the 20-way clause question, Jev and Opus were each right 74% of the time (46 test items). On yes/no clause checks, Opus led, 95% to 87%.

Speed and cost: not close. Median latency was 0.15 seconds for Jev and 6.5 seconds for my Opus arm, about 44x faster. At list price, a Jev decision cost $0.0000387 against $0.022 for Opus, about 570x cheaper. The entire Jev bill for all 1,020 calls in this run came to about four cents.

Two footnotes: my Opus latency includes claude -p start-up, and its cost is priced at list rates although I ran it on a subscription. Both flatter Jev's ratio.

Jev vs. my Opus arm: about 44x faster and about 570x cheaper per decision

How that compares with TypeSafe's own claim. The homepage says 193.6x faster and 444.6x cheaper, measured on TypeSafe's own workflow evals against the average of GPT-6 Astra and Fable 5.1. The launch post says those numbers are "on the higher end of real world gains" and that typical speed-ups "range from 40x-200x faster." Different comparator, so context, not a verdict: my cost gap came out larger than theirs, and my speed gap sits at the bottom of their range.

The "25x faster, 76x cheaper" figures you may have seen are not TypeSafe's. They are third-party arithmetic on one launch demo, and I quoted them myself before tracing them (see corrections below).

Where it doesn't hold up

Jev's confidence runs ahead of its accuracy. Across all 273 test decisions its average stated confidence was 66%. It was right 53% of the time. The expected calibration error (ECE: the average gap between what the model says and how often it is right, weighted by how many answers sit at each confidence level) was 0.156. My pre-registered bar called anything above 0.10 "not calibrated as claimed."

I judge per class, because a pooled number can hide a failing one. Jev is not calibrated as claimed on 4 of 5 classes, in all three repeat runs. The fifth, yes/no clause checks, read "roughly calibrated" in two runs and missed the line by 0.002 in the other. I call that one draw-sensitive, not calibrated.

On clause typing, where Jev ties Opus on accuracy, its ECE was 0.20. Use the label. Don't set a threshold on the number.

A standard fix, one temperature fitted on held-out data, barely moved it: 0.156 to 0.150.

Reliability diagram: confidence against accuracy for Jev, Opus 5.5 and Qwen3-14B

Opus is not a stable reference either. I asked Opus the same 100 questions three times. On 21 of them it did not give the same answer all three times. The second ask changed the first answer on 16 items, the third on 13. On the same 100 questions, Jev was not identical on 7. Opus's pooled confidence looks close to honest, but its errors cancel within classes: on clause typing its average confidence matched its accuracy almost exactly, yet its ECE was 0.16, overconfident in some bins and underconfident in others.

The local model is wildly overconfident, and that part is fixable. Qwen3-14B averaged 91% confidence while being right 51% of the time, an ECE of 0.40. One temperature, fitted on the held-out calibration split, brought that to 0.067. That fixes the confidence, not the accuracy, which stayed at 51%. Its accuracy trailed the best non-local model on all five classes, so it did not earn the local tier on any of them.

The result that surprised me

No model could route my own operations work from bare text. On "who wrote this post", "which program owns this" and "which lane owns this promise", the best model (Opus) was right 21%, 58% and 44% of the time. Jev was lower on all three.

I suspect context before the model: those items had exactly what a real router would use stripped out. Part 2 tests it directly, with the most similar earlier records from my memory store added. Its pre-registration went public before a single context call. I don't know the answer yet, and I will publish it either way.

How I kept myself honest

  • Bars before data. Thresholds, metrics and verdict words were fixed in the pre-registration. Any later change is timestamped there with its reason.
  • A second reviewer re-derived the headline tables from the raw files. The calibration verdicts, the local-model verdicts, the speed and cost figures, and the pairwise comparisons all matched. An independent cross-vendor reviewer checked the method separately and built its own confidence-versus-accuracy table. The repeat-run flips, the temperature fixes and the routing coverage are still single-author, and I label them that way.
  • I corrected myself on the record, and kept the originals. My first cost ratio said about 1,080x; it compared two different item populations, and the fix is 570x. My pre-registration quoted 25x and 76x as TypeSafe's claim; they are not TypeSafe's numbers. And my own disclosure stated the local model ran at temperature 0, which was wrong; the run used the model card's recommended settings. Each correction is a dated commit, and the original wording stays in the history.

Check your own model in 20 lines

If your model returns probabilities, you can run this today. Run on Jev's 273 test decisions, it prints accuracy 0.531, mean confidence 0.662, ECE 0.145. That is with equal-width bins; my headline 0.156 uses equal-mass bins, the pre-registered method.

import json
 
def ece(confidences, correct, n_bins=15):
    """Expected calibration error with equal-width bins.
    confidences: the probability the model gave its own answer (0..1)
    correct:     1 if that answer was right, else 0"""
    bins = [[] for _ in range(n_bins)]
    for c, y in zip(confidences, correct):
        bins[min(int(c * n_bins), n_bins - 1)].append((c, y))
    total, n = 0.0, len(confidences)
    for b in bins:
        if b:
            avg_conf = sum(c for c, _ in b) / len(b)
            accuracy = sum(y for _, y in b) / len(b)
            total += len(b) / n * abs(avg_conf - accuracy)
    return total
 
# one JSON object per line: {"probabilities": {...}, "answer": "...", "gold": "..."}
rows = [json.loads(line) for line in open("decisions.jsonl")]
conf = [max(r["probabilities"].values()) for r in rows]
hit = [int(r["answer"] == r["gold"]) for r in rows]
print(f"accuracy {sum(hit)/len(hit):.3f}  mean confidence {sum(conf)/len(conf):.3f}  ECE {ece(conf, hit):.3f}")

Under 0.05 is calibrated. Over 0.10, don't build a threshold on it until you have measured it per class.

What I'm doing with it

Jev is now my pick for contract-shaped classification. It matches Opus on clause typing at a tiny fraction of the cost and time, and on yes/no checks a confidence gate can auto-accept about 81% of items at 95% precision. I will not treat its probabilities as calibrated until I have measured them per class, and neither should you.

No local tier on Qwen3-14B. No routing model until Part 2 reports.

TypeSafe built something fast and cheap and genuinely useful. The number it hands you is a good start. It is not yet a promise.

The eval code, the public contract items and every per-call record for them are at github.com/nxtg-ai/jev-calibration-eval. The pre-registration is here.

Ready to build?

Ship AI you can trust

Forge gives you agents, governance, and verification — so your AI ships with confidence, not hope.

Newsletter

Enjoyed this article?

Get more insights like this delivered straight to your inbox.

Email subscription coming soon. Follow along on LinkedIn in the meantime.

Follow on LinkedIn