Article

Does a decision model (like Jev) know when it is guessing? (Post 1)

Sharath Shankaranarayana
October 1, 2026

A model can be wrong and unsure, or wrong and so confident that nobody checks.

A new kind of AI model, the decision model (also called a System One model), is built for automated decisions. It doesn't write text. It picks from a fixed set of options and returns a probability for each one. Software can act automatically when the model is sure and send the case to a person when it isn't. We tested the best-known one, Jev from TypeSafe AI.

Diagram: a confidence gate in production. A message goes to Jev, which returns probabilities. If confidence is above the threshold the system acts automatically, otherwise it sends the message to a person. A model that knows and is sure is automated correctly. A model that has no idea but sounds sure passes the gate and causes a wrong action that no one reviews.
Figure 1. The confidence gate. A common way to decide what to automate, usually alongside business rules and monitoring.

Picture a bank routing thousands of support messages an hour. Nobody reads each one, so the confidence number decides which ones a person sees. If the number drops when the model doesn't know the answer, the second model in Figure 1 gets stopped. If it doesn't, the mistake goes straight through. TypeSafe says Jev's confidence does drop. Its launch post promises:

"Calibrated decisions: answers with epistemically honest probabilities on System One tasks." TypeSafe, Introducing System One models and Jev, accessed 29 September 20261

That sentence makes two promises. To see what they mean, here is one of our calls, as sent and as returned (token counts omitted):

// request
"state": {"question": "What number will come up on a single roll of a fair six-sided die?"},
"questions": {"answer": {"type": "choice",
                         "instructions": "Choose the correct answer to the question.",
                         "criteria": {"one": null, "two": null, "three": null,
                                      "four": null, "five": null, "six": null}}}

// Jev's response
"model": "jev-1.13.0",
"answers": {"answer": {"type": "choice", "choice": "one", "confidence": 0.73,
                       "probabilities": {"five": 0.01, "six": 0.09, "two": 0.01,
                                         "one": 0.77, "three": 0.06, "four": 0.06}}}

Jev picked "one" and gave it a probability of 0.77. The response also has a confidence score of 0.73, but that's not a separate judgement. It's the 0.77 rescaled so that spreading evenly over the six faces (1/6 each) would read 0 and certainty would read 1. The formula is roughly (6 × 0.77 − 1) / 5, which gives 0.73 up to rounding. Only the 0.77 can be checked against how often Jev is right, so in this post "confidence" means that number, the probability of the answer Jev picks.2

The first promise is calibration. Of all the answers Jev gives 77%, about 77% should be right. The second promise is about this answer alone. The die is fair, so each face has a one-in-six chance and nothing in the question favours "one". By saying 77%, Jev claims to know something it can't. This wasn't one unlucky call. We asked the same question with the six options in all 720 possible orders, and Jev picked "one" every time, with 0.80 on average. A model can keep the first promise and still break the second. If cases like this are rare, the overall numbers still look calibrated, yet each such case clears a gate set at 50%. The gate depends on the second promise.

We tested both claims in about 575,000 Jev calls.3 Four results stood out:

  • On familiar closed-choice tasks, confidence tracked accuracy well and supported useful filtering.
  • When the prompt contained no basis for an answer, or the question fell beyond the observed knowledge boundary, confidence often stayed high.
  • Asking "Do you know?" seemed at first to detect these gaps. Once we controlled for names that looked made up and for dates in the questions, it did no better than ordinary answer confidence.
  • Questions about the case itself held up better. Jev could tell whether the evidence it was given was enough, and whether an outcome was already settled.

1. We changed what Jev could know

Reading a probability on its own can't test the second promise. A 50% can mean the model is missing a fact, or that the outcome is a coin toss. The first is epistemic uncertainty, which more information can reduce. The second is aleatoric uncertainty, which no amount of information removes. So we built situations where we knew the answer. In each one we changed exactly one thing about what Jev was told and watched what its probabilities did. In other tests we stated the odds outright.

Diagram: two different situations, a fair coin toss and an unknown fact, both produce the same output, 50 percent. The output alone cannot say which kind of uncertainty it is.
Figure 2. The same probability can come from chance or from missing knowledge, and the output alone can't tell you which.

Our first test used customer-support messages from two public datasets, Banking77 and CLINC150. Take this one: "I ordered a card and I still haven't received it. It's been two weeks. What can I do?" The right category is "card arrival". We replaced every category name with a meaningless code, such as INTENT_67, so that without help nobody could tell which code was which. Then we gave Jev different amounts of help: nothing, a few examples per code, the real names, irrelevant text, or examples attached to the wrong codes. Click through the settings:

With one example per category, Jev gets it wrong. The only example for "getting virtual card" was "my virtual card has not came yet!", which sounds a lot like our customer. Two examples fix it, and with eight Jev is 96% sure and right. Across all 2,270 messages, more information meant higher accuracy and higher confidence:4

Now click No hints. For our question this is the setting that matters most, because here the right answer is unknowable. Jev was still 32–36% sure of its top pick on average. An even split across the options would be about 1%, and that's roughly how often it was right. Mostly it went for whichever code looked like a default, such as INTENT_00:

We made this test artificial on purpose, so there is no doubt the answer was unknowable. It shows that Jev can sound sure with nothing to go on.

How the test was set up, and two comparison models
Pipeline diagram of the first stage: 2,270 real test messages; category names hidden behind codes reshuffled per message; nine settings for how much help Jev gets; each asked three times, about 61,300 calls; confidence compared with accuracy in every setting. The same messages also go to a five-student ensemble and to an open rival model, and everything is compared message by message.
The main test on one page. Nine settings on all 2,270 messages (2,270 × 9 × 3 = 61,290 calls). The message and the right answer never change, so any change in Jev's probabilities has to come from what it was told. A tenth setting, the no-hints test with letter codes instead of numbers, was run separately.
The message (identical in every setting)
"I ordered a card and I still haven't received it. It's been two weeks. What can I do?"
Right answer card arrival
↓ the same 77 options, hidden behind codes, with different amounts of help ↓

No hints

INTENT_00 (nothing)INTENT_01 (nothing)… 75 more

1, 2, 4 or 8 examples each

INTENT_01 "my virtual card has not came yet!"INTENT_67 "What is the status of my card's delivery?"…

Category names

INTENT_01 getting virtual cardINTENT_67 card arrival…
↓ Jev returns a probability for every code ↓
What Jev saw. The two trick settings were irrelevant filler, as much text as eight examples but unrelated, and swapped examples, real examples from the wrong category.

Jev followed swapped examples on 90–98% of items. So the swap changed the mapping Jev used. Scoring those answers against the original labels is not, by itself, evidence of miscalibration.

We ran two comparison models on the same messages. An open model built for the same job, GLiNER2.5-Decide, stayed close to an even split with no information (1.7% on its top option) but was less accurate than Jev with category names. Five small models trained separately on the same examples (our "students") were less accurate and less confident. With one banking example per category, their confidence ran 10 points ahead of their accuracy, against 15 for Jev.

The open model is fastino/GLiNER2.5-Decide, run locally without training. The students are five fine-tuned copies of a 22M-parameter encoder that differ only in random seed. We use their disagreement as a rough measure of missing knowledge, a common heuristic rather than ground truth.

2. Where confidence worked

None of this makes Jev's confidence useless. When Jev had what it needed, as on standard multiple-choice benchmarks, its confidence tracked its accuracy closely:

Take trivia questions (TriviaQA, turned into four-option multiple choice). Keep only the answers Jev was at least 90% sure of, and you keep 82% of the questions, with 99.75% of those answers right.5 That's a strong result for this dataset, though it's no guarantee for other tasks.

It wasn't perfect, even here. On Banking77, where some categories are near twins, Jev was 83% confident but only 68% accurate with one example per category. It was also overconfident on MMLU-CF, a benchmark built to avoid questions a model may have seen in training.

More charts: confidence against accuracy for intent routing, and quiz clues revealed one at a time

Questions about real people, films and places gave the same result. From the famous to the obscure, confidence matched accuracy. Then we asked the same kinds of questions about 1,600 made-up subjects ("Who wrote The Velbri Tide?"), where none of the four options is right. Jev had nothing to go on, and it still put about half its probability on one option.

3. Later news, higher confidence

Made-up subjects are still a bit artificial. News gives a natural test, because a model's knowledge stops somewhere in time. We asked Jev yes/no questions about real events from January 2020 to July 2026, like "Will X happen by March 2025?". Its accuracy dropped sharply around late 2024. We call that point the observed knowledge boundary. We found it in the data, and it isn't a confirmed training cutoff.

Before the boundary, Jev was right 70% of the time and 75% confident on average. Beyond it, Jev was right only 51% of the time, no better than a coin flip, yet its confidence rose to 82%. It answered "no" to 92% of the later questions, even though about half of those events did happen. It behaved as if not having heard of something meant it hadn't happened.

The dates in the questions drove much of that "no". When we took them out, Jev said "no" much less often to later events that did happen (57% instead of 92%), while earlier questions barely changed. Removing a date can change what a question means, though, so we read this as a clue to how Jev responds, not as a measure of accuracy.6

How we checked the date edits, and their limits

An AI assistant assessed 240 original and edited question pairs against a written protocol. These labels are provisional and haven't yet been checked by a person.

About 35% kept their answer, 32% lost a deadline that could change a "no", and 33% no longer clearly identified the event. On the pairs judged to keep their answer, later "no" answers fell from 92% to 65% (37 pairs).

Can recalibration fix it?

The usual fix is recalibration, which learns a mapping from the model's confidence to how often it's right. It works only if the same confidence means the same thing on new cases, and here it didn't. At the same confidence, later questions were right 17–31 points less often than earlier ones. A mapping trained on earlier months still left later questions about 20 points overconfident.7 The number alone carries no sign of the shift.

4. What we learned so far

Jev's confidence is useful on familiar ground and unreliable beyond it, and correcting the number afterwards doesn't help. In Part 2 we ask Jev directly whether it knows, test what that question reacts to, and look at the questions that did hold up.

5. Methods, limitations and provenance

Scope and measurement
  • One model version. Requests used the jev-latest alias; every response reported jev-1.13.0. Other versions and model families may behave differently.
  • Scale. About 575,000 API calls (575,442 by the machine-generated ledger) across 15 public datasets and six generated task families. A call can carry several questions, so calls, questions and unique items are different counts.
  • Aggregation. Most items were sent three times; we averaged the returned distributions, then took the top option and its probability. With one call per item instead, estimated calibration error changed by less than about one point; that doesn't mean every individual decision would be the same.
  • Calibration. We report top-label calibration error (smooth ECE), confidence–accuracy gaps, reliability charts and proper scores. Low top-label error doesn't imply calibration per class, per subgroup or in deployment.
  • Intervals. Item-clustered bootstrap intervals, month clusters for dated news, and moving blocks for analyses over time.
  • Multiple choice is easier than open questions. Most free-text benchmarks were converted to four options with automatic distractors.
Construct validity and exploratory analyses
  • False premises. Made-up-subject questions have no correct option, so they have no calibrated 25%-per-option target; an even split is only a forced-choice reference.
  • Observed knowledge boundary. The boundary is a statistical change in accuracy, not an identified training cutoff; topics, style and difficulty can change over time.
  • Date edits. Removing dates can change a question or lose the event. The 240-pair assessment was provisional, done by an AI assistant, with independent human checking pending.
  • Stand-ins for knowledge. Made-up/real status and period are not direct labels of what the model knows.
  • Exploratory controls. The surface-cue controls, date edits, verification variants and recalibration bounds were added after the main studies. The findings over time have not been confirmed on a fresh set of news.
  • Decision value. The routing comparisons don't establish value for every mix of questions, cost of errors or kind of shift.
First stage: key numbers for Jev
Setting B77 acc. B77 conf. B77 cal. error CLINC acc. CLINC conf. CLINC cal. error
No hints 0.010 0.320 0.310 0.008 0.363 0.355
1 example 0.681 0.828 0.170 0.911 0.904 0.024
2 examples 0.796 0.879 0.099 0.958 0.953 0.019
4 examples 0.861 0.914 0.066 0.969 0.964 0.015
8 examples 0.899 0.927 0.037 0.979 0.974 0.016
Category names 0.821 0.893 0.084 0.926 0.933 0.025
Names + 8 examples 0.909 0.939 0.039 0.981 0.978 0.015
Irrelevant filler 0.019 0.312 0.293 0.007 0.226 0.219
Swapped examples 0.001 0.932 0.904 0.001 0.974 0.957

Accuracy, average confidence (top probability) and calibration error (smooth ECE: 0 means confidence matches accuracy at every level; it is not the same as confidence minus accuracy) per setting. 770 Banking77 and 1,500 CLINC150 messages per setting, each the average of three identical requests. The swapped-examples row is scored against the original labels; Jev followed the swapped examples, so it shows that the swap took effect rather than a calibration failure.

The code that ran and analysed every experiment is open source at Syntheme/beyond-answer-confidence. Per-item results are packaged separately, without any dataset text. The full study is in the paper, Beyond Answer Confidence: A Controlled Audit of Self-Knowledge in a Black-Box Decision Model.

Authorship and assistance. The authors are affiliated with Synthpop.AI. An AI coding assistant helped implement the experiments and draft text, and produced the provisional date-edit labels. The authors take responsibility for the study. Written critiques of earlier drafts were informal feedback, not formal peer review.

Disclosure pending author confirmation: any commercial or financial relationships with TypeSafe, TypeLLM or other relevant model providers.

  1. Quoted from TypeSafe's release post, "Introducing System One models and Jev" (accessed 29 September 2026), which also says Jev "always communicates confidence and uncertainty with every output. Calibrated: higher confidence means higher accuracy." Its documentation describes Choice probabilities as "the full probability distribution across every option" and System One probabilities as "optimized against outcomes to reflect uncertainty" (docs.typesafe.ai, accessed 30 September 2026). "Honest" here describes how the probabilities behave, not intent. ↩
  2. We sent this request on 1 October 2026 (request ID req_01a0f72fbc7f7526b82d55876fd05fc3). The open-source code rebuilds it exactly: it is unit device:die_words:0 of the stated_odds experiment, in which all 720 orders gave the 0.80 average. In general the confidence score is close to (K·p − 1)/(K − 1) for K options, so the same probability gives different scores for different numbers of options: 0.77 reads about 0.72 with six options but 0.54 with two. It carries the same information as the probabilities, and calibration needs the probability scale. ↩
  3. 575,442 paid API calls by our machine-generated ledger, about 1.15 billion input tokens on one model version (jev-1.13.0), an estimated $48 at the listed input price. Public benchmarks were turned into Jev's question types, usually four-option multiple choice with distractors from the same dataset. ↩
  4. The spread of Jev's probabilities (entropy divided by its maximum) fell by 0.66 (Banking77) and 0.54 (CLINC150) from no examples to one, then by only about 0.02 and 0.01 for each further doubling. Three different random draws of examples gave near-identical curves. ↩
  5. 8,132 of 9,960 questions kept, 20 wrong. A 90% threshold doesn't mean the kept answers are 90% right; most sit well above it. Calibration error (smooth ECE, 0 is perfect) was 0.016 on MMLU-Redux, 0.031 on TriviaQA and 0.010–0.026 for CLINC150 with 5 to 150 options. On MMLU-CF, confidence was 91% at 78% accuracy. ↩
  6. The boundary is the single change point in monthly accuracy, at November 2024; the drop and the rising confidence appear whichever month is used, and within every news category. The date test used 1,680 later and 1,680 earlier questions with every date phrase removed by rule ("Will the Fed cut rates by the end of March 2024?" became "Will the Fed cut rates?"). Telling Jev today's date, either the study date or a date just after each event, changed neither later-event accuracy (51%) nor the "no" answers (91–92%). When asked to check proposed answers, Jev favoured "it didn't happen" whether asked "is this answer correct?" or "is this answer wrong?". ↩
  7. Standard temperature and Platt scaling left the gap beyond the boundary at 24 points. Maps fitted on some months and tested on others left 15–18 points; maps fitted only on earlier months left 19–23 points. An exact bound shows no order-preserving map can match accuracy in both periods on this data. The recalibration appendix of the paper gives the full procedure. ↩