Article

Does a decision model (like Jev) know when it is guessing? (Post 2)

Sharath Shankaranarayana
October 1, 2026

Decision models like Jev, from TypeSafe AI, pick from a fixed set of options and return a probability for each one. Software uses that probability as a gate: act automatically when the model is sure, and send the case to a person when it isn't.

In Part 1 we checked whether Jev's confidence drops when it doesn't know the answer. On familiar tasks it tracked accuracy well. With nothing to go on, or with news beyond its observed knowledge boundary, it stayed high, and correcting the number afterwards didn't fix that.

So in this part we asked Jev questions instead. Here is one of our calls. There is no book called The Velbri Tide, so none of the four options is right. Alongside the question we asked Jev whether it knew the answer:

// request
"state": {"question": "Who is the author of The Velbri Tide?"},
"questions": {
  "answer": {"type": "choice",
             "instructions": "Choose the correct answer to the question.",
             "criteria": {"A": "Nicholas Negroponte", "B": "Steve Gerber",
                          "C": "Erich Segal", "D": "Ruthann Friedman"}},
  "known": {"type": "noul",
            "instructions": "Do you know the answer to this question
                             for certain, rather than having to guess?"}
}

// Jev's response
"model": "jev-1.13.0",
"answers": {
  "answer": {"type": "choice", "choice": "D", "confidence": 0.58,
             "probabilities": {"B": 0.24, "C": 0.05, "A": 0.03, "D": 0.68}},
  "known": {"type": "noul", "noul": 0.23}
}

Jev put 0.68 on Ruthann Friedman as the author of a book that doesn't exist, enough to pass a gate set at 50%. The second question is a yes/no question, and noul is Jev's probability of "yes". At 0.23, Jev mostly said it didn't know. The answer sounds sure, but the follow-up seems to know better.1 That looks like a model that knows when it is guessing. This part tests whether it is.

Across the whole study, about 575,000 Jev calls,2 two results from this part stood out:

  • Asking "Do you know?" seemed at first to detect these gaps. Once we controlled for names that looked made up and for dates in the questions, it did no better than ordinary answer confidence.
  • Questions about the case itself held up better. Jev could tell whether the evidence it was given was enough, and whether an outcome was already settled.

1. What was "Do you know?" picking up?

If the number can't flag these gaps, maybe Jev can tell us directly. So alongside each question we also asked, "Do you know the answer for certain?" It said yes less often for made-up subjects than for real ones, and less often for later news than for earlier news. On both comparisons it scored about 0.91, on a scale where 1.0 means it tells the two groups apart perfectly and 0.5 means it's guessing.3

That looked like a model that knows what it doesn't know. But there was a simpler explanation to rule out first. Jev might be reacting to how the questions looked rather than to what it knew. Our made-up names looked made up, and the news questions carried their dates. A simple text classifier that never calls Jev spotted the made-up names perfectly, and the dates alone picked out the later news almost perfectly. So we ran three more tests:

  • Realistic made-up names. We built fake people from real first names and surnames. The text classifier was now at chance, and "Do you know?" dropped to 0.74, about the same as Jev's ordinary answer confidence (0.76).
  • No dates. With the dates removed from the news questions, "Do you know?" fell to 0.58, again about the same as answer confidence.
  • The answer in the text. When we added one sentence giving the answer about a made-up subject, Jev's "yes" rose from 20% to 81%.

So "Do you know?" mostly reacts to surface clues and to what's written in front of it. Once we controlled the clues we could measure, it told us nothing Jev's answer confidence didn't already tell us. That doesn't prove Jev can't tell what it doesn't know. It does mean this question isn't a reliable way to find out.

Spotting the signs of uncertainty is not the same as knowing whether you have the answer.

More charts: how every follow-up question scored before these tests

2. Asking about the case worked better

"Do you know?" asks Jev about its own memory. We also tried two follow-up questions about the case in front of it, meaning the question and any text that comes with it. Jev can answer those by reading the input. Both worked.

"Is there enough information?"

Take a question that needs two facts, such as which of two people was born first. We gave Jev the question with two paragraphs. Sometimes the paragraphs held both birth dates. Sometimes they were about related people and didn't. Then we asked, "Do the paragraphs given contain enough information to answer the question for certain?"

Jev's "yes" told the complete paragraphs from the incomplete ones with a score of 0.95, on the same scale as before (1.0 is perfect, 0.5 is guessing). Answer confidence scored 0.85 on the same questions. Longer text didn't explain the difference.4

This question tells you whether the documents are complete, not whether the answer is right. Jev often knew the answer without the paragraphs, so a "no" often came with a correct answer. To pick which answers to trust, answer confidence was still the better filter.

"Is the outcome already settled?"

Here is one of our made-up events. Jev saw "Who is the winner of the Roskal Cup final?" with four made-up teams, Ornwen, Athdra, Nesmal and Jorkal. One version added "The Roskal Cup final was played last spring." The other said it "will be played next spring." We asked whether the answer was already fixed, so that someone with more information could know it for certain.

For last spring's final, Jev's "yes" was 0.75. For next spring's, it was 0.10. Across 360 events like this one, the question told past from future perfectly when only the tense differed, and almost perfectly (0.996) from dates when we also gave today's date.5

The main question tells a different story. Jev picked Ornwen both times, with 0.63 for last spring's final and 0.76 for next spring's. Nothing in the input favours any team, so 0.25 each is the honest answer in both versions. Jev was most confident about the match that hasn't been played yet. It can tell a missing fact from a random event when asked, but its answer probabilities don't reflect the difference.

Both questions are useful checks on the input. A workflow can send a case to a person when the documents are incomplete or the outcome is still open. Neither question shows that Jev knows what it doesn't know.

3. A fair die, and 80% on "one"

The strangest result came from chance, the other source of uncertainty. Another team found that if you ask Jev about a roll of a fair die, it puts far more than one-sixth of its probability on "one".6 We checked it across every ordering of the six options, and with digits, colours and card suits:

Jev put about 80% on "one", whatever order the options came in. Given stated odds of 40%, 30%, 20% and 10%, it put 98% on the 40% option. Rewording didn't help either. "Which number will come up?", "which is most likely?" and "which would you bet on?" all got the same answer. TypeSafe describes these probabilities as "optimized against outcomes", so it's natural to read them as the chance of each outcome. For random events, they aren't. What did work was asking one yes/no question per outcome:

Asked one outcome at a time, Jev's answers were within about three points of the stated odds on average.7

4. What helped, and what it cost

Some changes to how you ask made Jev's outputs more useful. None of them was free.

Give it a way to say "I don't know"

Jev does notice missing information when it can see the gap. In 600 policy cases with a key fact deleted, it answered "cannot tell" every time.8 A yes/no question doesn't offer that way out, so for the news questions we added a third answer, "not known to me":

It worked, mostly by refusing to answer. Confident mistakes almost disappeared, but Jev declined 93% of the later questions and 36% of the earlier ones. Whether that's worth it depends on your mix of questions. When half came from beyond the boundary, the option beat a simple confidence threshold by five points at the same answer rate. When very few did, the threshold did a little better.9 A plain rule that skipped questions with recent dates did almost as well. And adding today's date to the prompt made things worse.

Tell it the documents may be irrelevant

Give Jev two related-looking paragraphs that don't contain the answer, and its accuracy falls from 84% to 75% while its confidence stays put. One sentence in the instructions brought accuracy back to 84% on this benchmark: "The paragraphs may be irrelevant or incomplete. Use them only if they actually answer the question."10

Describe your options

The cheapest fix comes from our first test, the one with coded categories. Give the categories real names and a few examples. That doesn't teach Jev to recognise what it doesn't know. It gives Jev what it needs to know. With eight examples per category, sending the least confident 10% of messages to a person halved the errors on the rest, from 10.1% to 4.8%.

More charts: Jev's most common mistakes with full information

Use a second check sparingly

Asking Jev in a second call whether its first answer was correct shrank the overconfidence on later news from 30 points to 6. But it also made Jev doubt easy answers that were already right. Use it where you expect trouble, not everywhere.

More charts: calibration before and after the second check

5. A production example

Here is what we would build for a refund decision where one deciding fact is missing:

// one Jev call: the decision, with a way out
{
  "state": {
    "policy": "Refunds are given within 30 days of delivery if the item is unused.",
    "facts": ["Order 5812 was delivered on 2 September 2026.",
              "The customer asked for a refund on 20 September 2026."]
  },
  "questions": {
    "answer": {
      "type": "choice",
      "instructions": "Is order 5812 eligible for a refund under the policy? Use only the policy and the facts given.",
      "criteria": {
        "refund": "eligible for a refund",
        "no_refund": "not eligible for a refund",
        "cannot_tell": "cannot tell from the information given"
      }
    }
  }
}

The facts don't say whether the item is unused, so "cannot tell" is the right answer. The routing rule then only needs Jev's answer and its probability:

// illustrative, not an evaluated policy
// p_answer is the probability of the chosen option, not the API's separate confidence field
p_answer = probabilities[choice]
if choice == "cannot_tell":             send to a person
elif p_answer >= validated_threshold:   act automatically
else:                                   send to a person

Pick the threshold on real examples from your own workflow, with the cost of mistakes in mind, and test the whole policy together. Given what we found, we wouldn't add "Do you know?" as a gate.

6. What we learned

So, does Jev know when it's guessing? On familiar tasks, its confidence tracks its accuracy and is useful. When information is missing, it isn't a reliable warning, and recalibrating doesn't fix that. Asking about the case, such as whether the evidence is complete or the outcome settled, works well. Asking the model whether it knows doesn't, at least not in any way we could separate from surface clues.

If you use a decision model, test each signal on your own cases before you rely on it.

7. Methods, limitations and provenance

Scope and measurement
  • One model version. Requests used the jev-latest alias; every response reported jev-1.13.0. Other versions and model families may behave differently.
  • Scale. About 575,000 API calls (575,442 by the machine-generated ledger) across 15 public datasets and six generated task families. A call can carry several questions, so calls, questions and unique items are different counts.
  • Aggregation. Most items were sent three times; we averaged the returned distributions, then took the top option and its probability. With one call per item instead, estimated calibration error changed by less than about one point; that doesn't mean every individual decision would be the same.
  • Calibration. We report top-label calibration error (smooth ECE), confidence–accuracy gaps, reliability charts and proper scores. Low top-label error doesn't imply calibration per class, per subgroup or in deployment.
  • Intervals. Item-clustered bootstrap intervals, month clusters for dated news, and moving blocks for analyses over time.
  • Multiple choice is easier than open questions. Most free-text benchmarks were converted to four options with automatic distractors.
Construct validity and exploratory analyses
  • False premises. Made-up-subject questions have no correct option, so they have no calibrated 25%-per-option target; an even split is only a forced-choice reference.
  • Observed knowledge boundary. The boundary is a statistical change in accuracy, not an identified training cutoff; topics, style and difficulty can change over time.
  • Date edits. Removing dates can change a question or lose the event. The 240-pair assessment was provisional, done by an AI assistant, with independent human checking pending.
  • Stand-ins for knowledge. Made-up/real status and period are not direct labels of what the model knows.
  • Exploratory controls. The surface-cue controls, date edits, verification variants and recalibration bounds were added after the main studies. The findings over time have not been confirmed on a fresh set of news.
  • Decision value. The routing comparisons don't establish value for every mix of questions, cost of errors or kind of shift.
First stage: key numbers for Jev
Setting B77 acc. B77 conf. B77 cal. error CLINC acc. CLINC conf. CLINC cal. error
No hints 0.010 0.320 0.310 0.008 0.363 0.355
1 example 0.681 0.828 0.170 0.911 0.904 0.024
2 examples 0.796 0.879 0.099 0.958 0.953 0.019
4 examples 0.861 0.914 0.066 0.969 0.964 0.015
8 examples 0.899 0.927 0.037 0.979 0.974 0.016
Category names 0.821 0.893 0.084 0.926 0.933 0.025
Names + 8 examples 0.909 0.939 0.039 0.981 0.978 0.015
Irrelevant filler 0.019 0.312 0.293 0.007 0.226 0.219
Swapped examples 0.001 0.932 0.904 0.001 0.974 0.957

Accuracy, average confidence (top probability) and calibration error (smooth ECE: 0 means confidence matches accuracy at every level; it is not the same as confidence minus accuracy) per setting. 770 Banking77 and 1,500 CLINC150 messages per setting, each the average of three identical requests. The swapped-examples row is scored against the original labels; Jev followed the swapped examples, so it shows that the swap took effect rather than a calibration failure.

The code that ran and analysed every experiment is open source at Syntheme/beyond-answer-confidence. Per-item results are packaged separately, without any dataset text. The full study is in the paper, Beyond Answer Confidence: A Controlled Audit of Self-Knowledge in a Black-Box Decision Model.

Authorship and assistance. The authors are affiliated with Synthpop.AI. An AI coding assistant helped implement the experiments and draft text, and produced the provisional date-edit labels. The authors take responsibility for the study. Written critiques of earlier drafts were informal feedback, not formal peer review.

Disclosure pending author confirmation: any commercial or financial relationships with TypeSafe, TypeLLM or other relevant model providers.

  1. We sent this request on 1 October 2026 (request ID req_01a0f75af8c87bf8aa9b3106fdb0f195). It is unit fab:author:7 of the open-source knowledge_boundary experiment, which sent it three times: 0.68 to 0.72 on "D", and 0.22 to 0.23 for "yes". As in Part 1, we use the probability of the chosen option (0.68), not the confidence score (0.58). ↩
  2. 575,442 paid API calls by our machine-generated ledger, about 1.15 billion input tokens on one model version (jev-1.13.0), an estimated $48 at the listed input price. Public benchmarks were turned into Jev's question types, usually four-option multiple choice with distractors from the same dataset. ↩
  3. The score is the AUROC: the chance that, for a random pair with one case from each group, the question ranks them the right way round (ties count half). It measures how well two groups are told apart, not the chance that a particular answer is wrong, and we did not test whether an 80% "yes" means an 80% chance that Jev knows. The text classifier looks only at letter patterns. The realistic names join a real first name to a real surname from two different obscure people (600 names). On the same questions, "Do you know?" minus answer confidence was −0.02 (95% interval −0.05 to +0.01) for realistic names and +0.02 (−0.01 to +0.06) without dates. These controls were added after the first results. ↩
  4. 951 HotpotQA comparison questions; every condition with paragraphs had exactly two. Within 10-word length bands, the question still scored 0.95 while length alone was at chance; its advantage over answer confidence was 0.10 (95% interval 0.09 to 0.11). As a filter on which answers to trust, it was 2–6 points less accurate than answer confidence. ↩
  5. 360 made-up events, each in a past and a future version that differ only in tense or date: perfect separation from tense, 0.996 from dates when today's date was given. In the generated worlds, even where the right answer was an even split in both cases, Jev was 76% sure on the future draw and 49% on the unknown fact. The Roskal Cup numbers are averages of three calls each, units tense:sports:0:past and tense:sports:0:future of the open-source inferred_settledness experiment. ↩
  6. TypeLLM's post "Can Jev roll a die?" reported 83–86% on "one" across all 720 orders. We measured 80% with the words one to six, 74% with digits and 62% for "red" with coloured faces. Asking for "the probability of each" in the question helped only partly (66% on a six that comes up half the time). ↩
  7. 480 generated draws with the chances stated: about 3 points off on average, and about 1.2 after rescaling each draw's answers to sum to 100%. This tests reporting of given odds, not forecasting of real events. ↩
  8. Inspired by the contrastive curation of Bespoke Labs' Nimble dataset; "cannot tell" received 91% of the probability on average. Jev followed a single changed fact 99.7% of the time, but flagged contradictions only 42% of the time. ↩
  9. 1,680 later and 1,680 earlier news questions, reweighted to any share of later questions and compared at exactly the same answer rate: +5.1 points at a 50% share, +1.0 at 27% (interval −0.2 to +4.3), −3.4 at 5%. Beyond the boundary only 6.6% of questions were still answered, 63% of them correctly. With today's date added, 49% of the answered later questions were right. ↩
  10. 951 HotpotQA comparison questions, each asked with no paragraphs and with exactly two. The unhelpful paragraphs are the dataset's own distractors, chosen because they look related. ↩