A boss rejects candidates over two confidence questions—the evidence stops there

In November 2024, an anonymous software manager was reported to use a two-part “confidence” test in every interview and reject candidates whose responses failed his standard. The available account still does not identify the employer, define what constitutes failure or provide evidence connecting the test with performance at work.
The distinction matters because confidence calibration is a legitimate subject of research, but a single confidence estimate is not automatically a validated hiring assessment. Current official guidance instead emphasizes job-related questions, consistent administration and predefined standards for evaluating every candidate.
What the two-question test actually asks
The exercise begins with a factual problem: the candidate may be asked to calculate 23 multiplied by 37 or spell “surveillance.” Only after an answer has been committed does the interviewer ask how confident the candidate is that it is correct.
The multiplication answer is 851, while “surveillance” is the correct spelling. Yet correctness was reportedly not the manager’s only concern. The interviewer watched whether candidates sought clarification, requested permission to use a calculator and expressed absolute or qualified certainty.
The November 2024 account of the test says the unnamed manager worked in software development and believed the combination of answer and confidence revealed something about problem-solving, reporting information and customer service. It also presents the practice as a rejection rule, but supplies no threshold separating a pass from a failure.
That omission leaves several materially different responses open to interpretation. A correct answer delivered with high confidence might indicate knowledge, careful checking or familiarity with the problem. The same confidence attached to a wrong answer might suggest poor error detection, but one mistake cannot establish that this is a stable professional trait.
Confidence is not the same as calibration
Calibration describes how closely expressed confidence corresponds to actual accuracy across judgments. Someone who assigns 80% confidence repeatedly would be well calibrated if roughly eight out of ten comparable answers proved correct. One answer followed by one percentage cannot establish that pattern.
A candidate saying “100%” is therefore not inherently better or worse than one saying “99.9%.” The first response may reflect ordinary conversational emphasis; the second may indicate careful qualification, false precision or an attempt to satisfy what the candidate thinks the interviewer wants. Without repeated comparable questions and an announced scale, the interviewer is reading meaning into language rather than measuring calibration.
The underlying concept nevertheless has empirical value in the right setting. A longitudinal PLOS ONE study of mental-arithmetic judgments defined calibration as alignment between confidence and accuracy and found that better calibration among children in Grade 5 predicted larger accuracy gains through Grade 8 after initial performance was controlled. That research involved repeated judgments by schoolchildren, not employment interviews, so it cannot validate a one-item hiring decision.
Why a pass-or-fail decision is difficult to defend
The test combines at least three possible constructs: factual skill, willingness to use tools and awareness of uncertainty. Those are not interchangeable. Asking to use a calculator may be efficient in a workplace where verification matters, while avoiding one may be relevant when mental calculation is an explicit requirement.
The alternative spelling prompt introduces another comparability problem. Multiplication and spelling draw on different knowledge, and candidates receiving different tasks have not faced the same assessment. An unfamiliar word may also measure language experience more than the professional judgment the manager says he wants to observe.
A question about confidence could still support a broader interview when the employer first identifies a job-related competency. For example, a software role that requires communicating uncertain estimates might reasonably assess whether a candidate states assumptions, checks work and explains what evidence would change an answer. The scoring standard should reward those observable behaviors rather than a preferred confidence percentage.
The US Office of Personnel Management’s current structured-interview guidance says candidates should receive the same predetermined questions in the same order, with responses judged against the same rating scale and standards for acceptable answers. The viral test satisfies only part of that model by repeating a sequence; the published account provides neither job-linked criteria nor a common scoring rubric.
How the idea could become a fairer assessment
An employer interested in intellectual honesty does not need to discard the confidence question. It can be redesigned as one component of a structured work sample, where the candidate’s reasoning and response to new information are visible.
- Choose a short problem that resembles a real task in the advertised role.
- Give every candidate the same instructions, tools and time limit.
- Ask for an answer, a confidence estimate and the assumptions behind both.
- Provide a relevant new fact, then observe whether the candidate updates the answer appropriately.
- Score predefined behaviors such as verification, explanation and correction instead of treating one percentage as a personality verdict.
This version tests something closer to professional judgment: not whether a person projects certainty on demand, but whether they can distinguish knowledge from assumption and revise a conclusion when evidence changes. It also produces behavior that interviewers can compare instead of relying on intuition about what “99.9%” supposedly means.
What candidates should take from the question
A candidate encountering this exercise cannot know the interviewer’s hidden preference. The most defensible response is to solve the task, state confidence in ordinary language and briefly explain the basis for it. If tool use or assumptions affect the answer, saying so makes the reasoning inspectable.
That approach does not guarantee success with a manager who has an undisclosed pass-or-fail rule. It does, however, demonstrate the qualities the test ostensibly seeks: clear reporting, proportionate certainty and willingness to verify. The enduring lesson is narrower than the viral claim—confidence can be useful evidence when paired with accuracy and consistent scoring, but one confidence answer does not by itself establish character, professionalism or future job performance.
Also read:
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.