Skip to main content
eScholarship
Open Access Publications from the University of California

UCLA

UCLA Previously Published Works bannerUCLA

How far can a commercial decision model's probabilities be trusted? A calibration audit of Jev against open and general-purpose classifiers

Creative Commons 'BY' version 4.0 license
Abstract

Commercial "decision models" return a choice or a probability for a closed question about a passage of text, and they are marketed as calibrated judgments that software can act on directly. We audit one such system, TypeSafe's Jev 1.13.0, with 3,244 evaluations of 2,372 distinct texts from two public classification tasks, and compare it with two open zero-shot classifiers (one of them trained with examples from these tasks) and two small general-purpose language models. The top-label calibration of Jev's probabilities differs between tasks and question forms, so it has to be checked for each: the calibration error changes by 0.071 between two ways of posing the same question, the middle of the four-class scale is overconfident (stated 0.750, observed 0.612), and 71.6% of four-class answers sit at exactly 1.00, which leaves a threshold on the top-class probability 53 operating points. One fitted temperature removes most of the miscalibration where it exists, and split conformal prediction reaches its marginal target, although its single-label answers are wrong 5.93% of the time at a nominal 5%. Against an open checkpoint trained without examples from these benchmarks, Jev is more accurate in all 16 combinations of label wording and decision rule on both tasks, with an exploratory full-test-set estimate of +5.72 points (95% interval [+4.70, +6.75]) on AG News. On AG News it is less accurate than a checkpoint trained with examples from these tasks (-2.30 points [-3.33, -1.28]), and on sentiment the two are not distinguishable. The benchmarker's choices matter as much: the label descriptions alone moved the gap from +6.93 to +9.67 points, and the choice of open checkpoint moved it by ~8 points and reversed its sign. Returned scores should therefore be validated for the target task, question form and decision rule, and recalibrated where that validation shows it is needed.

Many UC-authored scholarly publications are freely available on this site because of the UC's open access policies. Let us know how this access is important for you.