Jev calibration and stability evidence: 1,000 outputs recalculated
A reproducible recalculation of two 500-item Jev calibration runs and a bounded review of three repeated calls on two fixed inputs.
What does this evidence establish?
On the author’s two public datasets, one prompt wording per task, and one hosted-model run, Jev was accurate but not perfectly calibrated: SST-2 probabilities were generally conservative, while the highest AG News probability bin was overconfident. This site’s recalculation matched every published aggregate.
The results do not establish the same accuracy or ECE for another business distribution, wording, model version, or access path, and they do not produce a production threshold.
Pinned provenance and integrity
- Author
- Colin McNamara
- Run context
- 2026-09-20 · Austin, TX
- Model / SDK
jev-1.13.0·typesafe-sdk 0.7.0- Pinned commit
e32717b1c592↗- t3_raw.json SHA-256
5a46d61b1074d350ed5c915ff36d0ceb0692fe1cf1f904494cc342f0af83ea1f- t1_results.json SHA-256
cac6a52f33c74dfd1c7c6f8bb21cba5cacb057b4bd029b13fedb43206af46a40
Run the downloaded script. It fetches only two JSON files from the pinned commit, verifies SHA-256 before parsing, and prints the recalculation. No Jev API key is required.
node recompute.mjs Recalculated metrics
| Task | Primitive | n | Accuracy | ECE | Brier |
|---|---|---|---|---|---|
| SST-2 validation | Noul | 500 | 0.9400 | 0.0934 | 0.0491binary |
| AG News test | Choice | 500 | 0.8820 | 0.0866 | 0.1013top-label only |
ECE uses 10 equal-width bins and weights |mean predicted probability − observed frequency| by bin count. Probability 1.0 is included in the final bin.
Brier is binary mean squared error on the positive-class probability for SST-2. For AG News it is top-label probability versus top-label correctness—not a full four-class Brier score.
Reliability-bin tables
SST-2 validation · Noul
| Probability bin | n | Mean predicted | Observed |
|---|---|---|---|
| 0.0-0.1 | 164 | 0.04 | 0.012 |
| 0.1-0.2 | 46 | 0.146 | 0.087 |
| 0.2-0.3 | 19 | 0.243 | 0.105 |
| 0.3-0.4 | 13 | 0.348 | 0.538 |
| 0.4-0.5 | 14 | 0.459 | 0.786 |
| 0.5-0.6 | 4 | 0.547 | 1 |
| 0.6-0.7 | 25 | 0.644 | 0.92 |
| 0.7-0.8 | 34 | 0.755 | 0.941 |
| 0.8-0.9 | 54 | 0.849 | 1 |
| 0.9-1.0 | 127 | 0.949 | 1 |
AG News test · Choice
| Probability bin | n | Mean predicted | Observed |
|---|---|---|---|
| 0.0-0.1 | 0 | — | — |
| 0.1-0.2 | 0 | — | — |
| 0.2-0.3 | 0 | — | — |
| 0.3-0.4 | 0 | — | — |
| 0.4-0.5 | 1 | 0.47 | 0 |
| 0.5-0.6 | 15 | 0.543 | 0.733 |
| 0.6-0.7 | 15 | 0.648 | 0.667 |
| 0.7-0.8 | 17 | 0.741 | 0.765 |
| 0.8-0.9 | 25 | 0.857 | 0.48 |
| 0.9-1.0 | 427 | 0.995 | 0.925 |
Tinted rows have n < 20 and should not be interpreted alone. The AG News 0.9–1.0 bin contains 427/500 items, with mean top-label probability 0.995 and observed accuracy 0.925.
Repeated calls: ranges, not a stability rate
The upstream author also called two fixed inputs three times each. Values were not identical, but the sample is far too small to estimate drift probability, tail behavior, or cross-version stability.
0.21 · 0.22 · 0.210.00 · 0.01 · 0.000.72 · 0.73 · 0.720.44 · 0.47 · 0.41A production stability test should pin state, questions, criteria, model version, and access path; run enough repeats; and record full distributions, errors, retries, and time. Never retry until a high-confidence answer appears.
What this report does not establish
- 01The site recalculated published outputs; it did not reproduce the model calls.
- 02One author-written question wording was used per task.
- 03SST-2 and AG News are famous benchmarks and may appear in training data.
- 04Several middle reliability bins contain fewer than 20 items.
- 05The three repeated calls per example do not establish a stability rate or guarantee.
- 06Results do not determine a production threshold for another workflow or data distribution.