EXTERNAL RUN · SITE RECALCULATION · 2026

Jev calibration and stability evidence: 1,000 outputs recalculated

A reproducible recalculation of two 500-item Jev calibration runs and a bounded review of three repeated calls on two fixed inputs.

SITE RECALCULATED EXTERNAL RUN 2026-09-20 jev-1.13.0 · typesafe-sdk 0.7.0
1,000item-level calibration outputs
94.0%SST-2 accuracy
0.0934SST-2 ECE
0.0866AG News ECE
00 / ANSWER

What does this evidence establish?

On the author’s two public datasets, one prompt wording per task, and one hosted-model run, Jev was accurate but not perfectly calibrated: SST-2 probabilities were generally conservative, while the highest AG News probability bin was overconfident. This site’s recalculation matched every published aggregate.

The results do not establish the same accuracy or ECE for another business distribution, wording, model version, or access path, and they do not produce a production threshold.

01 / PROVENANCE

Pinned provenance and integrity

Author
Colin McNamara
Run context
2026-09-20 · Austin, TX
Model / SDK
jev-1.13.0 · typesafe-sdk 0.7.0
Pinned commit
e32717b1c592 ↗
t3_raw.json SHA-256
5a46d61b1074d350ed5c915ff36d0ceb0692fe1cf1f904494cc342f0af83ea1f
t1_results.json SHA-256
cac6a52f33c74dfd1c7c6f8bb21cba5cacb057b4bd029b13fedb43206af46a40

Run the downloaded script. It fetches only two JSON files from the pinned commit, verifies SHA-256 before parsing, and prints the recalculation. No Jev API key is required.

node recompute.mjs
02 / CALIBRATION

Recalculated metrics

TaskPrimitivenAccuracyECEBrier
SST-2 validationNoul5000.94000.09340.0491binary
AG News testChoice5000.88200.08660.1013top-label only

ECE uses 10 equal-width bins and weights |mean predicted probability − observed frequency| by bin count. Probability 1.0 is included in the final bin.

Brier is binary mean squared error on the positive-class probability for SST-2. For AG News it is top-label probability versus top-label correctness—not a full four-class Brier score.

03 / BINS

Reliability-bin tables

SST-2 validation · Noul

Probability binnMean predictedObserved
0.0-0.11640.040.012
0.1-0.2460.1460.087
0.2-0.3190.2430.105
0.3-0.4130.3480.538
0.4-0.5140.4590.786
0.5-0.640.5471
0.6-0.7250.6440.92
0.7-0.8340.7550.941
0.8-0.9540.8491
0.9-1.01270.9491

AG News test · Choice

Probability binnMean predictedObserved
0.0-0.10——
0.1-0.20——
0.2-0.30——
0.3-0.40——
0.4-0.510.470
0.5-0.6150.5430.733
0.6-0.7150.6480.667
0.7-0.8170.7410.765
0.8-0.9250.8570.48
0.9-1.04270.9950.925

Tinted rows have n < 20 and should not be interpreted alone. The AG News 0.9–1.0 bin contains 427/500 items, with mean top-label probability 0.995 and observed accuracy 0.925.

04 / STABILITY

Repeated calls: ranges, not a stability rate

The upstream author also called two fixed inputs three times each. Values were not identical, but the sample is far too small to estimate drift probability, tail behavior, or cross-version stability.

TICKET A · NOUL0.21–0.220.21 · 0.22 · 0.21
TICKET A · CHOICE YES0.00–0.010.00 · 0.01 · 0.00
TICKET B · REFUND0.72–0.730.72 · 0.73 · 0.72
TICKET B · NOT REFUND0.41–0.470.44 · 0.47 · 0.41

A production stability test should pin state, questions, criteria, model version, and access path; run enough repeats; and record full distributions, errors, retries, and time. Never retry until a high-confidence answer appears.

05 / LIMITS

What this report does not establish

  1. 01The site recalculated published outputs; it did not reproduce the model calls.
  2. 02One author-written question wording was used per task.
  3. 03SST-2 and AG News are famous benchmarks and may appear in training data.
  4. 04Several middle reliability bins contain fewer than 20 items.
  5. 05The three repeated calls per example do not establish a stability rate or guarantee.
  6. 06Results do not determine a production threshold for another workflow or data distribution.