OPEN-SOURCE PROJECT

jev-calibration-audit

Audits Jev calibration, option-wording effects and Korean judgments through public APIs and datasets.

Stars
0
Forks
0
License
MIT
Last commit
2026-09-18

What Jev does here

Collects Noul and Choice outputs and compares them with labels for error, accuracy and stability.

Keeps per-call records and experiment notes to qualify conclusions.

Calibration varies with tasks, labels and answer options. These experiments do not guarantee reliable confidence in every domain. README and integration source reviewed at a fixed commit; not independently run or benchmarked by this site.