OPEN-SOURCE PROJECT

jev-eval

Compares Jev and OpenRouter models on labeled tasks for accuracy, calibration, latency and cost.

Stars
1
Forks
0
License
Last commit
2026-09-19

What Jev does here

Collects judgments on matched tasks and computes confidence intervals and repeat-input stability.

Publishes data processing, runner and statistics code with model-specific results.

The published comparison uses 300 items per task with specified models and dates. Advantages vary by task and do not demonstrate universal superiority or absence of bias. README and integration source reviewed at a fixed commit; not independently run or benchmarked by this site.