OPEN-SOURCE PROJECT
jev-eval
Compares Jev and OpenRouter models on labeled tasks for accuracy, calibration, latency and cost.
- Stars
- 1
- Forks
- 0
- License
- —
- Last commit
- 2026-09-19
What Jev does here
Collects judgments on matched tasks and computes confidence intervals and repeat-input stability.
Publishes data processing, runner and statistics code with model-specific results.
The published comparison uses 300 items per task with specified models and dates. Advantages vary by task and do not demonstrate universal superiority or absence of bias. README and integration source reviewed at a fixed commit; not independently run or benchmarked by this site.