Jev benchmark on 1,759 decisions
A launch-week field report comparing Jev with a frontier Gemini model across multilingual labeled cases and anonymized routing replays, with latency, cost, calibration and failure examples.
- Stars
- 0
- Forks
- 0
- License
- —
- Last commit
- —
What Jev does here
Evaluates whether confidence gating can separate safe automated routing from ambiguous cases, and documents failures involving dialect distinctions, missing ownership context, misleading prefills and one high-confidence misclassification.
Adds a concrete independent comparison and operational error analysis to the ecosystem, helping teams design evaluation sets and fallback thresholds instead of relying only on vendor headline numbers.
Self-reported Reddit study reviewed on 2026-09-24. The post describes 1,759 decisions and manual disagreement review, but this directory has not received the raw dataset or independently reproduced its reported accuracy, latency, cost or calibration results.