COMMUNITY PROJECT

Jev benchmark on 1,759 decisions

A launch-week field report comparing Jev with a frontier Gemini model across multilingual labeled cases and anonymized routing replays, with latency, cost, calibration and failure examples.

Stars
0
Forks
0
License
—
Last commit
—

What Jev does here

Evaluates whether confidence gating can separate safe automated routing from ambiguous cases, and documents failures involving dialect distinctions, missing ownership context, misleading prefills and one high-confidence misclassification.

Adds a concrete independent comparison and operational error analysis to the ecosystem, helping teams design evaluation sets and fallback thresholds instead of relying only on vendor headline numbers.

Self-reported Reddit study reviewed on 2026-09-24. The post describes 1,759 decisions and manual disagreement review, but this directory has not received the raw dataset or independently reproduced its reported accuracy, latency, cost or calibration results.