OPEN-SOURCE PROJECT

jev-rerank-bench

Compares Jev, dedicated rerankers and chat models on the same retrieved passages.

Stars
4
Forks
0
License
MIT
Last commit
2026-09-17

What Jev does here

Ranks candidate passages with Choice, Noul and rubric scores, then computes retrieval metrics.

Publishes raw responses, scoring code and per-dataset results for inspection.

The author reports equal-dataset nDCG@10 of 0.692 for Jev and 0.691 for Cohere on eight English datasets, without establishing a winner. Query weighting changes the comparison. README and integration source reviewed at a fixed commit; not independently run or benchmarked by this site.