OPEN-SOURCE PROJECT

typesafe-ai-benchmark

Compares Jev with other structured-output models on shared application tasks, recording errors, latency, Tokens, and estimated cost.

Stars
34
Forks
5
License
MIT
Last commit
2026-09-19

What Jev does here

Maps the same tasks to Jev Choice/Noul questions and normalizes answers to a shared result format.

Preserves comparison methods and results for inspecting model differences.

The current README describes limited synthetic workloads and separately captured results. The demo video is not measurement data, and results are not a ranking for all tasks. This site has not reproduced them.