OPEN-SOURCE PROJECT

jev-agent-failure-benchmark

A benchmark using Jev to attribute multi-Agent failures to an Agent, step and error type.

Stars
1
Forks
0
License
Apache-2.0
Last commit
2026-09-17

What Jev does here

Builds candidate sets from traces and submits three choice questions.

Provides evaluation scripts and author results; some baselines generate answers while Jev selects candidates.

Some Jev benchmark axes use constrained choices while paper baselines generate freely; not every metric is a like-for-like comparison.