PRODUCTION CASE STUDY · JEV × SEO

The 0.89-confidence suggestion was wrong

Loomaly uses Jev for per-page audits, internal-link suggestions, and conversion checks on real websites. Speed and cost improved sharply, but four sets of rules outside the model made the results usable.

Published 2026-09-26About 12 minutesLoomaly production notes
00 / BASELINE

What changed after moving to Jev

BEFORE · GENERATIVE AUDIT 53–170s

per page; on one day, 163 of 190 audits timed out.

  • Repeated runs produced different issue lists
  • It reported issues unsupported by the supplied evidence
  • Cost limited each run to 25 pages
AFTER · JEV DECISION LAYER 1–2s

per page at about $0.00003, making full-page coverage practical.

  • 150 of 153 decisions matched across two runs
  • A Jev run over the Pop Mart site cost about one cent
  • Internal-link acceptance rose from 14% to 80%

These figures come from Loomaly design documents, calibration logs, and production records from September 2026. They are not universal performance claims for Jev.

01
CONFIDENCE ≠ CORRECTNESS

High confidence is not correctness

Loomaly asked Jev whether a sentence should receive an internal link, then joined the answers with real accept and reject outcomes. Mean confidence barely separated the two groups.

ACCEPTED0.619mean confidence
REJECTED0.623mean confidence
HIGHEST CONFIDENCE0.89but the suggestion was wrong
FAILED CASE
“What are the alternatives to Google Analytics?”

The target article was relevant, but the source sentence was an H2. The model judged semantic relevance without the editorial rule that headings should not become inline links.

In the same reconciliation, the anchor-text Score separated outcomes better: 1.63 for accepted suggestions versus 1.39 for rejected suggestions. Loomaly therefore shows only suggestions above 1.4 in its free check. This does not mean Score always beats confidence; it means keeping the signal that actually predicts outcomes in your workflow.

02
INPUT BOUNDARY

It can judge only the evidence you provide

On the first Pop Mart audit, Jev marked 168 pages as thin and 56 live product pages as lacking a next step. The conclusions sounded plausible, but the input was incomplete: body content loaded through JavaScript and the crawl was nearly blank; button data was never supplied.

CODE

Decide whether the question is valid

Below 150 body characters, do not ask whether content is thin; report that the page is empty to a non-JavaScript crawler. If buttons were not captured, do not ask whether the page offers a next step.

JEV

Judge the remaining semantic question

Only after evidence passes deterministic prerequisites should Jev judge specificity, CTA clarity, or intent fit.

The open-source Jev SEO project uses the same boundary: code owns crawling, rules, and scoring, while Jev handles semantic judgments such as page type, intent, helpfulness, and specificity.

03
QUESTION DESIGN

Question design changes the answer

NEGATION

Avoid negative questions

“Is this page missing author information?” increased directional errors. Loomaly blocks negative phrasing before the call and asks “Does this page have author information?” instead.

NUMBERS

Translate numbers in code first

Calculate page length, link counts, and other numeric facts in code, convert them into explicit categories, and pass the semantic state to Jev.

SCOPE

Judge one page at a time

Do not pack many page rows into one cross-row ranking decision. The Loomaly article reports an external test dropping from 0.97 to 0.51 but does not link the primary benchmark, so this site preserves it as an attributed observation rather than an independently verified result.

04
PRODUCTION PARITY

On launch, 40 of 50 suggestions already existed

Internal-link acceptance improved from 14% to 78% in testing. Yet among the first 50 production suggestions, 40 already existed on the customer site. The test script and production system read existing links differently, so offline evaluation never exposed the gap.

TEST14% → 78%iterative suggestion tuning
→
FIRST RELEASE40 / 50duplicates; the batch was discarded
→
AFTER FIX80%production acceptance

The largest gain did not come from another model tweak. Code rejected candidates that should never appear: existing links, previously rejected suggestions, image captions, table-of-contents items, headings, and anchors cut from the middle of a phrase. Image captions alone accounted for 14 of 26 rejections in one round.

SHIP CHECKLIST

If you want to use Jev in production

  1. Decide who owns the final action.Put countable facts in code and reserve Jev for semantic questions code cannot resolve.
  2. Calibrate on real outcomes.Do not worship confidence; find the signal that separates outcomes in your workflow.
  3. Give code veto power.Deletion, publication, charges, and other irreversible actions never gain authority from a high model score.
  4. Inspect what the model actually saw.When the input is empty or fields are missing, stop the question instead of accepting a polished but unsupported answer.