Python — Model evaluation: measures whether a SQL ORDER BY over a Jev probability is defensible (pairwise inversion, Score ordinality against a human grade, calibration, wording invariants, sort-key ties) under a pre-registered gate that jev-1.13.0 passes on 20 Newsgroups topics and fails four of six conditions on Amazon ESCI product relevance, and shows a DuckDB extension's default 40-row batching fails the ranking gate that one row per request passes.
First discovered , last refreshed . Descriptions and stats are pulled from the project's own GitHub repo and refreshed automatically — they aren't independently verified by JevTracks beyond the initial eligibility check.