Python — Model evaluation: independent calibration test of Jev on 900 rule-generated support tickets it cannot have seen plus three public benchmarks, publishing every raw response, ECE against a simulated noise floor, temperature refit, and the per-type sign of miscalibration (Choice and Score overconfident, Boolean underconfident).
First discovered , last refreshed . Descriptions and stats are pulled from the project's own GitHub repo and refreshed automatically — they aren't independently verified by JevTracks beyond the initial eligibility check.