JevTracks

Submit your tool

← JevTracks

Jev V13, a decision-making system, was tested playing blitz chess against frontier LLMs like Fable 5.1 and GPT-6 Astra, with each move requiring one API call.

aimlapievaluation-benchmarking548.9K views · 2.3K likes

First discovered , last refreshed . Descriptions and stats are pulled from the project's own X post and refreshed automatically — they aren't independently verified by JevTracks beyond the initial eligibility check.