Model leaderboard
Benchmark another model
mbench run <llama-swap id> full suite, 2–5 hours, runs in the background mbench run <id> --quick 45–90 minutes, ranked as provisional mbench run <id> --reuse keep what still applies from the last run mbench run <id> --submit also submit to localmaxxing mbench compare <id> <id> which differences are real, task by task mbench status · mbench logs -f · mbench ls · mbench board --open