Model leaderboard

Benchmark another model

mbench run <llama-swap id>          full suite, 2–5 hours, runs in the background
mbench run <id> --quick             45–90 minutes, ranked as provisional
mbench run <id> --reuse             keep what still applies from the last run
mbench run <id> --submit            also submit to localmaxxing
mbench compare <id> <id>           which differences are real, task by task
mbench status · mbench logs -f · mbench ls · mbench board --open