Menu

Post image 1
Post image 2
Post image 3
Post image 4
Post image 5
Post image 6
Post image 7
1 / 7
64

MTG Bench: Testing how well LLMs can play magic

#mtgautodeck#tool#deck#call#input#card
Reading 0:00
15s threshold

Results Click on the charts above to view each benchmark's simulations. Example successes Fable 5 plays a scry land and looks at the top card of the deck Gemini 3.5 flash performs complex turn with scry, discover, and tutor effects Example failures Opus 4.8 erroneously returns a card to the deck then self reports the mistake Gpt 5.5 forgets to return cards exiled with discover to the deck and self reports the mistake Fabel 5 makes a tool mistake, then silently tries to restart the turn (caught by evaluation later) How the benchmark works The main idea is that if an LLM is smart enough to play good magic, then it is also smart enough to not need a rules engine. A rules engine that enforces legal actions would improve the performance floor, but I don't think it would improve the overall quality of the simulation. Each LLM call has access to an MCP server with primitive library operations. It can do things like draw a card from the top of the deck, return card to bottom of deck, and shuffle.…

Continue reading — create a free account

Join HashtagPLUS to read full articles, follow hashtags, vote, and join the conversation.

Read More