ANA β€” A Normal Adversary

A small action model that plays Magic: The Gathering like a human β€” it drafts at the rank you dial, then plays the games. These are the weights serving the live demo.

5,926,784 parameters. Runs on a CPU. Trained on 13.9M human draft picks from 17Lands' public MTG Arena data (sets TDM, EOE, DSK).

Measured (held-out, session-disjoint):

  • given the same draft decision β€” the same two cards, both available β€” ANA sides with a human 91% of the time. Two humans side together 75%. That is consensus-prediction, not "more human than humans" β€” each human is one draw from a spread. (98M matched decision pairs.)
  • skill dial: per-tier pick-match climbs bronze 64.9% β†’ mythic 70.9%, monotone, and calibrated β€” each rank setting predicts its own tier best (6 of 6).

Play it: alakazam.gg/ana-mtg Β· Essay: blog Β· Technical report (120pp, negative results included): PDF

Gameplay heads (also in this release β€” the full seat from the live demo):

  • in game, its first choice is the human's 2 turns in 3 (66.8%, 400 held-out decisions from a set it never trained on)
  • 1.6Γ— stock Forge AI at predicting people (names the human's exact play 62.5% vs 39.3%, same decisions, paired)
  • the attack head serves by sampling its own distribution β€” the readout, not just the weights, is part of acting human. Full story in the report.

Honest limits

Drafter: three sets (TDM/EOE/DSK). Gameplay: 11 training sets, evaluated on foreign holdouts; a two-line mana-curve heuristic still beats it on cast top-1 (the metric is nearly saturated β€” the report explains). Blocks are engine-side in the demo. Splits are session-disjoint, not player-disjoint (player identifiers do not exist in the public corpus). The rank dial conditions drafting; conditioning in-game play is an open problem, and our report says so.

This model is tier-level by design and is never trained per-player: per-player imitation models have been shown to re-identify individuals from ~100 games (McIlroy-Young et al.), and we deliberately do not build them.

Files

  • ana_drafter_pool_tokens_step5076.pt β€” drafter state dict (pool-attention arm, serving the live demo; 5.93M params; sha256 prefix 61e4c9ed61824dcd)
  • config.json / vocab.json β€” drafter config + 1,244-entry vocabulary
  • ana_player_cast_mull_cap_d256_step12000.pt (+.config.json) β€” gameplay cast + mulligan heads (5.59M params; sha256 prefix b7e2b28bb263ae00)
  • ana_player_attack_v9_joint_step12000.pt (+.config.json) β€” the joint-turn attack head served with sampled decode (1.14M params; sha256 prefix d0ccace64da119fa)
  • vocab_v3.json / scryfall_attrs_v3.npz (+ manifest) β€” gameplay vocabulary (8,680 entries) and the 79-dim card-attribute matrix the gameplay heads consume
  • Whole system: ~12.7M parameters across three models, CPU-served

Drafter inference code

  • modeling_ana.py β€” the drafter architecture, standalone. Imports torch and the standard library and nothing else; it does not need our training repo. Loads the checkpoint with weights_only=True, so these weights cannot run code on your machine.
  • example_draft_pick.py β€” card names and a tier in, pick probabilities out. Run it with no arguments for a real contested pick scored at three ranks.
  • reproduce_gate.py + fixtures/ β€” the receipt below, runnable by you.
pip install torch numpy
python reproduce_gate.py

Receipt. modeling_ana.py is a copy of the training code, and a copy can be wrong in ways that still read fine. So the check ships with it. fixtures/reference_scores.json holds per-row scores produced by the original training repository β€” its own checkpoint loader, its own model builder, its own dataset class, importing nothing from this repo. reproduce_gate.py scores the same 1,000 held-out validation rows here and compares. In a venv containing only torch 2.13.0 and numpy 2.5.1, on CPU, with this repo copied out of the source tree and PYTHONPATH cleared:

rows                 1000
top-1 (this code)    0.7020  (702/1000)
top-1 (reference)    0.7020  (702/1000)
random baseline      0.2355  (mean 1/pack_len on these rows)
argmax agreement     1000/1000
max logit deviation  5.07e-07

Every one of the 1,000 argmaxes matches; the residual is the 6-decimal rounding of the stored reference logits, not a numerical disagreement. The fixture is a fixed-seed 1,000-row sample drawn across all shards of the 776,147-row validation split, so its top-1 carries roughly a Β±3 pp sampling interval. It is a self-test that fails loudly if this code drifts from these weights, not a second measurement of the model.

One pick costs about 4.5 ms on a single CPU thread.

Gameplay loaders are not in this drop. The two gameplay heads are different architectures over the 8,680-entry vocabulary and the card-attribute matrix, and they do not load into modeling_ana.py; the technical report documents them.

Attribution

Data from 17Lands (CC BY 4.0). Card attribute data from Scryfall; portions are property of Wizards of the Coast. Unofficial Fan Content permitted under the Wizards of the Coast Fan Content Policy. Not approved/endorsed by Wizards. Β©Wizards of the Coast LLC.

License: CC BY-NC-SA 4.0 β€” non-commercial. Want it in a product? Talk to us.

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support