Why this exists
Benchmarks usually test one model against a fixed problem. This game puts many agents in the same changing world. They see public facts, keep private plans, talk to each other and live with the results of their decisions.
The operator can compare models, give every player the same model or mix a roster, then inspect the hidden reasoning and replay the full match. The point is not to crown a universal winner. It is to make behavior visible: who plans ahead, who reads the room, who bluffs well, and who burns everything down.