privatepanda.co

projects · flagship

MSPoker in active build

solo · design + implementation · Rust engine · Python training & search

One brain, several seats: a poker AI for the open multiplayer, multi-seat version of No-Limit Hold’em.

The problem

Two-player No-Limit Hold’em is effectively solved: counterfactual regret minimization (CFR) produces strategies that cannot be meaningfully exploited. Add more players and those guarantees disappear. MSPoker studies a sharper version of the multiplayer game: a single agent controls k of the n seats at a table.

The core idea is to play the controlled seats as one joint player rather than k independent bots. All k seats share one information set, one reward, and one policy. The set of joint strategies strictly contains the product of the per-seat strategy sets, so the best joint strategy weakly dominates independent play. The pooled hole cards also give the coalition extra information about every opponent’s range. That is why MSPoker is multi-seat instead of several copies of a heads-up bot.

This is a research system. It is built and evaluated in simulation, and this page describes the architecture and the design decisions behind it.

Architecture

seat 1 seat 3 seat 5 sharedhole cards coalition brain1.4M-param seat-attn CFR blueprint PPO fine-tune CFR+ live search jointaction

Rust game engine

The engine is written in Rust and exposed to Python through PyO3/maturin. Money is an i64 in milli-big-blinds, so no floating-point chips exist anywhere, and every hand is checked against chip-conservation invariants. Hands are deterministic and replayable. The hand evaluator is a sort-and-categorize design packed into a u32 rank instead of a 130 MB lookup table: no static data, fully auditable, and still within the microsecond budget, with optimization kept behind a stable evaluate_5/evaluate_7 API.

Seat-attention network

A network of about 1.4M parameters attends across seats, so the controlled seats are represented together rather than as separate inputs. It is the policy and value core the blueprint and the search both rely on.

Blueprint, then exploitation

Training starts with a DREAM/CFR-family blueprint, which gives a provable floor on exploitability. PPO then fine-tunes against weaker opponents with a KL anchor back to the blueprint. One method alone gives either safety or exploitation; the combination gives both.

Joint real-time search

At decision time, CFR+ search refines the blueprint. A coalition range gadget computes the team’s counterfactual values jointly. Summing per-seat values would be wrong, because card-removal and blocker effects cross between the coalition’s known cards; the spec treats a per-seat sum as a build-breaking defect.

Opponent modelling

Beta-binomial statistics track each opponent’s tendencies, so exploitation adapts to the players actually at the table instead of a fixed population.

Built like flight software

The system is built spec-first against about 155 ordered, verifiable tasks. Each task has a contract, the engineer and auditor roles are kept strictly separate, and phases close only on a human-signed git tag. A correctness-critical AI system is built auditably rather than in one pass.

Where it is now

The design is complete (fourth generation, after three earlier iterations). The configuration layer, the Rust engine, the data models, and the neural network are implemented. Multi-seat coordination, training at scale, and live search are the next phases.

Rust · PyO3 / maturin · Python · neural networks (attention) · CFR / CFR+ / MCCFR · DREAM · PPO · Bayesian opponent modelling

back to projects →