Mandate 2038: Balance and exploitability
Status: active simulation governance Machine authorities:
lab/contracts/balance-contract.jsonlab/contracts/experiment-matrix.json
The 0.8.0-rc.17-test simplification baseline retains the map, shared Programs, Research Protection, Training, Mega-Clusters, trades, Audit, Dossiers, Networks, and component supply. All numerical findings below that name removed programs or an earlier rules identity are historical diagnostics, not balance authority for the current baseline. The current package must earn fresh three-, four-, and five-player evidence before any faction or profile is described as balanced.
Win condition
Mandate 2038 should remain learnable but strategically resistant. No faction, seat, persona, opening, decision backend, or familiar matchup trick should win reliably from identity alone. Strong play should invite counterplay; counters should admit adaptation; several materially different routes should remain competitive.
“NP-hard” describes the desired lived search space, not a theorem this project claims. Finite experiments cannot prove optimality, computational hardness, permanent balance, or that no machine can solve the game. They can falsify a candidate by exposing dominance, brittle matchups, collapsed choice, loops, or missing coverage.
Faction truth precedes balance
Faction parity is constrained by the reason a player chose the institution. content/data/factions.json#factionTruthContract is the semantic authority for the seven relative identity dimensions, each faction's protected strength, plausible institutional bottleneck, and forbidden correction. Ratings run from one to five and describe relative institutional identity; they are not a point-buy budget and do not imply equal starting resources.
| Faction | Protected strength | Balance around |
|---|---|---|
| Dovetalis Labs | Deals, coalitions, institutional flexibility | Willing partners and individually rational terms |
| Loopfold AI | Installed Customers, distribution, market reach | Trust, exposure, deployment consequences |
| Mirevanta Works | Compute, Research reliability, Capability | Public validation, commercialization, claim timing |
| Kestralyn | Infrastructure velocity, capital, mobility | Scrutiny, rigidity, political legitimacy |
| Orisonix | Safety, Trust, controlled deployment | Slower expansion and fewer aggressive conversions |
| Corthaven | Compute abundance and supplier leverage | Opponent-count scaling and rivals' demand |
A proposed correction must name both the protected strength and the bottleneck it changes. It is rejected if it makes Mirevanta compute-poor or less capable because Research succeeded, Kestralyn slow at building, Corthaven a weak supplier, Loopfold bad at distribution, Dovetalis independent of partners, or Orisonix's Safety and Trust intrinsically losing. A numerically effective identity violation is a failed design, not a promotion candidate. Probes must also report whether a conversion reward makes one Core Action compulsory.
One experiment frame
All automated games belong to one seven-factor frame:
- rules configuration;
- player count;
- faction and Initiative seat;
- strategy roster;
- seed and Mandate mode;
- decision backend; and
- resulting negotiation behavior.
Cooperation is not a cohort. It is an observed result inside a match: non-binding promises may be made, superseded, fulfilled, broken, or remain unexercised. Power may be traded or withheld. A supplier receives causal credit only when removing its imported Power would make the buyer’s selected powered demand infeasible.
Rules probes change exactly one registered lever and reuse common seeds. Deterministic weighted and greedy backends supply bulk coverage. The matrix uses four backend regimes: all-weighted, all-greedy, and both alternating-seat patterns. Homogeneous regimes ask whether a strategy remains strong against opponents with the same degree of decisiveness. Alternating regimes expose backend interactions, but a strong backend beating a deliberately stochastic one is not by itself a rules defect. Provider-backed decisions appear only in a separate, metered, preregistered holdout.
The rules gate uses seat, faction, strategy, interaction, and head-to-head cells inside homogeneous regimes. Pooled and alternating-regime dominance is published as a diagnostic, not silently discarded, but cannot alone nominate a board-game rule change. Adaptive allocation includes regime-specific uncertainty so initial homogeneous coverage cannot be abandoned in favor of better-populated mixed cells.
When a matrix contains multiple rules configurations, each inference cell retains its configuration id. Paired cells use the same seed, faction rotation, persona rotation, seat, backend, and Mandate mode; adaptive batches advance every arm in that pair together. The report publishes both configuration-specific partially pooled intervals and bounded paired deltas. Pooling rules arms or changing their random streams invalidates an A/B claim.
After multiple one-lever candidates have independent receipts, one explicitly registered package-interaction comparison may combine them. It requires one empty canonical arm and one package arm, preserves common seeds, and reports that it validates interactions rather than discovering a new causal effect. Package evidence may reject the combination; it cannot retroactively replace the isolated evidence or conceal several changes inside a nominal single lever.
Sampling and uncertainty
The matrix does not full-cross every combination. It:
- gives each registered cell initial coverage;
- rotates factions, seats, strategies, and backends;
- uses common random numbers for controlled comparisons;
- then allocates batches toward large uncertainty and thin factor coverage;
- stops only at its registered precision target or match cap.
Sparse cells are partially pooled toward their family mean. Reports publish raw rates, empirical-Bayes posterior means and approximate intervals, and bounded Hoeffding confidence sequences. Adaptive looks use quadratic alpha spending; families share alpha through a Bonferroni split.
The approximate shrunk interval is diagnostic. A dominance flag requires the minimum exposure and both the shrunk interval and the multiplicity-safe confidence sequence to clear the registered bound. A noisy maximum from a thin matchup cell is never a rule-change authority.
What balance means here
- Faction and seat equity: identity and order do not decide the winner.
- Strategy and backend robustness: neither authored priorities nor the decision implementation should create an unanswerable line. Backend strength is interpreted within its regime: homogeneous games are the strategy test; mixed games diagnose interaction and opponent-quality effects.
- Interaction robustness: faction×strategy, faction×backend, and strategy×backend cells remain bounded after shrinkage.
- Matchup counterplay: head-to-head dominance and credible cyclic metas are reported directly.
- Opening and path diversity: no compulsory opener or single winning route.
- Information tension: an early lead does not settle the game.
- Negotiated viability: a necessary Power supplier can cooperate without reliably forfeiting a competitive finish.
- Procedural integrity: state remains finite and nonnegative, actions remain bounded, and provider fallbacks are visible.
Balance qualification, not absolute proof
No finite study can prove this game balanced for every strategy, player population, negotiation culture, and future metagame. A release may instead be described as balance-qualified only against a frozen, falsifiable contract. Failure to detect a difference is insufficient: confidence intervals for the registered faction, seat, strategy, and matchup effects must fit entirely inside their equivalence tolerances at the required precision.
Qualification requires all of the following for the exact named release:
- canonical rules, browser play, rich simulation, and batch simulation agree on legal decisions and state transitions;
- clean common-seed evidence can detect the registered practical imbalance;
- conclusions survive greedy, weighted, negotiation-capable deterministic, and independently controlled LLM policies without pooling their authority;
- four-player authority passes together with separate three- and five-player guardrails;
- suspicious mechanisms receive one-variable causal isolation;
- targeted AGI policies create enough Commit, Hedge, paid, and ineligible Dossier paths to evaluate the choice, while natural matrices independently measure the game-level matching-token emergence rate;
- a disjoint, preregistered seed population confirms selected conclusions;
- experienced human players rotate faction and seat while negotiation behavior is recorded; and
- adversarial policies search for loops, kingmaking, collusion, denial, and dominant openings.
Every conclusion remains bound to rules, engine, simulator, policy, study, source commit, and exclusions. The strongest defensible claim is therefore that a named version is balance-qualified for its declared three-, four-, and five-player configurations under the frozen tolerances and evidence classes. Stronger play may still contradict that maintained empirical claim.
Supported player-count contract
Mandate 2038 supports three, four, and five players.
- Four players is the authoritative balance configuration. Faction, strategy, opening, interaction, and counterplay decisions are diagnosed here first.
- Three players must preserve viable negotiation, credible refusal, meaningful scarcity, and faction viability despite the smaller market.
- Five players must preserve bounded faction scaling, reasonable Initiative and seat effects, useful negotiation, and acceptable congestion without allowing opponent-triggered abilities to compound uncontrollably.
Every unified promotion audit must cover all three counts. A four-player candidate is rejected if either adjacent supported count produces an integrity failure or credible faction, seat, strategy, or interaction dominance regression. Two- and six-player evidence is exploratory and cannot promote the current product.
Numeric bounds are provisional hypotheses. Their machine source is the balance contract, and their provenance must remain hypothesis until a cited evidence receipt supports a replacement.
The unified gate evaluates available outcome bounds separately for every supported player count in the candidate rules configuration. It directly checks faction win-share range, action entropy, opening entropy and top share, winning-path entropy and top share, policy fallbacks, and forced-no-op rate. Canonical data remains the paired baseline in a rules comparison, but a known canonical defect does not masquerade as a candidate failure. Missing supported counts, integrity failures, and registered dominance retain their separate, stronger verdicts.
Diagnostic funnels and faction swaps
AGI evidence separates each secret Dossier choice, final payment eligibility, Scrutiny cost, supported evidence modules, Capability threshold, Publication strength, selected Mandate rank, each deterministic tiebreak, and whether the supported claim overrode the provisional winner. The game-level emergence rate remains the primary rarity measure; no rate is assumed by rule.
The agi_claim_window_v1 compatibility scenario isolates final Dossier payment. Immediately before the first Era IV action selection it creates common-seed arms with four commitments and sufficient Compute or one fewer Compute. Passing this scenario qualifies payment reachability and deterministic policy choice only. Natural sweeps must separately test secret commitment patterns, supported-evidence frequency, claim strength, and deterministic tiebreaks.
Faction diagnosis uses realized ability values and preregistered common-seed swaps. A swap holds every non-focal input constant and changes only the faction at the focal seat. It reports paired win credit, Mandate, and rank deltas, but cannot promote a rule. A suspected ability must then enter a second, fresh-seed, one-lever unified audit. Old holdout seeds are retired after the decision they informed and may not be reused to select later corrections.
Foundry selection evidence
The 2026-07-27 preregistered Foundry scaling study selected one change: Everybody Gets a GPU scores one Mandate per four rivals. Across 3,976 common-seed pairs, the lever reduced Foundry win share by 6.70 percentage points with a multiplicity-adjusted interval below zero and removed the credible six-player Foundry×greedy signal from that rule arm. Starting Compute and New Architecture were tested independently and retained because their paired effects were not distinguished from noise.
That historical conclusion is no longer current balance authority. Executable 0.7.0 applied Shovels after Core Actions only, while the qualifying two-Compute spends occurred in Escalations. The paired GPU result remains a faithful observation of that defective executable, not of the physical rules.
The corrected 0.7.1 three-arm matrix retained the divisor of four: restoring two rivals per Mandate increased Foundry win share by 6.094 percentage points and mean score by 0.870, while producing a credible six-player Foundry×greedy cell. Reducing Shovels from twice to once per Era changed no standings and essentially no score, so the cap of two remains a hypothesis rather than receiving evidence provenance.
The historical receipt is 2026-07-27-foundry-scaling-rule-selection.md. The correction is 2026-07-27-foundry-shovels-executable-correction.md. The current selection receipt is 2026-07-27-corrected-foundry-final-selection.md.
The supported-count follow-up is 2026-07-28-foundry-supported-count-conversion.md. Disabling the remaining Era IV Mandate changed Foundry by -4.365 win-share percentage points and -0.939 mean score at five players, with exactly zero effect at three and four. The discontinuity is real but cannot explain Foundry's broader strength, which is already present where the ability scores zero. Removal is rejected as a standalone correction. Starting Compute and supplier leverage remain protected; a later candidate must make victory conversion depend on realized rival demand rather than weaken the supply fantasy.
Adversarial diagnostic
The unified audit embeds a small best-response, counter-response, and holdout slice. It uses common seeds and partially pooled rates, but remains diagnostic: bounded mutation cannot prove that a strongest response was found. It may identify a one-lever probe; it cannot promote that probe.
The older npm run simulate:adversarial command remains available for historical comparison. Its unshrunk worst-cell maxima are calibration evidence only.
LLM negotiation holdout
Provider calls default to zero. A holdout requires:
- a committed preregistration under
evidence/studies/simulation/preregistrations/; - explicit
--allow-llm; - named seats, profiles, backend, model, seed, stages, and a declared provider allowance; the engine separately rejects any repeated immediate-trade formal window and enforces that rule-derived packet ceiling;
- strict holdouts set
strictLlmEvidence: true: provider failures, malformed decisions, gated configured-model stages, and historical fallback cache entries quarantine the match rather than selecting a deterministic fallback. Anullstage list means the configured model owns every legal decision for its seat; deterministic seats remain explicitly declared opponents. - CLI version, model, prompt hash, selected decision, fallback, and source provenance in the report.
Fresh calls test robustness and may capture successful decisions in write-only mode, which never reads an older answer. Cached calls test reproducibility. A read-only cache miss fails rather than silently becoming a fresh call. Provider failures retain attempted-provider provenance and a stderr hash even when play continues through the deterministic fallback. The initial holdout is stage-gated to negotiation decisions and is descriptive, not a provider comparison or balance estimate.
Promotion gate
No simulation changes rules automatically. A rule adjustment requires:
- a clean committed source;
- a preregistered or explicitly identified hypothesis;
- a raw archived report and tracked dated receipt;
- minimum exposure and uncertainty-aware evidence;
- one rule lever changed on common seeds;
- an audit of rulebook, semantic graph, engine, UI, player aids, schemas, tests, and physical-test protocol;
- a new immutable release; and
- explicit user approval.
Physical playtests remain the authority for teachability, duration, emotional fairness, memorable betrayal, and whether Realignment feels dramatic rather than tedious.
Passing means “no registered exploit exceeded these bounds in this bounded search.” It never means “solved,” “balanced,” “machine-proof,” or “fun.”
An automated balance result is invalid when deterministic policies exhaust more than 3% of ordinary action opportunities without a legal resolution. Selection packets omit choices that cannot resolve from current state and cannot gain a resolution through one legal accepted pre-Act trade. They label trade-required choices separately. Later target loss and rejected required trades remain legitimate commitment failures and are reported independently from impossible-choice filtering.
Current canonical screen
The corrected homogeneous-backend screen ran 11,920 canonical matrix matches plus 70 bounded adversarial diagnostics from a clean source. No homogeneous seat, faction, strategy, interaction, or head-to-head cell cleared both registered dominance gates. Precision remained incomplete, and three pooled Mirevanta Works×Capability Rusher cells remained diagnostic. No rules value was changed.
The tracked receipt is 2026-07-27-final-homogeneous-backend-screen.md.
Current faction-conversion evidence
The preregistered progress-conversion matrix ran 11,986 clean matches across three, four, and five players. Removing one Capability when Scientific Method saved a run reduced Demis's paired win share by 3.60 percentage points but violated faction truth, remained imprecise, and was rejected. Awarding one Mandate when Industrial Velocity actually reduced the cost of a completed Facility increased Elon's paired win share by 6.61 percentage points and mean score by 1.34 without materially changing Build selection.
The Elon result is a nomination for independent fresh-seed confirmation, not a physical rule. The next Demis candidate targets public validation while preserving Compute, Research reliability, and Capability. The broad matrix remained inconclusive_precision_not_reached; no current simulation certifies the game as balanced.
The tracked receipt is 2026-07-28-faction-progress-conversion-calibration.md.
The independent fresh-seed confirmation then ran 11,986 clean matches. Elon's realized Industrial Velocity Mandate reproduced: paired win share increased 6.71 percentage points, mean score increased 1.36, mean placement improved 0.258 ranks, and four-player partially pooled win rate moved from 17.78% to 25.47% without increasing Build selection. It is promotion-eligible but not promoted without explicit approval.
The identity-safe Demis candidate preserved every Capability point and withheld 1,043 threshold Mandate, but improved parity by only 2.18 percentage points in aggregate and left Demis above parity at every supported count. The exact candidate is rejected as insufficient; a later diagnostic must isolate the remaining public-validation conversion without weakening science.
The tracked confirmation receipt is 2026-07-28-faction-public-validation-confirmation.md.
The subsequent demand-validation matrix ran 11,986 clean matches on fresh common seeds. Adding two Scrutiny whenever Scientific Method protected a run changed Demis by only -0.43 win-share percentage points and -0.09 Mandate in aggregate; it preserved Capability but did not make public exposure a meaningful bottleneck. The candidate is rejected.
Making Jensen's New Architecture self-Compute equal to one plus accepted rival licenses preserved every offer, sale, and payment while reducing Foundry win share by 1.07, 0.84, and 1.33 percentage points at three, four, and five players. The direction is valid but the exact formula is too weak. A later one-lever probe may test license-only self-Compute while preserving four starting Compute, supplier offers, revenue, and the three-Compute ceiling.
The tracked receipt is 2026-07-28-faction-demand-validation.md.
The next prestige-and-demand matrix also ran 11,986 clean matches. Reducing Nobel Effect from two Trust to one changed Demis by less than 0.4 win-share percentage points at every supported count because the ability activated in only 62/67/57 of 768/888/1,018 appearances. That candidate is rejected; the remaining Demis correction must target late public validation without weakening science.
Removing New Architecture's unconditional base Compute while retaining one self-Compute per accepted license moved Foundry from 35.14%/28.73%/25.25% to 30.18%/24.82%/22.79% at three, four, and five players. Every offer, accepted license, Runway payment, and rival Compute grant remained unchanged. The candidate is selected for combined-package confirmation, not yet promoted.
The tracked receipt is 2026-07-28-faction-prestige-demand.md.
The isolated Demis late-validation matrix ran 7,990 clean matches. Reducing Capability 9 and 12 from two Mandate to one moved Demis by -6.89, -9.26, and -10.17 win-share percentage points at three, four, and five players, with a multiplicity-adjusted aggregate interval entirely below zero. Research, Scientific Method, Scaling-Law Capability, and threshold attainment remained effectively unchanged.
The candidate produced competitive score bands at three and four but overcorrected five because late-threshold attainment rises with player count. The scalar rule is rejected. Its public-validation direction is retained for a narrow follow-up in which four rival institutions can restore the second Capability-12 Mandate through broad validation.
The tracked receipt is 2026-07-28-demis-late-validation.md.
The registered peer-validation refinement then ran 7,990 clean matches. Demis retained full Capability while Capability 9 and 12 scored one Mandate; at five players, four rival institutions restored Capability 12's canonical second Mandate. Paired win share moved by -9.17, -8.64, and -7.54 percentage points at three, four, and five players.
The resulting Demis win shares were 42.38%, 26.82%, and 19.20%. The four-player balance-authority result entered the leading band, while the five-player restoration avoided the scalar candidate's 13.54% overcorrection. Research selection, Scientific Method Capability, and Scaling-Law Capability remained effectively unchanged.
The structured schedule is selected for a combined-package interaction audit, not promoted to physical rules. The broad matrix still reported inconclusive_precision_not_reached.
The tracked receipt is 2026-07-28-demis-peer-validation.md.
The selected three-lever package then ran 11,998 clean matches with 5,964 common-seed pairs. At the four-player balance authority, all six candidate factions fell inside a 9.28 percentage-point win-share range. Three- and five-player ranges were 13.03 and 12.64 points. No registered faction, strategy, backend, or head-to-head dominance appeared.
Demis retained every technical output while late public validation reduced his paired win share by 9.28 points at four players. Elon's realized-build Mandate improved his four-player paired win share by 7.44 points without increasing Build selection. Jensen retained every rival license, payment, and Compute grant while demand-coupling reduced his four-player win share by 4.81 points. Every effect kept its isolated direction at three and five.
The package is selected for explicit physical-rule approval, not automatically promoted. It is not overall balance certification: matrix precision remained incomplete, and winning-path entropy was 0.572 at four players and 0.547 at five, below the provisional 0.60 floor even though no path exceeded the separate 55% top-share ceiling. That remaining path-concentration question must be diagnosed separately from faction conversion.
The tracked package receipt is 2026-07-28-selected-faction-conversion-package.md.
The corrected-gate confirmation repeated the package on 11,998 fresh, clean-source matches. Faction ranges were 13.41, 12.19, and 13.36 percentage points at three, four, and five players. No registered dominance, pairwise dominance, meta cycle, integrity failure, fallback, action collapse, or opening collapse appeared.
Overall balance correctly remained red. Winning-path entropy was 0.571/0.600/0.535, below the provisional 0.60 floor, while Adoption represented 46.53%/46.28%/50.53% of classified winners. Those shares pass the separate 55% ceiling; the failure says that two large routes and several rare hybrids do not yet provide enough measured strategic diversity. The next diagnostic must attribute Mandate sources, actions, faction, and outcomes to winning paths before a global lever is proposed.
The enforced-gate receipt is 2026-07-28-selected-faction-package-gate-confirmation.md.
Winning-path reports now attribute each winner-credit share to its Mandate sources, Core Action selections, final Capability, Customers, Facilities, Trust, faction, persona, backend, and World Ending. This evidence distinguishes a scoring conversion defect from a classifier or deterministic-policy defect; only a later fresh-seed one-lever experiment may change a physical value.
The first attribution diagnostic finds that Adoption winners receive 2.29/2.20 more Customer Mandate than Research winners at four and five players, while Research recovers only 0.93/0.67 through Capability. Trust, Era-objective, faction, and backend composition do not explain the gap.
The eligible next probe preserves every Customer and its income, costs, requirements, Mark's starting Customer, and distribution identity. It tests only diminishing public recognition: the first three Customers retain two Mandate, while the fourth and fifth score one. That schedule is not a physical rule unless a fresh one-lever matrix improves path diversity without creating faction, action, or matchup regressions.
The attribution receipt is 2026-07-28-winning-path-attribution.md.
The isolated late-Customer recognition probe then ran 11,998 clean matches. Keeping two Mandate for Customers one through three and scoring one for Customers four and five improved winning-path entropy at every supported count without changing Customer income, requirements, or distribution mechanics. It did not independently repair the known faction ranges, so it was retained only for an explicitly registered package-interaction test.
The receipt is 2026-07-28-late-customer-recognition.md.
The four-lever package combined that schedule with the selected Demis, Elon, and Jensen conversions. Across 11,998 clean matches, faction ranges fell to 12.61, 12.60, and 12.84 percentage points at three, four, and five players. Every enforced result passed except five-player winning-path entropy, which remained 0.5731 under a classifier that recognized a hybrid only when two lane scores tied exactly.
The package receipt is 2026-07-28-selected-faction-late-customer-package.md.
The preregistered margin-validity diagnostic found that 56.70% of five-player package winner credit finished with its top two lanes within one point, while only 19.08% tied exactly. The exact-tie rule was therefore rejected as a brittle measurement contract. Infrastructure was also confirmed as an engine that converts into Research or Adoption, matching the authored Infrastructure Compounder, rather than a required standalone scoring lane.
The diagnostic receipt is 2026-07-28-winning-path-margin-validity.md.
Both executable runners now use the same lane-margin-v1 classifier. A winner is labeled as a hybrid when the top two lane signals finish within one action, Customer, Facility, or comparable end-state unit; AGI remains a separate path, and lane-pair names have one canonical order.
The fresh preregistered confirmation ran 11,998 matches from clean commit 72678155, including 5,964 common-seed pairs. The selected package passed every observed three-, four-, and five-player faction-range, action-diversity, opening-diversity, winning-path, fallback, and forced-no-op bound. Candidate faction ranges were 12.67, 12.00, and 12.72 percentage points. No registered or diagnostic dominance, pairwise dominance, integrity failure, or fallback appeared.
This selects the package for controlled physical testing; it does not certify the game as solved or permanently balanced. The top-level result remains inconclusive_precision_not_reached because thin, multiplicity-controlled interaction cells still have wide intervals. The physical rules remain unchanged pending explicit approval, and human play remains mandatory for negotiation, Realignment, duration, fairness, and fun.
The confirmation receipt is 2026-07-28-winning-path-tolerance-confirmation.md.