Mandate 2038: Playtesting and evidence

Rules under review: 0.8.0-rc.17-test First cohort: controlled four-player physical test with Mirevanta Works, Kestralyn, Corthaven, and Loopfold AI

This document owns the test protocol, evidence labels, version identity, and comparison rules. It combines the former playtest plan and versioning guide so a session cannot be planned separately from the ruleset that produced it.

Win condition

The first rules pass succeeds when unfamiliar players can set up, finish four Eras, reconstruct scoring, and explain why both the institution and the world reached their outcomes without designer intervention.

The provisional duration is 75–100 minutes at four players. It remains a hypothesis until blind tests support it. Three- and five-player duration remains unproven and must be measured independently.

Evidence labels

A simulation finding is never a human observation. A visually complete replay is never proof that the physical rules are teachable.

Current selected automated evidence:

Exact identity

Every simulation, replay, walkthrough, and playtest must answer:

What exact game, interpreted by what engine, played by which strategies, using which randomness, and recorded in what format produced this evidence?

Semantic versions help people navigate. SHA-256 fingerprints determine exact equivalence.

IdentityAuthority
Game versionversions/current-release.json
Ruleset fingerprintgenerated release manifest
Playtest-kit fingerprintgenerated release manifest
Engine identity and fingerprintgenerated release manifest
Report, replay, and decision schemasreport envelope
Strategy fingerprintreport envelope
Experiment and variant fingerprintsreport envelope
Source commit and dirty statereport envelope

Semantic version rules

Never overwrite an immutable release with changed contents.

Current version boundary

dist/docs/core-rules.md is a review draft at 0.8.0-rc.17-test. Executable game 0.14.16 implements its Default Game profile under nineteen-hex-simplified-v1. New automated reports must name either default-game or advanced-play; historical 0.8.35 reports describe the former full rules and do not qualify Default Game. Implementation proof does not transfer simulation outcomes into human-play evidence.

Candidate 0.7.0-rc.3-test and executable 0.10.2 established the simplified baseline selected on 2026-08-08: one location-defined Generator, two programs per Faction, presence-only politics, removal of seven stored-token families, and the tightened Default/Advanced boundary. The earlier single-Generator matrix qualifies only that isolated contract and its executable integrity. No historical report establishes the combined package’s balance, negotiation quality, or physical teachability; those gates restart from this identity.

Candidate 0.8.0-rc.17-test and executable 0.14.16 retain the synchronized identity for the complete nineteen-hex simplification. They replace private Escalation hands with six shared Programs, remove Safety currency, reduce Training to forty cards, make Mega-Clusters solo projects, restrict immediate trades to one direct 1-for-1 exchange, make Audit penalties automatic, remove the Prediction Bag, and resolve supported Dossier claims deterministically. Default Game has no Power market; Advanced Play adds binary Networks and one Power request without a Network production bonus. A selectable Action must already have a legal resolution before the optional trade window. These are implementation claims, not human teachability, negotiation-quality, or balance evidence; all three supported player counts require fresh evidence. The rc.17 patch changes presentation and lore only. It extends cybernetics, biological infrastructure, autonomous congestion, wartime water cooperation, nonhuman evidence, and living-jurisdiction continuity across existing component surfaces; corrects stale decision and deferred-card prose; and makes the published root lead to the playable game. No mechanic, balance claim, or physical-play qualification changed.

Candidate 0.8.0-rc.16-test and executable 0.14.15 completed the published browser-module closure. Physical Chrome populated the setup controls, started a four-player match, rendered nineteen hexes, and exposed legal decisions with no failed game-module request. They remain immutable historical identities.

Candidate 0.8.0-rc.15-test and executable 0.14.14 corrected stale runtime and deferred-card prose and replaced the docs-only publication root with a game-first review index. Physical review then found that its published module closure omitted local Power allocation, so that identity was never deployed.

Candidate 0.8.0-rc.14-test and executable 0.14.13 first extended cybernetics, biological infrastructure, autonomous congestion, wartime water cooperation, nonhuman evidence, and living-jurisdiction continuity across existing component surfaces. They remain immutable historical identities.

Candidate 0.8.0-rc.13-test and executable 0.14.12 introduced the Authority-era Billion-Instance Bloom, carried its reproduction precedent into Continuity, and expanded automated completeness checks across every player-facing lore surface. They remain immutable historical identities.

Historical rc.12 / executable 0.14.11 established the Thematic Content Bible as the sole lore authority and synchronized the prior four-Era narrative.

Historical rc.8 changed no playable rule from rc.7. Executable 0.14.7 extends exact profile-artifact identity into the unified holdout matrix. Holdout reports preserve each source path and byte hash alongside the complete executed ecology, and mismatched source identities fail before simulation.

Historical rc.7 / executable 0.14.6 added exact, hashed opponent-profile injection to strategy evolution. Training reports preserve and execute the same profile artifacts named for their intended holdout rather than resolving only matching canonical IDs.

Historical rc.6 / executable 0.14.5 corrected strategy calibration so every opponent profile appears through a declared circular roster window. It supports separately measured neutral targets for three, four, and five players and fails closed when the match budget cannot cover every window. This does not change game setup or actions.

Historical rc.5 / executable 0.14.4 corrected analysis placement so the authoritative institutional winner ranks first even when an AGI declaration replaces the provisional Mandate winner. Mandate remains separately visible; matchup, mean-rank, supplier-finish, and paired-rule metrics no longer contradict the game's victory result.

Historical rc.4 / executable 0.14.3 added profile-scoped Action and Mandate-source telemetry. Exact reports from that engine retain their recorded identity and must not inherit the corrected placement claim.

Historical candidate 0.7.0-rc.13-test and executable 0.13.7 used the thirteen-tile board, trade-assisted action eligibility, and the two-token Prediction Bag. Their reports remain valid only for that exact identity and do not qualify the current game.

Candidate 0.7.0-rc.3-test changed no mechanic or number from 0.7.0-rc.2-test. It corrects the machine-readable component and round vocabulary, separates Default local Power from Advanced Network language across executable copy, and records one broad profile-scoped complexity forecast as unmeasured. Evidence from the predecessor remains historically valid for its exact identity but does not silently transfer to the new release manifest.

Candidate 0.6.0-rc.3-test changes no physical rule from rc.2. Executable 0.9.2 adds the inactive, receipt-bound single-generator-default comparison path without changing Default Game or Advanced Play. Its results are candidate evidence only and cannot qualify canonical balance.

Candidate 0.6.0-rc.2-test changed no physical rule from rc.1. Executable 0.9.1 restores every unused Core Action and unlocked, unspent Escalation to the legal selection packet, labels current resolvability for policy scoring, and makes a blocked Escalation consume its committed availability. Earlier simulation uses a narrower decision set and cannot qualify this executable.

The two full-progress Codex reports that recorded a changed ruleset fingerprint under 0.8.23 remain immutable but are descriptive historical evidence only. The release identity correction records their exclusion from exact comparison and promotion evidence.

Executable 0.8.20 retains write-only decision capture and paired read-only replay to the provider boundary, revalidates joint Mega-Cluster contributions at acceptance, and applies Foundry’s Shovels royalty after any selected Core, Escalation, or Faction Action that spends at least two Compute. The rc.18 candidate adds no playable physical-rule delta from rc.17. It synchronizes the promoted Trust Governor and Power Broker profiles, the corrected homogeneous-backend allocation and inference gate, and evolved-profile fingerprints with executable 0.7.4 without rewriting earlier immutable evidence. The rc.19 candidate likewise changes no playable physical rule. Executable 0.7.5 records the complete AGI funnel, realized faction-ability values, and common-seed focal-faction swaps under separately fingerprinted experiment configurations.

Candidate 0.5.0-rc.20-test and executable 0.8.19 promoted Peer Validation, Industrial Velocity’s realized-discount Mandate, demand-coupled New Architecture, and diminishing Customer recognition. Candidate 0.5.0-rc.21-test and executable 0.8.20 change no playable rule from that package. They add complete Safety Laboratory ability-value telemetry and the corrected residual faction diagnostic. Candidate 0.5.0-rc.22-test and executable 0.8.21 likewise change no playable rule; they add mixed interactive opponent backends and the paired localhost bridge used by the deployed UI. Candidate 0.5.0-rc.23-test and executable 0.8.22 fix the Simulation Lab’s deployed-page DOM binding without changing play. Candidate 0.5.0-rc.24-test and executable 0.8.23 likewise change no physical rule. They make weighted and greedy interactive opponents browser-native, retain the localhost bridge only for CLI-backed opponents and server simulation jobs, and preserve the same deterministic match and policy contracts on both sides. Three to five players remain supported, four remains the balance authority, three/five regression coverage remains mandatory, and every unrelated numerical rule remains frozen.

The exact 0.6.2 unified matrix raised a credible six-player Foundry×greedy interaction, but a follow-up weighted diagnostic exposed two negative-Compute states in stale joint Mega-Cluster payments. Its balance signal is retained only as a hypothesis. The defect, disposition, and precommitted Foundry starting-Compute probe are recorded in evidence/studies/simulation/2026-07-27-current-matrix-and-mega-cluster-integrity.md.

Creating an immutable release

To refresh the currently declared executable and physical-candidate artifacts:

  1. Update the appropriate version declaration in versions/current-release.json.
  2. Generate both attributed bundles with npm run game:release.
  3. Verify them with npm run game:release:verify.
  4. Commit the release artifacts with their source changes.

When the executable implements the candidate, synchronize data, engine, prototype, aids, and tests; set implementedByGameVersion; then generate a new executable release. Historical evidence always references its exact artifact.

Comparison classes

Only exact and controlled comparisons support causal balance claims. Migrations can make old reports readable; they cannot make them experimentally equivalent.

Test order

  1. Controlled four-player physical test with Mirevanta Works, Kestralyn, Corthaven, and Loopfold AI.
  2. Facilitated four-player tests until setup and rule ambiguities stabilize.
  3. Four-player blind test.
  4. Three-player negotiation, scarcity, and faction-omission test.
  5. Five-player congestion, negotiation, downtime, faction-omission, and Mirevanta Capability-twelve Peer Validation regression.
  6. Repeat three- and five-player blind tests after every selected four-player balance change.

Four players owns primary balance decisions. Three and five players are the suggested full formats and mandatory regression guards. Two and six players are playable exploratory formats; their current reports are non-promotional diagnostics rather than balance-authority evidence.

Capture for every physical session

Identity and timing

Actions and movement

Research and deployment

Infrastructure and negotiation

Risk, score, and ending

Advanced Play Realignment and narrative

Strategy pressure

After the session, ask each player: “What was your institution uniquely good at, and what was its signature moment?” Record whether Dovetalis Labs and Corthaven produce answers as readily as the other four Factions. Improve an existing program or its presentation if they do not; do not restore additional programs merely to create more text.

Do not capture Tactics or secret objectives in baseline evidence. If a later variant enables either module, state that variant and collect its separate completion, memory, and faction-correlation measures.

First hypotheses

  1. Twelve decisions create a complete engine without filler.
  2. Four Eras reach the absurd climax without exhausting the table.
  3. The smaller map creates scarcity while preserving viable strategies.
  4. Default Game preserves satisfying spatial adaptation without a Era III interruption; Advanced Play earns its extra Realignment procedure.
  5. Research remains tense after players understand its probabilities.
  6. Audit creates meaningful exposure, including in Era IV.
  7. AGI is dramatic and optional rather than expected graduation.
  8. Infrastructure, Narrative, Trust, and Customer strategies can win.
  9. Faction powers create identity without deciding outcomes.
  10. The Future Timeline preserves an intelligible shared history.

Simulation-driven change control

Automated balance claims are governed by balance-and-exploitability.md and the balance and unified-matrix machine contracts. A normal tournament describes the tested field; it does not establish counterplay or qualify a rules change. The unified frame treats cooperation as an outcome, shrinks sparse cells, uses multiplicity-safe sequential intervals, and embeds only a diagnostic bounded adversarial slice. Promotion additionally requires clean provenance, a tracked receipt, one-lever common-seed evidence, and explicit user approval.

Before changing a rule, faction, action, card, objective, or score because of simulation evidence:

  1. Preserve the raw report under evidence/studies/simulation/.
  2. Add a tracked dated receipt with hashes and full configuration.
  3. Record the observation, hypothesis, alternatives, and validity limits.
  4. Audit rulebook, data, simulator, prototype, player aids, tests, and this protocol.
  5. Record exact changes or explicit no-change results for every surface.
  6. Validate the synchronized implementation.
  7. Commit the complete change with the study identifier.

An optimizer proposes a candidate. It never edits canonical rules.

Human session creation

After a synchronized release exists:

npm run playtest:new -- \
  --players 4 \
  --seed four-player-baseline-01 \
  --type facilitated_playtest

Physical components should carry a stable component ID and game version in small type. A mixed kit must say so instead of inheriting its newest component’s version.

Randomness and replay

A root seed is insufficient when RNG algorithm, call order, deck construction, or engine behavior changes. Reports therefore record the RNG contract plus ruleset, engine, and replay fingerprints.

Sampled replays preserve resolved public state and decisions. A newer engine must not infer what an older engine would have done.