Version: 2.1 Date: 2026-08-02 Owner: Design / Combat Status:
Evidence report. Companion to
03-RIGGING-AUDIT.md. States what has been measured, what
has not, and how much weight each number can carry.
[!important] One-line coverage summary All 17 live prototypes now have a bespoke driver and sit in band at n=200 (mean 58.7% win, 8.1 turns). Two need design attention rather than tuning: V04 is still structurally untunable, and V12's difficulty is a cliff. One number (V28) is a weak floor for driver reasons and should not be read as a verdict on its feature.
[!danger] Every balance number in version 1.0 of this document was wrong They were produced by a generic DOM-clicking bot that ignored each variant's signature mechanic. Enemy HP had been lowered until that bot could scrape a win. Re-measured with drivers that play the actual rule, the same builds scored 90–100% — V22 and V25 were unloseable. The v1.0 figures are struck, not adjusted. See §2.
[!warning] "Functionally healthy" is NOT "balance-tested" All 17 live prototypes pass the functional health check. That result means exactly one thing: the prototype does not crash. No JS errors, no unrecoverable deadlock, HP values move, and the fight terminates in Victory or Defeat.
It says nothing about whether the fight is winnable, whether it is winnable at the right rate, whether it lasts the right number of turns, or whether the feature under test does anything at all. Do not cite a health-check pass as evidence that a design works.
Balance-tested, in this document, means: 200 simulated games driven by a policy script that plays the prototype's actual rule through its own JS API, producing a win rate and a mean turn count, followed by tuning into the target band (35–80% win, 6–12 turns).
The v1.0 table rated nine shared-engine variants "medium-high confidence". That rating was too generous by a wide margin, and the failure is the same one that produced V01's fake 0% — just pointing the other way.
The generic bot places cards anywhere empty, clears whatever gate appears, rolls, and ends the turn. It never reads the omen, the stance, the enemy's declaration, the river's head, or the pressure dial. Tuning against it lowered enemy HP until a player who ignores the mechanic could win ~55% of the time. Hand the mechanic to a driver that uses it and the fights collapse:
| Prototype | v1.0 "tuned" | Same build, rules-aware driver |
|---|---|---|
| V12 Weather Die | 60.5% | 95.0% |
| V20 Blank Bargain | 54.0% | 90.0% |
| V22 Forecast | 62.0% | 100.0% |
| V23 Tributary | 56.0% | 95.0% |
| V24 Pressure Dial | 65.0% | 90.0% |
| V25 Gatehouse | 66.5% | 100.0% |
This is good news about the designs and bad news about the process: V22's omen slot and V25's gate pattern are strong levers — playing them correctly wins outright. The mechanics work. The numbers measuring them did not.
The Etch line's v1.0 figures (59.0 / 53.3 / 59.7) describe a
build that no longer exists. That snapshot had enemy HP 70,
energy max(1, emptySlots), and a lockedSlots()
that required a number and an element match. All three
have since changed. Those numbers are not comparable to anything current
and should not be cited.
| Prototype | Engine tier | Win rate | Turns | Enemy HP | Confidence | Note |
|---|---|---|---|---|---|---|
| V01 House Rules | bespoke | 53.5% | 9.96 | 58 | High | — |
| V04 Painted Ground | bespoke | 78.5% | 11.07 | 150 | Low — see §4 | HP is not a lever |
| V07 One Wick | bespoke | 61.5% | 7.33 | 57 | High | was 22.5% at HP 75 |
| V09 Their Dice | bespoke | 43.5% | 6.84 | 45 | High | — |
| V10 The Written Die | bespoke | 47.0% | 7.33 | 70 | High | — |
| V12 Weather Die | shared-engine | 58.0% | 7.42 | 37 | Medium — see §5 | cliff, not a slope |
| V13 Manthan | shared-engine | 48.0% | 7.65 | 60 | High | — |
| V14 Kalari Stance | shared-engine | 62.0% | 7.42 | 48 | High | — |
| V20 Blank Bargain | shared-engine | 59.5% | 8.57 | 44 | High | — |
| V21 Counterweight | shared-engine | 59.0% | 8.22 | 46 | High | — |
| V22 Forecast | shared-engine | 59.0% | 8.32 | 56 | High | — |
| V23 Tributary | shared-engine | 68.0% | 8.13 | 60 | High | — |
| V24 Pressure Dial | shared-engine | 64.5% | 7.43 | 47 | High | — |
| V25 Gatehouse | shared-engine | 60.5% | 8.08 | 51 | High | — |
| V26 Etch | Etch line | 67.5% | 7.88 | 70 | High | control: etching off → 0.0% |
| V27 Etch + Known Face | Etch line | 66.5% | 7.84 | 70 | Medium — see §6 | omen now used by the driver |
| V28 Etch + Element Styles | Etch line | 44.0% | 9.96 | 70 | Low — see §6 | weak floor, driver-limited |
Mean 58.7% win, 8.1 turns. All 17 inside the band.
Engine tiers. bespoke = V01, V04, V07, V09, V10, each its own codebase. shared-engine = V12, V13, V14, V20–V25. Etch line = V26–V28, generated from one template. Cross-tier comparison of raw throughput is not valid; compare within a tier.
The Scorched cap (SCORCH_AGE_CAP = 3) did real work: V04
came down from 93.3% to a band-legal figure. It did not
restore HP as a lever.
| Enemy max HP | Win rate | Turns |
|---|---|---|
| 150 | 75% | 10.81 |
| 170 | 72% | 11.67 |
| 185 | 71% | 12.31 |
| 200 | 71% | 12.95 |
+50 HP buys 4 points of win rate and 2.1 turns of length. Adding HP still adds duration, not difficulty. HP 150 is simply the only value where both metrics happen to sit in band, and it gets there by being short rather than by being balanced.
[!danger] Do not quote V04's 78.5% as a tuned result The cap bounded the quadratic escalation term. Something else in the model is still flat against HP. Until that is found and fixed, V04 is design-blocked, not tuned.
| Enemy HP | 34 | 36 | 37 | 38 | 40 | 42 | 46+ |
|---|---|---|---|---|---|---|---|
| Win rate | 97% | 72% | 58% | 43% | 25% | 10% | 0% |
Roughly 11 points of win rate per point of HP, against ~5 for its siblings. A usable window exists at 37, but ±2 HP leaves the band entirely.
A design with a difficulty window that narrow is fragile: any content
change — one card retuned, one intent adjusted — knocks it out, and a
human even slightly better or worse than the driver falls off one edge.
V12's energy formula (4 − occupied slots) couples the
player's economy directly to the thing they are optimising, which is the
likely cause. Flagged for design review.
V26's mechanic is load-bearing, proven by control. Running with etching disabled gives 0.0% win rate against 78% with it on, same driver, same HP. The mechanic carries the fight.
The engine is rolling, not board-building — and this is not obvious from the UI. An etched face fires the instant the die lands on it, so every roll is a damage event. Measured on V26 at n=100:
| Policy | Win rate |
|---|---|
| etch, then roll while affordable | 78–81% (flat for 0–4 slots filled) |
| etch, then roll only if nothing fired | 4–12% |
| no etch, roll while affordable | 0.0% |
Order matters just as much: etch before placing scores 81%, placing before etching 3.8% — same driver, same HP, only the sequence differs. Placing first empties the hand into slots, so the etch step finds nothing to write. If V26–V28 ever get a tutorial, this is the thing it has to teach.
V27's number is now feature-valid. The v1.0 figure
was flagged invalid because the policy ignored the revealed face. The
current driver uses S.omenN both ways — it etches onto the
revealed face when blank, and otherwise loads the slot that face names.
Confidence is Medium rather than High only because a scripted policy
still cannot represent how a human uses foresight.
[!warning] V28's 44.0% is a weak floor, not a verdict on element styles V28 is a strict superset of V26 and V27, so a lower score is suspicious on its face. The likely cause is the driver: V28's Jal style requires a two-step "pick an etched face, then pick a blank one" interaction that the driver almost certainly fumbles, wasting the action. Before concluding that element styles hurt, V28 needs human playtesting.
npm i playwright && npx playwright install chromium
node tools/prototype_drivers.js --games 200 --all
node tools/prototype_drivers.js --hp 56 V22-forecast.html # tuning sweep
node tools/prototype_drivers.js --json # machine-readable
tools/prototype_drivers.js holds one driver per
prototype. Each seeds Math.random, drives the prototype's
own JS API rather than its DOM, and plays that
variant's actual rule. V20–V25 are the same source behind a
cfg.id switch, so they share a driver that branches on
cfg.id exactly as applyVariant() does.
--hp rewrites the enemy-HP constant in a temporary copy and
asserts the substitution took effect; it never edits a file on disk.
The Etch line is generated — change
tools/build_etch_prototypes.py, regenerate, and verify with
--check. Never hand-edit V26–V28.
node tools/check_prototype_health.js V01-house-rules.html ...
ROUNDS=420 node tools/check_prototype_health.js V07-one-wick.html
All 17 pass. See §1 for what that does and does not mean.
Not a measurement problem, but it surfaced while writing the drivers and belongs on the record.
In all six V20–V25 builds, turn income is a flat 2, the first roll costs 1, and a reroll costs 2. A reroll is therefore never affordable — not rarely, never, on any turn.
For a project whose thesis is rigging the dice, six of seventeen prototypes price the rigging out of reach entirely. The dice in those six are pure weather, and the only lever the player actually holds is where to place before the dice are seen. Every policy in the shared driver is consequently a placement policy.
Compare V12, whose income is 4 − occupied slots: that is
a real trade-off (a fuller board means more targets and fewer rerolls),
and it is the reason V12 needed its own driver.
This interacts with charter C6 ("rerolls cost") and
with the C9 amendment proposed in 03-RIGGING-AUDIT.md.
Recommend either raising income or lowering reroll cost in V20–V25, then
re-measuring — the current HP values assume a player who never rerolls,
because none can.
Method: take the competent driver, randomise exactly one decision, leave everything else intact, re-measure at n=120. The drop is the price of not mastering that decision. This is the V26 etching-off control generalised, and it deliberately avoids a naive-bot comparison, which cannot separate "this decision matters" from "the bot could not operate the UI".
[!note] Reading the numbers At n=120 the standard error on a ~55% rate is ~4.5 points, so a difference of under ~13 points is not distinguishable from zero. Treat those as no effect.
An exactly 0.0 result is different in kind and is not noise: it means the ablated branch never executed in any of 120 games, i.e. the decision is structurally unreachable.
| Prototype | base | reroll | targeting | card choice | signature feature | order |
|---|---|---|---|---|---|---|
| V01 House Rules | 53.3% | 7.5 | — | 5.0 | — | — |
| V04 Painted Ground | 76.7% | 0.0 | — | 5.9 | — | — |
| V07 One Wick | 60.0% | 19.2 | 36.7 | — | 40.0 (colour) | — |
| V09 Their Dice | 42.5% | 28.3 | — | — | 27.5 (declaration) | — |
| V10 The Written Die | 45.0% | 25.0 | — | — | 45.0 → 0% (inscribe) | — |
| V12 Weather Die | 55.0% | 0.0 | — | — | -4.2 (budget) | — |
| V13 Manthan | 45.8% | -7.5 | — | — | -7.5 (memory) | — |
| V14 Kalari Stance | 60.8% | 0.0 | — | — | 0.0 (stance) | — |
| V20 Blank Bargain | 55.0% | 0.0 | -8.3 | 36.7 | — | — |
| V21 Counterweight | 59.2% | 0.0 | 0.0 | 28.4 | — | — |
| V22 Forecast | 55.0% | 0.0 | 39.2 | 27.5 | — | — |
| V23 Tributary | 65.8% | 0.0 | 18.3 | 31.6 | — | — |
| V24 Pressure Dial | 61.7% | 0.0 | -9.1 | 40.9 | — | — |
| V25 Gatehouse | 62.5% | 0.0 | -2.5 | 30.8 | — | — |
| V26 Etch | 65.0% | 0.0 | — | — | 65.0 → 0% (etch) | 15.0 |
| V27 Etch + Known Face | 64.2% | 0.0 | — | — | 64.2 → 0% (etch) | 17.5 |
| V28 Etch + Element Styles | 40.0% | 0.0 | — | — | 40.0 → 0% (etch) | -17.5 |
Etch-line roll-aggression axis: V26 63.3, V27 61.7, V28 35.8. Known-face (omen) axis: V27 -1.6, V28 -1.7.
reroll returns exactly 0.0 on V12, V13,
V14 and all of V20–V25. The branch never fires in 120 games because
income never covers the cost: a flat 2 (or 4 − occupied), a
first roll at 1, and a reroll at 2. This is broader than §9 recorded —
it is not six prototypes, it is nine.
For a project whose thesis is rigging the dice, over half the lab cannot rig. Where the reroll is affordable it is immediately one of the deepest decisions available (V09 28.3, V10 25.0, V07 19.2), which is the strongest possible argument for fixing the economy in the other nine.
V14 Kalari Stance scores exactly 0.0 on both axes. The enemy announces its stance every turn, which is good design intent — but acting on it requires a reroll the player can never afford. The information is real and completely inert.
V12 and V13 likewise show no axis above noise; V13's two axes are negative, meaning the "chase a PURE relation" policy is no better than random. Either the wheel does not reward what the design says it rewards, or the driver's policy is wrong. Flagged rather than concluded.
For V20, V21, V24 and V25, the variant's own
lever (targeting) measures at or below zero, while
generic card choice measures 28–41 points. Strip the
theming and what the player is actually doing is "play your biggest
card" — the shared engine, not the variant.
Only V22 Forecast (39.2) and V23 Tributary (18.3) have targeting levers that carry real weight. Those two are worth keeping from the expansion wave on current evidence.
Ablating the omen changes V27 by -1.6 and V28 by -1.7 — both inside noise. V27's base (64.2%) is also statistically identical to V26's (65.0%). On present evidence V27 is V26 with a decoration. The caveat in §6 stands but sharpens: this is now a measured null, not an untested feature. If foresight is meant to matter, it needs a stronger hook than a revealed face, and a human playtest should confirm before the idea is dropped.
| Rank | Prototype | Deepest decision | Depth | Character |
|---|---|---|---|---|
| 1 | V26 / V27 Etch | etching | 65 / 64 pts → 0% | three real decisions (etch, order, roll aggression) |
| 2 | V10 Written Die | inscribing | 45 pts → 0% | two real decisions |
| 3 | V07 One Wick | colour | 40 pts | three real decisions — the richest spread |
| 4 | V22 Forecast | targeting | 39 pts | one strong lever + card choice |
| 5 | V09 Their Dice | reroll / declaration | 28 pts | two real decisions |
V07 is the most interesting result in the study: three independent decisions worth 40, 37 and 19 points. No other prototype spreads its depth that evenly, and it is the only one where the rigging lever, the targeting lever and the signature mechanic all measure.
Hard to master: five of seventeen (V07, V09, V10, V26, V27), plus V22 and V23 partially. Six are shallow (V01, V04, V20, V21, V24, V25) — one trivial decision, "play the big card". Three have nothing to master (V12, V13, V14).
Easy to understand is still unmeasured. No human has rated comprehension; the survey holds only QA test rows. Depth and comprehension are different axes, and V26 illustrates the gap: it has the most depth in the lab and its central rule — an etched face fires the instant the die lands on it, so rolls are the damage — appears nowhere in the interface. It took a policy bake-off to find. High depth, probably poor comprehension.