← Back to the Lab

Combat Variants — Balance Test Coverage

Version: 2.1 Date: 2026-08-02 Owner: Design / Combat Status: Evidence report. Companion to 03-RIGGING-AUDIT.md. States what has been measured, what has not, and how much weight each number can carry.

[!important] One-line coverage summary All 17 live prototypes now have a bespoke driver and sit in band at n=200 (mean 58.7% win, 8.1 turns). Two need design attention rather than tuning: V04 is still structurally untunable, and V12's difficulty is a cliff. One number (V28) is a weak floor for driver reasons and should not be read as a verdict on its feature.

[!danger] Every balance number in version 1.0 of this document was wrong They were produced by a generic DOM-clicking bot that ignored each variant's signature mechanic. Enemy HP had been lowered until that bot could scrape a win. Re-measured with drivers that play the actual rule, the same builds scored 90–100% — V22 and V25 were unloseable. The v1.0 figures are struck, not adjusted. See §2.


1. The distinction that matters most

[!warning] "Functionally healthy" is NOT "balance-tested" All 17 live prototypes pass the functional health check. That result means exactly one thing: the prototype does not crash. No JS errors, no unrecoverable deadlock, HP values move, and the fight terminates in Victory or Defeat.

It says nothing about whether the fight is winnable, whether it is winnable at the right rate, whether it lasts the right number of turns, or whether the feature under test does anything at all. Do not cite a health-check pass as evidence that a design works.

Balance-tested, in this document, means: 200 simulated games driven by a policy script that plays the prototype's actual rule through its own JS API, producing a win rate and a mean turn count, followed by tuning into the target band (35–80% win, 6–12 turns).


2. Why the old numbers were wrong

The v1.0 table rated nine shared-engine variants "medium-high confidence". That rating was too generous by a wide margin, and the failure is the same one that produced V01's fake 0% — just pointing the other way.

The generic bot places cards anywhere empty, clears whatever gate appears, rolls, and ends the turn. It never reads the omen, the stance, the enemy's declaration, the river's head, or the pressure dial. Tuning against it lowered enemy HP until a player who ignores the mechanic could win ~55% of the time. Hand the mechanic to a driver that uses it and the fights collapse:

Prototype v1.0 "tuned" Same build, rules-aware driver
V12 Weather Die 60.5% 95.0%
V20 Blank Bargain 54.0% 90.0%
V22 Forecast 62.0% 100.0%
V23 Tributary 56.0% 95.0%
V24 Pressure Dial 65.0% 90.0%
V25 Gatehouse 66.5% 100.0%

This is good news about the designs and bad news about the process: V22's omen slot and V25's gate pattern are strong levers — playing them correctly wins outright. The mechanics work. The numbers measuring them did not.

The Etch line's v1.0 figures (59.0 / 53.3 / 59.7) describe a build that no longer exists. That snapshot had enemy HP 70, energy max(1, emptySlots), and a lockedSlots() that required a number and an element match. All three have since changed. Those numbers are not comparable to anything current and should not be cited.


3. At a glance — current state (n=200, 2026-08-02)

Prototype Engine tier Win rate Turns Enemy HP Confidence Note
V01 House Rules bespoke 53.5% 9.96 58 High —
V04 Painted Ground bespoke 78.5% 11.07 150 Low — see §4 HP is not a lever
V07 One Wick bespoke 61.5% 7.33 57 High was 22.5% at HP 75
V09 Their Dice bespoke 43.5% 6.84 45 High —
V10 The Written Die bespoke 47.0% 7.33 70 High —
V12 Weather Die shared-engine 58.0% 7.42 37 Medium — see §5 cliff, not a slope
V13 Manthan shared-engine 48.0% 7.65 60 High —
V14 Kalari Stance shared-engine 62.0% 7.42 48 High —
V20 Blank Bargain shared-engine 59.5% 8.57 44 High —
V21 Counterweight shared-engine 59.0% 8.22 46 High —
V22 Forecast shared-engine 59.0% 8.32 56 High —
V23 Tributary shared-engine 68.0% 8.13 60 High —
V24 Pressure Dial shared-engine 64.5% 7.43 47 High —
V25 Gatehouse shared-engine 60.5% 8.08 51 High —
V26 Etch Etch line 67.5% 7.88 70 High control: etching off → 0.0%
V27 Etch + Known Face Etch line 66.5% 7.84 70 Medium — see §6 omen now used by the driver
V28 Etch + Element Styles Etch line 44.0% 9.96 70 Low — see §6 weak floor, driver-limited

Mean 58.7% win, 8.1 turns. All 17 inside the band.

Engine tiers. bespoke = V01, V04, V07, V09, V10, each its own codebase. shared-engine = V12, V13, V14, V20–V25. Etch line = V26–V28, generated from one template. Cross-tier comparison of raw throughput is not valid; compare within a tier.


4. V04 Painted Ground — still a design problem

The Scorched cap (SCORCH_AGE_CAP = 3) did real work: V04 came down from 93.3% to a band-legal figure. It did not restore HP as a lever.

Enemy max HP Win rate Turns
150 75% 10.81
170 72% 11.67
185 71% 12.31
200 71% 12.95

+50 HP buys 4 points of win rate and 2.1 turns of length. Adding HP still adds duration, not difficulty. HP 150 is simply the only value where both metrics happen to sit in band, and it gets there by being short rather than by being balanced.

[!danger] Do not quote V04's 78.5% as a tuned result The cap bounded the quadratic escalation term. Something else in the model is still flat against HP. Until that is found and fixed, V04 is design-blocked, not tuned.


5. V12 Weather Die — a cliff, not a slope

Enemy HP 34 36 37 38 40 42 46+
Win rate 97% 72% 58% 43% 25% 10% 0%

Roughly 11 points of win rate per point of HP, against ~5 for its siblings. A usable window exists at 37, but ±2 HP leaves the band entirely.

A design with a difficulty window that narrow is fragile: any content change — one card retuned, one intent adjusted — knocks it out, and a human even slightly better or worse than the driver falls off one edge. V12's energy formula (4 − occupied slots) couples the player's economy directly to the thing they are optimising, which is the likely cause. Flagged for design review.


6. The Etch line

V26's mechanic is load-bearing, proven by control. Running with etching disabled gives 0.0% win rate against 78% with it on, same driver, same HP. The mechanic carries the fight.

The engine is rolling, not board-building — and this is not obvious from the UI. An etched face fires the instant the die lands on it, so every roll is a damage event. Measured on V26 at n=100:

Policy Win rate
etch, then roll while affordable 78–81% (flat for 0–4 slots filled)
etch, then roll only if nothing fired 4–12%
no etch, roll while affordable 0.0%

Order matters just as much: etch before placing scores 81%, placing before etching 3.8% — same driver, same HP, only the sequence differs. Placing first empties the hand into slots, so the etch step finds nothing to write. If V26–V28 ever get a tutorial, this is the thing it has to teach.

V27's number is now feature-valid. The v1.0 figure was flagged invalid because the policy ignored the revealed face. The current driver uses S.omenN both ways — it etches onto the revealed face when blank, and otherwise loads the slot that face names. Confidence is Medium rather than High only because a scripted policy still cannot represent how a human uses foresight.

[!warning] V28's 44.0% is a weak floor, not a verdict on element styles V28 is a strict superset of V26 and V27, so a lower score is suspicious on its face. The likely cause is the driver: V28's Jal style requires a two-step "pick an etched face, then pick a blank one" interaction that the driver almost certainly fumbles, wasting the action. Before concluding that element styles hurt, V28 needs human playtesting.


7. How to reproduce

Balance measurement

npm i playwright && npx playwright install chromium
node tools/prototype_drivers.js --games 200 --all
node tools/prototype_drivers.js --hp 56 V22-forecast.html     # tuning sweep
node tools/prototype_drivers.js --json                        # machine-readable

tools/prototype_drivers.js holds one driver per prototype. Each seeds Math.random, drives the prototype's own JS API rather than its DOM, and plays that variant's actual rule. V20–V25 are the same source behind a cfg.id switch, so they share a driver that branches on cfg.id exactly as applyVariant() does. --hp rewrites the enemy-HP constant in a temporary copy and asserts the substitution took effect; it never edits a file on disk.

The Etch line is generated — change tools/build_etch_prototypes.py, regenerate, and verify with --check. Never hand-edit V26–V28.

Functional health check (proves no crash, nothing more)

node tools/check_prototype_health.js V01-house-rules.html ...
ROUNDS=420 node tools/check_prototype_health.js V07-one-wick.html

All 17 pass. See §1 for what that does and does not mean.


8. Confidence caveats — read before quoting any number here


9. Open design issue — the reroll economy in V20–V25

Not a measurement problem, but it surfaced while writing the drivers and belongs on the record.

In all six V20–V25 builds, turn income is a flat 2, the first roll costs 1, and a reroll costs 2. A reroll is therefore never affordable — not rarely, never, on any turn.

For a project whose thesis is rigging the dice, six of seventeen prototypes price the rigging out of reach entirely. The dice in those six are pure weather, and the only lever the player actually holds is where to place before the dice are seen. Every policy in the shared driver is consequently a placement policy.

Compare V12, whose income is 4 − occupied slots: that is a real trade-off (a fuller board means more targets and fewer rerolls), and it is the reason V12 needed its own driver.

This interacts with charter C6 ("rerolls cost") and with the C9 amendment proposed in 03-RIGGING-AUDIT.md. Recommend either raising income or lowering reroll cost in V20–V25, then re-measuring — the current HP values assume a player who never rerolls, because none can.


10. Depth study — is each system hard to master?

Method: take the competent driver, randomise exactly one decision, leave everything else intact, re-measure at n=120. The drop is the price of not mastering that decision. This is the V26 etching-off control generalised, and it deliberately avoids a naive-bot comparison, which cannot separate "this decision matters" from "the bot could not operate the UI".

[!note] Reading the numbers At n=120 the standard error on a ~55% rate is ~4.5 points, so a difference of under ~13 points is not distinguishable from zero. Treat those as no effect.

An exactly 0.0 result is different in kind and is not noise: it means the ablated branch never executed in any of 120 games, i.e. the decision is structurally unreachable.

Depth per decision (win-rate points lost when that decision is randomised)

Prototype base reroll targeting card choice signature feature order
V01 House Rules 53.3% 7.5 — 5.0 — —
V04 Painted Ground 76.7% 0.0 — 5.9 — —
V07 One Wick 60.0% 19.2 36.7 — 40.0 (colour) —
V09 Their Dice 42.5% 28.3 — — 27.5 (declaration) —
V10 The Written Die 45.0% 25.0 — — 45.0 → 0% (inscribe) —
V12 Weather Die 55.0% 0.0 — — -4.2 (budget) —
V13 Manthan 45.8% -7.5 — — -7.5 (memory) —
V14 Kalari Stance 60.8% 0.0 — — 0.0 (stance) —
V20 Blank Bargain 55.0% 0.0 -8.3 36.7 — —
V21 Counterweight 59.2% 0.0 0.0 28.4 — —
V22 Forecast 55.0% 0.0 39.2 27.5 — —
V23 Tributary 65.8% 0.0 18.3 31.6 — —
V24 Pressure Dial 61.7% 0.0 -9.1 40.9 — —
V25 Gatehouse 62.5% 0.0 -2.5 30.8 — —
V26 Etch 65.0% 0.0 — — 65.0 → 0% (etch) 15.0
V27 Etch + Known Face 64.2% 0.0 — — 64.2 → 0% (etch) 17.5
V28 Etch + Element Styles 40.0% 0.0 — — 40.0 → 0% (etch) -17.5

Etch-line roll-aggression axis: V26 63.3, V27 61.7, V28 35.8. Known-face (omen) axis: V27 -1.6, V28 -1.7.

Finding 1 — the reroll is unreachable in NINE of seventeen

reroll returns exactly 0.0 on V12, V13, V14 and all of V20–V25. The branch never fires in 120 games because income never covers the cost: a flat 2 (or 4 − occupied), a first roll at 1, and a reroll at 2. This is broader than §9 recorded — it is not six prototypes, it is nine.

For a project whose thesis is rigging the dice, over half the lab cannot rig. Where the reroll is affordable it is immediately one of the deepest decisions available (V09 28.3, V10 25.0, V07 19.2), which is the strongest possible argument for fixing the economy in the other nine.

Finding 2 — three systems have no decisions at all

V14 Kalari Stance scores exactly 0.0 on both axes. The enemy announces its stance every turn, which is good design intent — but acting on it requires a reroll the player can never afford. The information is real and completely inert.

V12 and V13 likewise show no axis above noise; V13's two axes are negative, meaning the "chase a PURE relation" policy is no better than random. Either the wheel does not reward what the design says it rewards, or the driver's policy is wrong. Flagged rather than concluded.

Finding 3 — four expansion-wave variants are the same game wearing hats

For V20, V21, V24 and V25, the variant's own lever (targeting) measures at or below zero, while generic card choice measures 28–41 points. Strip the theming and what the player is actually doing is "play your biggest card" — the shared engine, not the variant.

Only V22 Forecast (39.2) and V23 Tributary (18.3) have targeting levers that carry real weight. Those two are worth keeping from the expansion wave on current evidence.

Finding 4 — the known-face feature does nothing measurable

Ablating the omen changes V27 by -1.6 and V28 by -1.7 — both inside noise. V27's base (64.2%) is also statistically identical to V26's (65.0%). On present evidence V27 is V26 with a decoration. The caveat in §6 stands but sharpens: this is now a measured null, not an untested feature. If foresight is meant to matter, it needs a stronger hook than a revealed face, and a human playtest should confirm before the idea is dropped.

Finding 5 — the deep systems, ranked

Rank Prototype Deepest decision Depth Character
1 V26 / V27 Etch etching 65 / 64 pts → 0% three real decisions (etch, order, roll aggression)
2 V10 Written Die inscribing 45 pts → 0% two real decisions
3 V07 One Wick colour 40 pts three real decisions — the richest spread
4 V22 Forecast targeting 39 pts one strong lever + card choice
5 V09 Their Dice reroll / declaration 28 pts two real decisions

V07 is the most interesting result in the study: three independent decisions worth 40, 37 and 19 points. No other prototype spreads its depth that evenly, and it is the only one where the rigging lever, the targeting lever and the signature mechanic all measure.

Answer to "easy to understand, hard to master?"

Hard to master: five of seventeen (V07, V09, V10, V26, V27), plus V22 and V23 partially. Six are shallow (V01, V04, V20, V21, V24, V25) — one trivial decision, "play the big card". Three have nothing to master (V12, V13, V14).

Easy to understand is still unmeasured. No human has rated comprehension; the survey holds only QA test rows. Depth and comprehension are different axes, and V26 illustrates the gap: it has the most depth in the lab and its central rule — an etched face fires the instant the die lands on it, so rolls are the damage — appears nowhere in the interface. It took a policy bake-off to find. High depth, probably poor comprehension.


Changelog