ADR 0054: Use Player-System Hypotheses, Not Fun Scores
Status
Accepted.
Context
ADRs 0052 and 0053 moved gameplay intent and direct play to human-controlled gates. They did not yet give agents and designers a sufficiently concrete way to reason about systems-heavy games between those gates.
Systems-heavy prototypes can be mechanically complete while the player cannot read the state, perceive viable alternatives, connect action to consequence, form a mental model, or find a reason to repeat the loop. Lanternworks and Followspot demonstrated this failure. Their internal architecture and automated routes remained useful engineering evidence, but did not create a comprehensible or enjoyable product.
Established research offers useful constructs for motivation, challenge, choice, interface design, game feel, learning, accessibility, telemetry, and playtesting. The evidence does not support one universal formula, checklist, or scalar score for fun. Many constructs are experienced and must be measured through people rather than inferred from source or simulation.
Decision
For materially new or changed systems-heavy gameplay:
- Start from the human-approved product promise, target player, complexity budget, and intended experience established by ADR 0053.
- Trace the repeated player-system loop as
perceive -> decide -> act -> read consequence -> learn -> adapt. - Treat UI, feedback, onboarding, pacing, progression, balance, accessibility, and ethics as parts of that loop rather than downstream polish categories.
- Build the smallest complete decision through production input, presentation, consequence, explanation, recovery, and continuation.
- Use automation for rules, state, curves, distributions, reachability, latency, feedback coverage, presentation mechanics, and encoded-policy outcomes.
- Label predictions about meaningful choice, clarity, pacing, balance, learning, accessibility, engagement, or enjoyment as hypotheses until appropriate human evidence exists.
- Observe owner play without coaching and preserve observation, player report, interpretation, proposed change, and retest result separately.
- Require representative target players when the claim widens beyond the owner, first-party designer, or specifically observed participant.
- Do not collapse telemetry, retention, session time, theory checks, simulation, or human reports into a generic fun score.
- Treat ethical player control and accessibility barriers as design concerns during the loop, not only release polish.
This is conditional practitioner guidance. Bounded fixes and faithful implementation of settled gameplay do not require a new systems-design brief.
Consequences
- Agents gain concrete questions and mechanical work without receiving product authority they cannot legitimately exercise.
- UI and feedback must support player reasoning before system breadth is treated as complete.
- Progression and balance are evaluated against a declared design goal rather than equal numbers or activity metrics by default.
- Human tests happen around a complete decision, so rejected hypotheses are cheaper to revise.
- Telemetry and simulation retain high diagnostic value while their claim boundary remains explicit.
- The workflow costs human attention earlier, but reduces the risk of multiplying a coherent, well-tested, unwanted game.
Evidence Boundary
The decision synthesizes MDA, GameFlow, self-determination theory, the Player Experience Inventory, direct-manipulation and ecological-interface research, game-feel and onboarding studies, accessibility guidance, games-user-research practice, and practitioner postmortems. Those sources support individual lenses and methods. They do not validate this exact STAGE sequence or establish one optimal design process across genres and players.
The local evidence remains negative and same-owner: two mechanically healthy dogfoods failed product review, while The Circussy One's repeated human-agent correction loop produced accepted changes at substantial correction cost. A future trial beginning from an approved player-system hypothesis is required before STAGE can claim positive product evidence for this sequence.