Skip to main content

Systems-Heavy Gameplay Evidence

Systems-heavy games ask players to understand and steer interacting rules over time. Examples include strategy games, management games, builders, automation games, roguelites, deckbuilders, simulations, and economy-heavy survival games. Their appeal can come from discovery, planning, adaptation, expression, mastery, dramatic recovery, or watching a personally shaped system unfold.

STAGE does not treat those qualities as automatic consequences of having many systems. A coherent simulation can still be opaque, low-agency, repetitive, or emotionally flat. This guide turns the human-led gameplay design decision into a testable player-system hypothesis without pretending to calculate fun. It does not provide a universal formula for fun.

This is practitioner guidance. It synthesizes research and postmortems, but the exact sequence has not been established as a universal scientific model.

The Product Promise Comes First

Before specifying an economy, progression tree, event schedule, or information panel, state why operating this system should be worth a player's attention. The human designer chooses:

  • the player role and fantasy;
  • the intended forms of agency, tension, expression, and mastery;
  • the target player, prior knowledge, session context, and complexity budget;
  • the emotional arc and acceptable kinds of frustration;
  • what should remain uncertain, surprising, or deliberately unexplained; and
  • the ethical boundary around time, spending, compulsion, and player control.

Mechanics are candidates for producing that promise. Their existence is not evidence that the promise is present.

Tynan Sylvester's RimWorld postmortem provides a useful example of a strong design pillar: the product is framed as a story generator, and mechanics, failure, recovery, visual hierarchy, and scope are selected for the stories and emotions they generate. That does not prescribe RimWorld's solution for other games. It demonstrates how a product promise can constrain systems instead of being inferred from them afterward.

The Player-System Evidence Loop

Use this inner loop to reason about the experience:

Promise
|
Perceive -> Decide -> Act -> Read consequence -> Learn -> Adapt
^ |
+---------------- changed situation -----------------+

The corresponding development loop remains human-led:

Human intent -> Mechanical model -> Thin real slice -> Human play
^ |
+------------- revise, tune, or reject --------------+
|
Representative study when claims widen

The inner loop is healthy when the player can:

  1. Perceive the current goal, material state, pressure, opportunities, and constraints.
  2. Decide among alternatives that appear viable and differ in consequences relevant to the player's goals.
  3. Act through controls that respect the intended command without unnecessary friction.
  4. Read consequence at the immediate, causal, strategic, and emotional levels appropriate to the action.
  5. Learn why the result occurred well enough to update a mental model.
  6. Adapt by applying that model to a changed problem rather than repeating a solved chore.

The loop is a diagnostic model, not a score. A game can intentionally obscure information, delay consequences, constrain agency, or create routine. Those choices need to serve the approved experience rather than arise accidentally.

Motivation And Challenge Are Hypotheses

A systems-heavy loop needs a reason to operate it beyond the existence of rewards or content. Name the expected motive rather than assuming every player wants the same thing. Useful lenses include:

  • Autonomy: volition, ownership, and expression through meaningful control, not merely a large option count.
  • Competence: becoming more effective through learnable challenge, intelligible feedback, and skill or knowledge that transfers.
  • Relatedness: connection, care, cooperation, rivalry, or responsibility toward other people or represented beings when the game supports it.
  • Curiosity and mastery: uncertainty the player can reduce by observing, experimenting, predicting, and revising a mental model.
  • Identity and authorship: a build, deck, city, factory, story, or strategy that visibly becomes the player's own.
  • Rhythm and restoration: satisfying routine, flow, collection, or contemplation when those are part of the approved promise.

The self-determination theory game studies associate experienced autonomy, competence, and relatedness with enjoyment and future play. The reducible-uncertainty account of challenge offers another useful lens for why learning and prediction can be rewarding. Neither establishes a universal recipe, target value, or requirement that one loop satisfy every motive.

Declare the relevant kind of challenge: execution, planning, inference, resource pressure, risk, timing, coordination, or another approved demand. Increasing health, speed, prices, timers, or system count does not by itself produce competence. The player needs enough causality, control, and recovery to interpret failure and attempt a better or more expressive response.

Agents can inspect available control, challenge curves, error recovery, progress signals, and opportunities for different strategies. Only human evidence establishes felt volition, competence, connection, curiosity, ownership, pressure, or desire to continue. Extrinsic rewards and retention can change behavior without demonstrating any of those experiences.

Meaningful Decisions

For a repeated decision, inspect these properties:

PropertyQuestion
AvailableAre at least two alternatives actually possible in the current state?
PerceivedCan the intended player discover and distinguish them in time?
ViableIs more than one alternative credible under some relevant state, goal, or strategy?
ConsequentialDo the alternatives change something the player currently values?
ContextualCan the preferred option change with resources, timing, risk, build, location, or intent?
CostlyIs there an opportunity cost, commitment, risk, delay, or foregone alternative?
LegibleCan the player connect the choice to its important outcomes?
LearnableDoes feedback help the player make a better or more expressive future decision?

Not every choice needs every property. A repeated core decision should not collapse accidentally into one obvious answer, an unreadable guess, or an action whose consequences do not matter.

Agents can enumerate options, calculate expected values, search encoded policies, find unreachable or dominated candidates, and compare sensitivity under declared assumptions. They cannot establish which alternatives a person perceives, whether a tradeoff feels meaningful, or whether committing to it is satisfying.

For narrative choices, consequence, sociality, and moral relevance can affect perceived meaningfulness. For non-narrative systems, use that finding only as a prompt to examine what the decision changes for the player; do not generalize a single narrative study into a universal choice formula.

Feedback And Game Feel

Feedback in a systems-heavy game has several jobs:

LayerJobExamples
InputConfirm that the command was receivedpress, click, cursor response, command marker
OutcomeShow the immediate resultmovement, hit, purchase, placement, production pulse
CausalityExplain what changed and whymodifiers, blockers, source attribution, failure reason
StrategyShow how future possibilities changedforecast, range, queue, capacity, threat, route
EmotionCommunicate importance and tonemotion, sound, haptics, hit stop, celebration, alarm
PersistenceLet the player recover or compare laterhistory, log, tooltip, ledger, graph, replayable help

Calibrate intensity to consequence. Feedback that is too weak makes the system feel unresponsive or inexplicable. Feedback that is always maximal destroys hierarchy, raises fatigue, hides important events, and can reduce experience quality. A large action should usually receive more emphasis than a routine tick, and repeated low-value events should aggregate or yield to higher-priority signals.

Research on "juiciness" supports visual and tactile embellishment as a useful tool, not an unlimited good. A study with 3,018 action-RPG players found that both no juice and extreme juice underperformed medium or high conditions on several experience and behavior measures. Treat feedback dosage, timing, and semantic priority as design variables.

For each important event, verify:

  • a production event actually reaches every intended feedback channel;
  • the channels agree about timing, magnitude, source, and state;
  • feedback pauses, resets, and tears down with the owning lifecycle;
  • critical information is not encoded through one sensory channel alone;
  • frequent events have budgets, aggregation, or priority rules; and
  • the player can inspect a durable explanation when transient feedback is not enough.

Frame time and input latency are part of feedback quality. An asynchronous method signature, queued animation, or emitted event does not prove that the player saw a timely response.

UI As A Decision Surface

The purpose of systems-heavy UI is not to display every variable. It is to help the player perceive, compare, predict, act, explain, and recover.

Prefer:

  • current state beside the decision it affects;
  • visible constraints and causal relationships, not totals without meaning;
  • comparison views that expose costs, deltas, opportunity costs, and consequences;
  • stable vocabulary, units, iconography, color semantics, and ordering;
  • direct manipulation with rapid, incremental, reversible effects where the fantasy and platform support it;
  • progressive disclosure by decision context, with details available on demand;
  • explicit reasons for disabled, blocked, invalid, or inaccessible actions;
  • histories, forecasts, overlays, and tooltips when they reduce mental bookkeeping;
  • one discoverable help channel with contextual links instead of scattered tutorial messages;
  • mouse, keyboard, controller, touch, and assistive navigation parity for the inputs the product supports; and
  • pause, slowdown, review, or confirmation where time pressure would otherwise turn comprehension into an input-speed test.

Avoid:

  • exposing internal implementation values or units without player meaning;
  • hiding the decisive comparison behind several unrelated panels;
  • treating more labels, charts, or tooltips as automatically more legible;
  • generic failure messages when a specific causal reason is available;
  • a visual hierarchy in which routine data competes with urgent state;
  • hover-only information on controller or touch paths;
  • UI focus that diverges from the visible selection;
  • tutorial text that explains a mechanic before the player has a reason to care; and
  • assuming diegetic or integrated UI is inherently more immersive or usable than an overlay.

Ecological interface design is a useful lens: expose relationships and constraints that remain valid across situations so players can reason under novel conditions. Direct-manipulation research is another useful lens: continuous representation, concrete actions, reversibility, and immediate visible effects can reduce the gap between intention and system response. Neither lens requires one visual style.

Learning And Onboarding

Onboarding is successful when the player forms a useful mental model and can act on it, not when the tutorial has displayed every instruction.

For a systems-heavy first session:

  1. Represent the real game immediately rather than delaying its central verb.
  2. Teach one core concept through a meaningful goal with minimal surrounding complexity.
  3. Let the player act, observe a consequence, and vary the action.
  4. Teach concepts and constraints rather than a memorized exact solution.
  5. Offer recovery and experimentation so one misunderstanding does not end the learning opportunity.
  6. Introduce complexity when the current decision creates a reason to need it.
  7. Keep explanations searchable and replayable after the transient prompt.
  8. Test comprehension without developer coaching before interpreting the result.

Factorio's New Player Experience retrospective is a useful negative case: an older tutorial delayed automation, constrained players into one solution, disconnected levels, created purposeless grind, and scattered too much information. Its redesign aimed to represent the game early, teach central concepts with meaningful goals, support experimentation, and unify information channels.

Tutorial presence alone is not a usability result. Studies of first-time user experience and implicit tutorials show context-dependent effects and different responses between novice and experienced players. Test the target players and the actual first-session route.

Useful first-session observations include:

  • time to infer the immediate goal;
  • time to the first meaningful decision and first completed core loop;
  • actions attempted before success;
  • requests for explanation;
  • repeated errors or unrecognized state;
  • whether the player can predict the next consequence;
  • whether the player can explain the system in their own words; and
  • whether they choose to repeat or vary the loop.

These observations do not explain themselves. Preserve what happened, the player's retrospective account, and the team's interpretation as separate records.

Pacing, Progression, And Rewards

Systems-heavy games operate on several clocks:

ClockTypical concern
ResponseTime from input to perceivable acknowledgment
DecisionTime and attention between consequential choices
Encounter or cyclePressure, escalation, payoff, recovery
Session or runBuild formation, reversals, climax, closure
Long-term progressionNew possibilities, mastery, identity, return

Do not optimize all clocks toward speed. Deliberation, anticipation, routine, and recovery can be valuable. Diagnose accidental dead time, forced waiting, unbroken pressure, decision floods, and long stretches in which the player's best action is already solved.

Good progression usually changes the decision space as well as the numbers. It may:

  • unlock a new verb, constraint, synergy, or counterplay relationship;
  • make an earlier decision more expressive;
  • create a new tension instead of merely removing an old one;
  • let a build, city, deck, factory, or strategy gain recognizable identity;
  • convert prior learning into new leverage; and
  • provide short-term payoff while preserving future agency.

Reward speed affects perceived competence and enjoyment even when objective performance looks similar. Points, levels, and progress indicators can improve task performance without necessarily improving autonomy, competence, or intrinsic motivation. Treat reward cadence and power curves as hypotheses to playtest, not motivational laws.

Randomness is useful when it creates adaptation, surprise, and varied planning. It becomes weak when the player cannot understand odds, influence exposure, recover from bad outcomes, or connect the result to a meaningful response.

Balance And Emergence

Define the balance goal before tuning. Possible goals include:

  • multiple options having a legitimate place;
  • distinct strategy identities;
  • no option warping the rest of the system;
  • fair counterplay or recovery;
  • readable risk and reward;
  • interesting adaptation across contexts; and
  • difficulty appropriate to the target player and intended experience.

Equal pick rates, win rates, damage, or resource efficiency are not universal balance goals. Distinguish:

  • actual power under declared conditions;
  • perceived power;
  • aesthetic and thematic appeal;
  • ease of learning and execution;
  • frequency and context of availability;
  • cost, risk, counterplay, and opportunity cost; and
  • the fun or expressiveness of using the option.

Simulation and telemetry can reveal exploits, sensitivity, dead content, candidate dominant policies, strategy distributions, bottlenecks, and unexpected interactions. Use several encoded objectives or procedural personas rather than one optimizer when the game supports different player motivations. The result remains evidence about those models.

Slay the Spire's balance postmortem offers a practical discipline: aim for each card to have a place, avoid options that warp the environment, ship frequent playable builds, collect metrics and feedback, and treat data as evidence rather than a conclusion. RimWorld and the My Life as a King postmortem add a related warning: more autonomous behavior or simulation detail can reduce player control and enjoyment when it displaces the player's decisions.

For a tuning decision, combine:

  • authored formulas and sensitivity analysis;
  • telemetry conditioned on context and player population;
  • strategy-search or simulation under explicit objectives;
  • observed play and player explanation;
  • designer judgment about identity and intended experience; and
  • a replay of the affected decision in the production game.

Accessibility Is Part Of The System

Accessibility changes who can perceive, decide, act, and learn. Review it while designing the system, not after the loop is considered finished.

At minimum, examine:

  • redundant visual, audio, text, and haptic cues for critical information;
  • remapping, hold/toggle alternatives, sensitivity, dead zones, and input complexity;
  • focus order, focus visibility, navigation consistency, and modal recovery;
  • color contrast and information that does not depend on color alone;
  • text size, reading time, language, symbol explanation, and cognitive load;
  • motion, flashes, camera behavior, vibration, and feedback-intensity controls;
  • time limits, pause, slowdown, turn-based alternatives, and confirmation;
  • difficulty dimensions that can be adjusted independently when practical; and
  • whether procedural content can create inaccessible routes, signals, or timing demands.

Automated checks can find missing settings, unsupported input paths, contrast violations, focus traps, and timing assumptions. People with relevant access needs are required to establish lived accessibility.

Engagement Without Coercion

Engagement, retention, session length, and return rate are observations. They are not interchangeable with enjoyment, well-being, or product success.

Do not optimize systems toward:

  • fear of missing out or punitive absence;
  • obfuscated odds, costs, or consequences;
  • friction that prevents stopping, refunding, declining, or recovering;
  • excessive interruption or variable rewards whose purpose is to inflate a developer metric;
  • artificial scarcity unrelated to the accepted fantasy; or
  • dark patterns that benefit the developer by reducing informed player choice.

Ask who benefits from the behavior, whether the player understands it, whether they can decline it without disproportionate loss, and whether the design respects the intended player's time and volition.

Claim-Matched Evidence

QuestionAgent or automation can establishHuman evidence required
Is the state internally valid?Invariants, transitions, reachability, softlocks, reset, deterministic tracesNone for the narrow mechanical claim
Are choices mechanically distinct?Options, expected values, sensitivity, candidate dominance, encoded policy outcomesPerceived viability, meaning, satisfaction
Is feedback wired and timely?Event coverage, timing, ordering, latency, channel consistency, lifecycleImpact, clarity in motion, comfort, emotional tone
Is the UI mechanically usable?Overflow, focus, mappings, contrast thresholds, input parity, disabled reasonsDiscoverability, information sufficiency, cognitive burden
Is onboarding complete?Trigger coverage, route completion, prompt persistence, state checkpointsFirst-time comprehension without coaching
Is pacing within authored targets?Interval distributions, dead time, decision density, reward and power curvesTension, fatigue, boredom, anticipation, payoff
Is the system balanced under a model?Simulated outcomes, strategy distributions, exploit candidatesPerceived fairness, strategy identity, actual human adaptation
Is the experience accessible?Configuration and rule checks, automated presentation testsLived access for the relevant population
Is the game engaging or fun?No direct automated proof; telemetry can describe behaviorDirect play, observation, self-report, and representative research for wider claims

Development And Playtest Procedure

For a materially new or changed system:

  1. Approve the promise. The human designer states the role, target player, intended experience, repeated decision, and unacceptable outcome.
  2. Trace the player-system loop. Name what the player perceives, decides, does, sees, learns, and changes.
  3. Model mechanical uncertainty. Write rules, curves, simulations, and invariants only for questions they can answer.
  4. Build the smallest complete decision. Include production input, UI, feedback, consequence, explanation, recovery, and continuation.
  5. Preflight mechanically. Remove crashes, deadlocks, missing feedback, focus defects, impossible states, and known accessibility blockers that would invalidate a human test.
  6. Run owner play without coaching. Observe first, then ask neutral retrospective questions. Record comprehension, usability, engagement, desire to continue, and the intended experience dimensions.
  7. Separate records. Keep observation, player report, interpretation, proposed change, and retest result distinct.
  8. Tune or reject within the slice. Use RITE-style rapid iteration for clear usability defects, but recheck the result empirically.
  9. Widen evidence only with the claim. Recruit representative target players for novice, audience, accessibility, or general-usability claims.
  10. Expand only after acceptance. Reopen the gate when a material change alters the promise, decision, information, feedback, pacing, or player population.

Useful neutral prompts after play include:

  • What were you trying to do?
  • What information did you use?
  • What alternatives did you consider?
  • What did you expect to happen?
  • What surprised you?
  • What caused the result?
  • When did you feel in control or out of control?
  • What would you try next, and why?

Do not teach during a first-session comprehension test, ask whether a feature was "fun" as the only question, or treat a player's proposed implementation as the diagnosis. Players are authoritative about their experience and behavior; the team remains responsible for interpreting the system cause.

Research Basis And Limits

The guide draws on several evidence families:

These sources differ in method, population, genre, and strength. Some are validated instruments or controlled studies; some are theory, surveys, postmortems, or practitioner reports. They improve hypotheses and study design. They do not guarantee that a mechanic will be enjoyable, establish one ideal UI, or make an agent representative of a player.