CodexGameDev/research/hypotheses.md
ookami125 089d28869b 009
2026-08-18 02:25:03 -04:00

20 KiB
Raw Permalink Blame History

Hypotheses

Confidence is deliberately qualitative until there is playtest evidence.

H01 — Curiosity chains are a primary source of fun

  • Confidence: low as a sufficient cause; remains plausible when answers increase valued agency
  • Evidence for: stated interest in Antichamber, Portal, Minecraft modpacks, engineering, systemic interactions, and becoming powerful through understanding. In 002 the player tested ramming and revised heat-extraction tactics. In 004 they voluntarily tested whether the midpoint marker could be moved against a wall after all required rooms.
  • Evidence against: Experiments 000 and 001 produced no voluntary experimentation. The unintended workaround in 003 was boring. The clean self-chosen marker experiment in 004 was explicitly not satisfying.
  • Experiments: 000, 001, 002, 003, 004
  • Unresolved: Does a question become enjoyable when its answer creates useful power, rescue, or improvisational options in an ongoing activity?

H02 — Consistent laws plus feedback are intrinsically interesting

  • Confidence: low for the tested signal language; unresolved for richer physical systems
  • Evidence for: preference for programming/electronics-like magic and systematic exploitation.
  • Evidence against: Experiment 000 felt limited; Experiment 001 made consequences visible but still felt like brute-force work. Consistency alone has shown no intrinsic pull.
  • Experiments: 000, 001
  • Unresolved: Would the same laws become interesting when their effects are embodied, or is the substrate itself too familiar/limited?

H03 — Legible surprise is more valuable than raw complexity

  • Confidence: high as a design requirement; not sufficient as a source of fun
  • Evidence for: the desired curiosity loop requires an unexpected result that makes sense in hindsight.
  • Evidence against: no positive surprise occurred in Experiment 000. Experiment 002 created a ramming hypothesis but failed to make its nonexistent result readable. In 004, the visible midpoint marker was not noticed as a planning reference until after the difficult room; early rooms did not require the model and the transfer jump became brute force.
  • Experiments: 000, 002, 003, 004
  • Unresolved: Can consequences teach a model at the moment it becomes useful without stopping an ongoing activity for tutorial puzzles?

H04 — A general-purpose solution becomes boring when it covers too much

  • Confidence: high prior, untested here
  • Evidence for: Factorio main-bus example and explicit strategy-crystallization concern.
  • Evidence against: universal solutions can remain satisfying when embodiment, execution, or social context varies.
  • Experiments: not yet tested; later scenario variants should attack dominant networks.
  • Unresolved: What is the strategy half-life of a flux network?

H05 — Reusable abstractions should remove repetition without removing reasoning

  • Confidence: medium-high prior
  • Evidence for: interest in programming, modules, and deep systems with narrow interfaces.
  • Evidence against: saved modules may encourage whole-answer reuse or make recomposition feel like integration chores.
  • Experiments: planned only after the base interaction earns another test.
  • Unresolved: What is the useful granularity of a reusable magical module?

H06 — Changing requirements may preserve reasoning

  • Confidence: low when requirements map directly to counter-parts
  • Evidence for: adaptation can prevent complete solutions from crystallizing.
  • Evidence against: changes may invalidate work and feel arbitrary or annoying.
  • Experiments: 003 changed requirements, but Furnace prescribed Sinks and Gale prescribed restoration of cardinal Drives. The player largely replaced the loadout rather than adapting an architecture.
  • Unresolved: Can requirement changes preserve reasoning when the shared system has real compositional structure rather than one-dimensional counters?

H07 — Engineering may need consequences in a cared-about world

  • Confidence: medium; a visible consequence alone was insufficient
  • Evidence for: many preferred games give construction combat, survival, traversal, or world consequences. Experiment 000's abstract instruments produced no reason to improve once their binary predicates passed.
  • Evidence against: Experiment 001's physical familiar was only an output to steer and did not create a reason to care about refinement. In 005, explicit pod rescue and sanctuary integrity also produced no care. A genuinely developed world or relationship may differ from prototype labels and counters.
  • Experiments: 001, 005
  • Unresolved: Does an observable settlement/ecosystem add purpose or distract from the lab?

H08 — Pressure may make decisions emotionally meaningful

  • Confidence: low as a sufficient cause; may still support an already valued activity
  • Evidence for: enjoyment of recoverable chaos, co-op PvE, action games, and improvisation.
  • Evidence against: pressure can interrupt careful engineering and obscure whether the base system works.
  • Experiments: In revision 2, matched Calm received about 2:55 and Breach about 16:44. In 005, the player actively completed a pressured three-minute shift despite wanting to stop almost immediately.
  • Unresolved: Can pressure heighten an independently enjoyable action system without being mistaken for the source of enjoyment?

H09 — Problem randomness with solution choice is preferable to part randomness

  • Confidence: high prior
  • Evidence for: stated preference for Risk of Rain 2 with Artifact of Command.
  • Evidence against: unavailable parts can sometimes provoke creative substitutions.
  • Experiments: 002 offered the same calibration class with exact choice. Revision 1 heavily favored conversion. With easier input in revision 2, Breach choices broadened to capacity 33, conversion 24, and transfer 8, but all choices amplified the same dominant loop.
  • Unresolved: Which opportunity classes create meaningful choices without prescribing builds?

H10 — Some systems require a minimum expressive complexity before they are fun

  • Confidence: medium; adding complexity to the rejected automation activity is now specifically disfavored
  • Evidence for: feedback, timing, storage, and conversion need interactions before emergence appears.
  • Evidence against: adding primitives can disguise a weak core and increase learning cost. The effective limitation may have been the observation space rather than the operation set.
  • Experiments: 000 starts with six transformations and should be judged for missing expressiveness separately from mere lack of content.
  • Unresolved: Is a richer directly manipulated physical system interesting, or does complexity still become work?

H11 — A quality gradient is needed to motivate refinement

  • Confidence: low as a sufficient explanation; quality gradients may be irrelevant without a valued activity
  • Evidence for: in both prototypes, a merely adequate solution gave no reason to improve; the player explicitly cited absent incentive.
  • Evidence against: some systems provoke play without scores or optimization gradients; curiosity can motivate variation on its own.
  • Experiments: 001 provided continuously visible routes and still did not prompt refinement. 002's waves and power growth did prompt repeated refinement, but their causal contribution is confounded.
  • Unresolved: Does improvement become desirable only when it increases power/survival/options inside an engaging activity?

H12 — Direct agency is preferable to programming an autonomous solution

  • Confidence: medium
  • Evidence for: indirect control was described as awkward steering; the player expects adding AI inputs to make the task more tedious, not more fun. Many preferred games emphasize immediate character control.
  • Evidence against: Factorio, Minecraft building, Robocraft, and engineering preferences show that indirect construction can be enjoyed in other motivational contexts. The sensor/AI variant was predicted, not played.
  • Experiments: 002 produced sustained engagement and some tactical adaptation, but its direct control was bundled with pressure, progression, and a richer tradeoff. Revision 2 showed that direct control alone did not prevent strategic repetition.
  • Unresolved: Is building enjoyable when it changes the capabilities of an artifact the player directly pilots, rather than specifying an autonomous controller?

H13 — Optimization needs instrumental stakes

  • Confidence: high after 002, though the kind of stake remains unresolved
  • Evidence for: the player twice cited no incentive to make an adequate solution better. In 002, survival, wave clearing, and escalating power coincided with ten waves of continued tactical adjustment and strongly instrumental upgrade choices.
  • Evidence against: voluntary sandbox creativity can occur without external rewards when the expressive medium itself is compelling.
  • Experiments: 002 calm versus pressure
  • Unresolved: Which stakes create desire rather than obligation: survival, power, protecting others, scarcity, or social consequence?

H14 — The failures came from low knowledge leverage, not automation itself

  • Confidence: medium
  • Evidence for: both prototypes admitted obvious solutions whose results were largely insensitive to deeper laws. The player enjoys other games where constructed systems amplify capability at scale.
  • Evidence against: even imagining richer field-aware control sounded tedious to the player, although that prediction is untested.
  • Experiments: not directly tested; preserve as a competing explanation during and after 002.
  • Unresolved: What minimal system would let one discovery qualitatively expand capability rather than merely improve a score or satisfy a predicate?

H15 — Coupled state changes create meaningful tradeoffs

  • Confidence: medium; useful early, vulnerable to power collapse
  • Evidence for: extracting heat simultaneously created ammunition, slowed a threat, prepared brittleness, and depleted future environmental energy. The player stopped maximizing extraction blindly and made it conditional on enemy count and wall position.
  • Evidence against: wall impact remained dominant in revision 2, producing 176/207 defeats. Extreme conversion made a small heat supply sufficient and erased the scarcity/state coupling later in the run.
  • Experiments: 002; corrective 002 revision should preserve this coupling.
  • Unresolved: How many domain consequences are needed before a choice becomes fertile rather than opaque? Does the tradeoff remain interesting without player-damage pressure?

H16 — Rapid, chosen power growth can sustain engagement

  • Confidence: high as a continuation mechanism, low as evidence of durable fun
  • Evidence for: the player continued through wave 29 in revision 2 while capacity reached 1,255, transfer 220, and conversion 374. Later waves became very fast.
  • Evidence against: the player reported little more to say and wanted to move on. Multiplicative power sustained activity while reducing decisions to one dominant answer.
  • Experiments: 002
  • Unresolved: Can rapid power be a satisfying payoff after qualitative reasoning without replacing that reasoning?

H17 — Physical construction may be enjoyable when it changes a directly piloted capability

  • Confidence: low for weakly spatial modular loadouts; broader hypothesis unresolved
  • Evidence for: the player enjoys construction games where the built object becomes their practical capability. Experiment 001's construction instead specified autonomous movement and felt like manual AI programming.
  • Evidence against: 003 placement barely mattered, trials reduced to obvious compositions, and reorienting Drives added mild friction. Direct piloting did not rescue shallow construction.
  • Experiments: 003
  • Unresolved: Would actual load paths, torque, damage topology, or geometry make construction meaningful, or would it remain implementation work?

H18 — Inferring and transferring an unfamiliar stable law can be intrinsically rewarding

  • Confidence: low for an isolated abstract law; unresolved when knowledge grants consequential agency
  • Evidence for: several enjoyed games turn a newly understood rule into immediate capability. Experiment 004 did provoke one self-chosen question about moving the midpoint marker.
  • Evidence against: the predictive model did not form during 004's required rooms, Transfer became guessing, and answering the self-chosen marker question was not satisfying. An unintended workaround in 003 was also not fun.
  • Experiments: 004
  • Unresolved: Is the missing ingredient better scaffolding, or more importantly a consequence that makes acquired knowledge useful inside a valued activity?

H19 — Recoverable chaos can make systemic agency enjoyable

  • Confidence: low for indiscriminate accumulating chaos; refined version remains untested
  • Evidence for: preference history includes Helldivers 2, Lethal Company, Risk of Rain 2, Crab Champions, and co-op PvE; Experiment 002's pressure variant sustained far more activity than matched Calm and produced early tactical adaptation.
  • Evidence against: pressure in 002 sustained a shallow universal wall-blast loop. In 005 the chaos was readable but became unmanageable: accumulated wreckage made nearly every intervention cause collateral damage, charge reduced options further, and nominal rescue stakes created no care.
  • Experiments: 005
  • Unresolved: Can local mistakes remain recoverable when failure changes the immediate problem without permanently shrinking the action space?

H20 — Assertive, selective agency is more intrinsically valuable than custodial control

  • Confidence: low as a sufficient cause
  • Evidence for: the player's preference history includes combat and movement games with targeted, immediate verbs. Experiment 002's direct attacks sustained more interest than controller construction, abstract puzzles, rig management, or rescue shepherding.
  • Evidence against: In 006 the player was ready to stop after one optional replay. Strike was sufficient rather than intrinsically satisfying, tether worsened position, dash stayed outside attention, and reflection was interesting conceptually but required unpleasantly divided aim/timing/threat attention.
  • Experiments: 006
  • Unresolved: Does assertive action become valuable as the expression and payoff of chosen growth, team coordination, or broader consequence rather than as an isolated activity?

H21 — Qualitative chosen growth creates value through authorship and payoff

  • Confidence: high for short-run interest and continued build testing; medium for sustained fun
  • Evidence for: the player prefers Risk of Rain 2 with Artifact of Command and enjoys Warframe, Crab Champions, Minecraft modpacks, and becoming extremely powerful through understanding. In 002, rapid chosen power growth sustained the longest sessions. In 007, the mutation condition prompted a logged Arc-only replay plus at least three additional reported runs involving Bloom, the combined system, and all-Bloom. The player formed an inaccurate ranking from descriptions, observed unexpectedly strong cross-effect behavior, revised Bloom upward, explicitly contrasted this with boring independent numerical upgrades, and kept generating specific comparative builds. Watching an all-Bloom cascade erase the field was reported as “kind of fun.”
  • Evidence against: 002's upgrades all amplified one wall-blast answer and the player reported little continuing interest. In 003, changing parts was obvious loadout work rather than expressive authorship. Bloom in 007 appeared overtuned, so raw exponential power and spectacle remain confounded with composition. Long-term strategy half-life is untested.
  • Experiments: 002, 003, 007
  • Unresolved: Does a larger but still legible network of hit/kill/status interfaces sustain chained build questions under changing enemy ecologies, or crystallize into a full-synergy package or obvious counters?

H22 — Cross-effect interfaces matter more than independently useful upgrades

  • Confidence: high for the tested short comparison
  • Evidence for: In 007, both conditions offered exact chosen growth on the same fields. The player predicted boredom from Trial II's independently useful damage/rate/width options, but found Trial I interesting enough for multiple replays after Fork, Arc, and Bloom interacted through shared hit/kill events. They specifically said Trial II's upgrades would have made Trial I more fun but did not create interesting differences among themselves. Later runs mapped Bloom to small enemies and Arc to large enemies, showing conditional knowledge rather than only higher throughput.
  • Evidence against: Mutation also had more spectacle, automatic targeting through Bloom, and possibly greater effective power. Those factors were not independently controlled. Only three interacting families were tested.
  • Experiments: 007
  • Unresolved: Is the payoff driven by planning causal composition, by surprising emergent cascades, by reduced execution burden, or by audiovisual/raw-power expression?

H23 — Changing enemy ecology can preserve compositional reasoning

  • Confidence: low as implemented; ecology was legible but not consequential
  • Evidence for: Without prompting, the player concluded that Bloom excelled against small enemies while Arc performed better against large enemies. Conditional target profiles can make component knowledge transfer while preventing one exact build from answering every field.
  • Evidence against: Experiment 003's changing requirements collapsed into obvious prescribed counter-parts. In 007, taking all mutation families together may dominate, and all-Bloom's aggressive scaling may erase ecology distinctions through raw power.
  • Experiments: 003 is negative adjacent evidence; 007 generated the hypothesis; 008 first run supplied only one short field after the fourth choice and could not isolate it.
  • Evidence update: In 008 the player used three different builds and believed multiple approaches may exist, but reported that each expedition ended just as something cool began emerging. The fourth module was active for only 722 seconds. Conduit-first was a misunderstanding, not planned interface composition.
  • Revision 2 evidence: Added mature fields made choice effects easier to see and supported small movement-strategy refinements. Builds differentiated and Bloom+Arc stopped being universal. However, the player believed almost any build or even baseline fire could complete every expedition; no exclusion created a relevant limitation, so the cap generated little meaningful tradeoff.
  • Interpretation: Population variety cannot preserve reasoning when success is insensitive to the composition. Simply raising difficulty remains an untested and potentially confounded correction.

H24 — Enjoyable building needs capability-sensitive consequence

  • Confidence: high as a requirement; the suitable consequence remains unresolved
  • Evidence for: In 008 revision 2, longer exposure made three causal builds legible and building was explicitly reported enjoyable. It simultaneously felt meaningless because nearly any build or baseline fire appeared sufficient. Earlier 000/001/003 failures also lacked consequences sensitive to deeper construction, while 002's survival/wave stakes gave upgrades instrumental value before one strategy dominated.
  • Evidence against: Pure sandbox building can be enjoyable without external failure when the output is expressive enough. Experiment 007 generated repeated build tests from surprise and spectacle alone, so hard challenge is not always required.
  • Experiments: 000, 001, 002, 003, 007, 008
  • Unresolved: Can capability-sensitive enemies make exact qualitative building matter through multiple causal routes without collapsing into obvious counters or raw throughput pressure?
  • Unresolved: Can mixed/behavioral ecologies support several viable causal compositions, or does adaptation become “equip the labeled counter”?