Initial Commit

This commit is contained in:
ookami125 2026-08-18 00:52:29 -04:00
commit 35f3810632
90 changed files with 29267 additions and 0 deletions

187
research/current_model.md Normal file
View file

@ -0,0 +1,187 @@
# Current Preference Model
Status: **updated after Experiment 008's under-length first run; corrective revision 2 is validated and ready**
The leading long-term theory remains that fun may come from learning a compact set of consistent laws, constructing a system from them, and discovering consequences that create further self-directed questions. Experiment 000 did not provide positive evidence: its permissive abstract observations were solved in about three minutes, produced no voluntary experimentation, and did not make phase, timing, or cyclic behavior perceptible or necessary.
After Experiment 000, the working theory was:
> A law is not meaningfully legible because it is documented or rendered. It becomes legible when predicting it is necessary to explain and control an outcome the player can perceive and care about.
Binary completion without a consequential quality gradient gives little reason to improve a merely adequate system. A technical system also needs an open consequence space: after solving the stated problem, changing the machine must plausibly create another behavior worth seeing.
Experiment 001 made output spatial and continuous, but the player wanted to stop almost immediately. Moving a familiar through the field felt like brute-force steering through an indirect interface. The player predicts that field sensors or a more capable controller would still feel like tedious manual AI programming. That prediction is evidence about their current expectation, not a tested result; it should lower the priority of that extension without being treated as proof against automation generally.
The newest model is:
> In both tested prototypes, engineering had little value beyond reaching an arbitrary terminal condition, so construction and refinement felt like work without leverage.
The player may need to participate directly in an intrinsically engaging ongoing activity—combat, movement, cooperative crisis, or recoverable chaos—where systems knowledge is a means to greater agency, power, and improvisation. In that framing, optimization has an external purpose rather than asking the player to care about elegance for its own sake.
This is not permission to hide a weak system beneath action spectacle. A controlled calm/pressure comparison can test whether direct manipulation has its own interesting interactions and whether pressure helps or harms them. It is an exploration probe, not the newly assumed final direction.
Experiment 002 produced the first sustained engagement and voluntary tactical revision. The player stayed for about 16 logged minutes, cleared ten Breach waves, adapted extraction based on position and remaining fuel, concentrated upgrades into conversion efficiency, and tested an unimplemented ramming hypothesis. This supports a more specific candidate:
> Interest may arise when one operation changes several strategically relevant domains at once, creating a stateful tradeoff whose value depends on context.
Extracting heat was simultaneously resource acquisition, enemy slowing, brittle preparation, and depletion of future available energy. The player learned not to maximize it blindly. This is more informative than the action theme alone.
The first result was heavily confounded. Corrective revision 2 made Calm and Breach matched complete runs and fixed the major interface/physics defects. The second session lasted about 19:39, including 2:55 in Calm and 16:44 in Breach, but did not produce a new reported layer of interest. The player is ready to move on.
Revision 2 clarifies the failure shape. Easier mouse input broadened upgrade choices and injection use, but it did not broaden the strategy. Wall-directed conversion caused 176 of 207 defeats. All three upgrade dimensions eventually amplified the same throughput loop, conversion reached 374×, and the starting scale of heat became sufficient to erase later waves.
The updated model is:
> Instrumental stakes and rapid power can sustain activity, but continued input is not the same as continued reasoning. A system loses value when every improvement amplifies one complete answer rather than changing which answer fits the current problem.
Pressure probably increases persistence, but it has not yet been shown to create curiosity. The next experiment should test changing qualitative requirements and physical construction without autonomous programming or escalating numerical power.
Experiment 003 tested that proposal and was boring despite completion of all three trials in about 6:47. The player explicitly ruled out awkward controls as the main cause. Haul, Furnace, and Gale each collapsed into an obvious part response, and physical placement barely mattered. Requirement changes produced loadout replacement, not adaptation of a meaningful architecture.
The player nevertheless invented a two-tractor cargo-juggling workaround after overlooking Sinks and misunderstanding the heat bar. This yields an important methodological correction:
> Voluntary experimentation shows agency and problem solving, but it is not sufficient evidence of fun. The experiment itself must generate a question or consequence the player values.
The next high-value uncertainty is legible surprise. No experiment so far has cleanly required the player to infer an unfamiliar stable law and transfer that model. All have either disclosed the relevant laws, admitted obvious brute force, or offered direct counter-parts.
Experiment 004 tested that uncertainty and split it into two results. Rooms 1 and 2 were obvious without requiring a predictive model. Transfer jumped directly to a state-planning problem, so the player guessed through 94 commands and never felt able to predict the system. The yellow midpoint cross was visible but did not become meaningful until the open chamber. Rendering state is not the same as making it strategically legible.
The player then formed a genuine unrequired question: could the yellow cross be moved against a wall? They spent about a minute testing it. This is the cleanest self-directed experiment yet, but answering it was explicitly not satisfying. Therefore curiosity-like behavior is not sufficient, even outside an authored objective.
The updated leading model is:
> A systemic question is more likely to matter when its answer changes the player's agency inside an already valued, ongoing activity. An abstract answer, optimized marker, or completed predicate is not enough by itself.
This does not prove that pure discovery or spatial puzzles cannot be fun for the player; Experiment 004's transfer scaffold was flawed. It does lower the priority of polishing isolated rule-discovery puzzles. The higher-information next probe is direct improvisation under recoverable, changing circumstances, with systemic verbs that affect useful and dangerous objects together and no runaway stat progression.
Experiment 005 tested that probe and failed almost immediately at the motivational level. The player understood the instructions but did not find the rescue activity interesting and did not care what happened. They nevertheless actively completed the three-minute shift, reinforcing that duration and input volume cannot stand in for enjoyment.
Later failure had a distinct cause: the chaos was readable but unmanageable. Accumulating wreckage and indiscriminate radial force meant that attacking raiders also threatened pods and the sanctuary. The field's shared consequences did not create fertile tradeoffs; they erased credible good actions. Charge depletion further reduced agency without adding a valued decision.
The updated model is:
> Coupling is valuable only when it creates leverage the player can selectively exploit. If every intervention propagates unavoidable collateral damage, readable emergence becomes paralysis.
Experiment 005 also did not implement “recoverable chaos” successfully. Captured pods disappeared, debris accumulated, and failures shrank the future action space. A truly recoverable mistake should transform into a different manageable problem, not permanent clutter or irreversible loss.
The next high-information contrast is assertive rather than custodial agency: directly create outcomes with targeted movement/combat verbs in a compact field, with local resettable failures and no upgrades. This tests whether the player values expressive action itself before construction, progression, or world-care layers are added.
Experiment 006 tested that contrast and did not find an intrinsically enjoyable base verb. The player cleared all four arrangements and voluntarily replayed two, but the replays were investigations rather than a desire for mastery. Grouping enemies continued the prior “can these overlap?” edge-case test and ended in an accidental six-kill strike. Reflection was noticed late and deliberately revisited because the concept resembled returning a Minecraft ghast fireball, but the player was ready to stop afterward.
Reflection is an especially useful correction to the method. Its concept attracted attention, while its execution did not: aiming at a moving enemy, judging projectile distance, and monitoring other threats competed for focus. Thus discovery, voluntary replay, a familiar analogy, and successful use can all occur without the interaction being fun.
Strike was merely sufficient, tether usually worsened position by pulling danger closer, and dash never became necessary or cognitively available. This lowers confidence that stripped-down assertive action is an intrinsic foundation. It does not establish that action, counters, or dodging are globally unsuitable; it establishes that isolated verbs without a valued trajectory or context have now failed alongside isolated construction, discovery, and rescue.
The strongest remaining cross-experiment candidate is not a single mechanic but a **trajectory of authored capability**. Experiment 002's rapid upgrades sustained by far the most play, but numerical multiplication collapsed into one answer. The player's prior preferences also repeatedly combine direct participation with deliberately chosen builds and dramatic power expression. The next experiment should therefore contrast numerical improvement with qualitative, player-chosen transformations over the same short action substrate. Its purpose is to test whether composing a build and witnessing its consequences supplies value that the bare action lacks—not to assume progression will rescue it.
Experiment 007 implements that comparison as two randomized-order trials. Both use the same low-demand hold-to-fire combat, five fixed fields, health rules, and four exact choices. One trial offers only damage, cadence, and projectile-width amplification. The other offers stacking Fork, Bloom, and Arc effects whose products can interact. The key outcome is whether a choice creates anticipation, causal understanding, or desire for another consequence—not trial completion, choice count, clear speed, or effect volume.
Experiment 007 produced the first clear positive curiosity chain. The player first chose Fork because it sounded interesting, then Arc for the same reason, without predicting their interaction. They initially underestimated and skipped Bloom. After completing both conditions, they replayed mutation with Arc stacked four times to test its isolated strength, found it weak alone, and performed another unsaved run involving Bloom/all three effects. The combined behavior was much better than expected, caused them to revise Bloom from skipped to the strongest option, and made the mutation trial interesting enough to replay.
The numerical condition was predicted to be boring as soon as its menu was read. Its choices were recognized as useful additions to the mutation system, but not as relationships that created interesting differences by themselves. This makes the leading model more specific:
> Chosen growth becomes interesting when components alter the future value and behavior of other components through legible interfaces. Independent improvements increase output; compositional effects create new things to predict, observe, and retest.
Fork, Arc, and Bloom formed a bounded causal network through hit and kill events. The player could understand each component, guess incorrectly about their combined value, and revise their model through play. This resembles the original “deep systems, narrow interfaces” aspiration more closely than the globally entangled chaos of 005 or the obvious counter-parts of 003.
The combined result was not an endpoint. The player performed two further runs to test more builds, including all-Bloom. They reported that watching an aggressively scaling Bloom build erase the field after a few shots was “kind of fun,” and began mapping Bloom to small enemies and Arc to large enemies. This is the first prototype to produce both an explicit positive enjoyment statement and a sustained chain of self-selected comparative experiments.
The updated leading model is:
> Fun is emerging from authored causal composition: exact component choice, surprising but legible cross-effect behavior, conditional strengths against different problem classes, and a visible power payoff that makes revised understanding worth expressing.
This does not establish that the shooter substrate is independently fun. It may be functioning primarily as a fast, legible oscilloscope for a build. That is acceptable at this research stage: unlike the arbitrary scopes in 000, the output creates differentiated agency and a spectacle the player sometimes enjoys.
The next risk is strategy crystallization. Fork+Arc+Bloom may become a universal package, all-Bloom may simply overpower every ecology, or enemy distinctions may prescribe obvious counters as in 003. The next experiment should retain the same causal vocabulary while limiting simultaneous components and presenting changing mixtures/behaviors. It should test whether the player can transfer knowledge and still face genuine composition choices, not whether more upgrade content is automatically better.
Experiment 008 implements that extension. It preserves Fork, Bloom, and Arc, then adds Focus at the repeated-primary-hit interface, Conduit at the secondary-hit interface, and Resonance as a modifier of all secondary effects. Three immediately available expeditions emphasize small bodies, concentrated durable bodies, and spawning mixtures. Corrective revision 2 uses eight fields: four exact choices occur after fields one through four, then the complete four-choice build persists through fields five through eight.
The intended test is not whether more modules or enemy types are more entertaining. It is whether the player carries causal knowledge into a new population, forms a build prediction, encounters a result that revises that prediction, and generates another composition question. Different builds without that reasoning could be obvious counter-loadout work rather than the desired mechanism.
The first 008 run did not provide enough mature-build exposure to answer that question. The player selected three different builds but reports little planning; Conduit-first was a misunderstanding. More importantly, each fourth selection was followed by only one field. Complete builds existed for about 7 seconds in Shoal, 20 seconds in Bastion, and 22 seconds in Brood. The player repeatedly began to see something cool emerge just as the expedition ended.
This is a structural measurement failure, not evidence against the compositional model. Compared with 007, the module vocabulary and populations grew more complex while mature-build observation time did not. Corrective revision 2 adds three fields to every expedition without changing the four-choice budget or module mechanics. The four selections still occur after fields one through four, and the completed build now persists through fields five through eight. Only after this sustained observation can replay, adaptation, universal cores, or obvious counter-loadouts be interpreted.
## Methodological Guardrail: Reports Are Evidence, Not Ground Truth
The player is an imperfect observer of their own enjoyment, as every playtester is. Treat reports of boredom, friction, desire, and behavior as high-value evidence. Treat causal explanations and predictions about hypothetical versions as hypotheses.
Maintain distinctions between:
- **observed behavior:** when they stopped, what they did, what they voluntarily tried;
- **reported experience:** what felt tedious, obvious, confusing, or pointless;
- **player causal theory:** why they believe it felt that way;
- **designer interpretation:** alternate causes consistent with the same evidence.
Do not reflexively implement the player's proposed fix, and do not reflexively eliminate a whole design family because the player predicts a variant would fail. Seek prototypes that discriminate root causes.
Current alternate explanations for the two failures remain:
- The goals were arbitrary and had no downstream value.
- The systems had too little expressive depth or knowledge leverage.
- Solutions contained no meaningful tradeoffs and converged immediately.
- Feedback exposed state but not interesting consequence.
- Construction friction exceeded the value of its output.
- Static completion removed all incentive for refinement.
- Indirect control itself may be a poor fit, but evidence from enjoyed building games argues this depends on context.
The most promising problem shape currently has:
- stable laws and changing circumstances;
- randomness in the problem or opportunity, not in access to exact solutions;
- reusable components and mental models, without reusable whole answers;
- enough external consequence to make engineering choices matter, but not necessarily time pressure;
- legible feedback that makes surprising behavior understandable after inspection.
The leading threat remains **strategy crystallization**, but Experiment 000 failed earlier than that: arbitrary signal handling was competent enough that its central distinctions never entered the strategy. Before testing long-term crystallization, a prototype must make its laws causally discriminating without collapsing into one authored waveform puzzle.
## Evidence From Experiment 000
- All observations completed in approximately three minutes or less.
- No desire to improve or continue after completion.
- No unrequired experiment.
- Source cycles and vector directions were not perceived.
- Random timing and direction would have been sufficient for the observation predicates.
- The player described the available possibility space as limited.
## Evidence From Experiment 001
- Desire to stop was almost immediate.
- Signal construction felt like awkward steering.
- Route completion was brute-force directional movement, not meaningful system reasoning.
- No desire to refine a completed route or try another physical problem.
- The player predicts that adding world-aware inputs would create tedious manual AI programming; this extension was not actually tested.
- Making consequences visible was insufficient; the consequence itself was not desirable enough to optimize.
## Evidence From Experiment 002
- About 16 logged minutes, roughly 14 in Breach; ten complete waves cleared.
- Tactical revision based on temperature, position, remaining enemies, and reservoir state.
- 74 Breach defeats, 67 non-wasted discharges, 184 tether sessions.
- Upgrade choices were strongly instrumental: 19 conversion, four capacity, one transfer.
- Reservoir was nearly empty in 28% of Breach snapshots, yet capacity was not considered the main solution; output per unit was.
- Voluntary ramming experiment, despite player collision damage not existing.
- Later damage avoidance improved dramatically, but learning, runaway power, and wall exploitation are confounded.
- In the first session, Calm versus pressure was untested because Calm lacked equivalent waves and was perceived as a sandbox.
- The trackpad conflict is historical evidence from the first 002 run. The player now has a mouse, so it should not anchor interpretation of the next run.
- Revision 2 gave Calm a matched structure. Calm received about 2:55 versus about 16:44 in Breach, which weakly supports pressure as a continuation driver but not as a source of deeper reasoning or enjoyment.
- Revision 2 produced 207 defeats and 69 upgrade choices without a qualitative strategy change. Duration, waves, and input count must not be used as stand-alone fun measures.
## Unknowns With High Information Value
- Does the coupled heat/resource tradeoff remain interesting after controls and wall exploits are corrected?
- Does an equivalent goal without pursuit/player damage sustain tactical reasoning?
- Was continued play driven by property reasoning, wave progression, runaway power, or conventional action feedback?
- Does bodily participation in collision create useful agency, or only risk and confusion?
- Are transfer/conversion laws interesting when exploited tactically against varied physical properties?
- Does unrestricted access create agency, or remove a useful sense of progression?
- Does modifying a directly piloted physical construction feel different from programming an autonomous controller?
- Do qualitatively changing requirements cause useful recomposition, or merely repeated loadout work?
- Does discovering an initially unexplained but consistent environmental rule create interest when application, not construction, is the interface?
- Does a learned rule transfer to a new situation in a way that feels empowering rather than like solving a short authored riddle?
- Can a simple coupled physical verb support enjoyable improvisation when the field continually creates recoverable problems rather than terminal puzzles?
- Does a self-generated tactic feel rewarding when it protects, rescues, or weaponizes something in motion rather than only changing an abstract measurement?
- Are targeted, assertive verbs intrinsically more engaging than globally coupled rescue/maintenance verbs?
- Does local execution mastery create a reason to replay when there is no build progression or accumulating world state?
- Does deliberately choosing qualitative capability transformations create authorship and anticipation that numerical growth does not?
- Does a build interaction produce a desire to see the next consequence even when the underlying combat is only serviceable?