This commit is contained in:
ookami125 2026-08-18 02:25:03 -04:00
parent 35f3810632
commit 089d28869b
13 changed files with 11741 additions and 27 deletions

View file

@ -1,22 +1,22 @@
# Agent Handoff — Private Research Notes
Last updated: 2026-08-17, after the successful Experiment 007 result and validated Experiment 008 implementation.
Last updated: 2026-08-18, after the completed Experiment 008 interpretation and validated Experiment 009 implementation.
This file is written so another agent can continue the research program without reconstructing the reasoning. The player intends not to read it before playtesting, to avoid expectation effects. It contains design hypotheses, likely failure interpretations, and things to watch for.
## Current Handoff Snapshot — Read This First
Date: 2026-08-17.
Date: 2026-08-18.
The current playable is **Experiment 008 — Catalyst Ecology revision 2** in `experiments/008_catalyst_ecology/`. Corrective telemetry is inspected and the detailed player report is pending. Experiment 007 is complete and is the first successful probe.
The current playable is **Experiment 009 — Catalyst Ascent revision 1** in `experiments/009_catalyst_ascent/`. Experiment 008 is closed: the player enjoyed building but found it meaningless because nearly any build or baseline fire seemed viable. Experiment 009 tests the designer interpretation that construction needs capability-sensitive consequence; it does not assume that more difficulty is the solution.
Run it with:
```bash
./experiments/008_catalyst_ecology/run.sh
./experiments/009_catalyst_ascent/run.sh
```
Then open `http://127.0.0.1:8000`. The custom local server writes validated logs directly to repository `JSONL/` when the player presses **Save JSONL**. Port 8000 was free and all Codex validation processes were stopped at handoff.
Then open `http://127.0.0.1:8000/experiments/009_catalyst_ascent/prototype/`. The custom local server writes validated logs directly to repository `JSONL/` when the player presses **Save JSONL**. Validation artifacts were moved out of `JSONL/`; only player logs should remain there. The validation server and Chromium process were stopped at handoff.
### The actual research objective
@ -45,9 +45,9 @@ The strongest current inference is not “the player dislikes systems,” automa
The current compact theory is:
> A systemic question is more likely to matter when its answer increases agency inside an activity the player already values. Coupling is useful only while it creates selective leverage; indiscriminate consequences can erase good actions rather than create emergence.
> Fun can arise from building a causal capability, discovering how components amplify one another, and expressing that understanding as visible power. For the activity to remain meaningful, problems must distinguish capabilities through consequences while still admitting multiple causal routes.
This remains a hypothesis. Experiment 006 deliberately tests the lower layer first: whether immediate, assertive, targeted action has any intrinsic value before adding engineering, progression, world attachment, or multiplayer context.
This remains a hypothesis. Experiment 007 supplied the first successful curiosity chain and Experiment 008 isolated enjoyable building from meaningful consequence. Experiment 009 now tests whether capability-sensitive problems join those two qualities, while guarding against the alternate explanation that added resistance merely makes a serviceable shooter compulsory.
### Experiment history in one page
@ -269,7 +269,7 @@ Revision 2 validation completed three full synthetic runs using the first-sessio
After the corrective playtest, compare mature fields five through eight within each expedition. Ask whether interactions became understandable, whether any full build developed or flattened over those fields, whether a specific alternative build arose, and when the extra exposure shifted from useful observation to repetition. Do not compare total duration directly to revision 1 as enjoyment evidence; revision 2 deliberately contains more fields.
### Experiment 008 revision 2 telemetry; follow-up pending
### Experiment 008 revision 2 complete report
The player saved `JSONL/catalyst-ecology-7b7c7a92-42cb-4ab2-8a81-d1316ea972c5.jsonl`; analysis is in `experiments/008_catalyst_ecology/results/7b7c7a92-revision-2-preliminary-analysis.md`.
@ -281,7 +281,48 @@ All three eight-field expeditions completed first attempt without replay. Builds
Bloom+Arc was no longer universal. Fork appeared in every build but fed different consumers. Mature-build exposure increased to approximately 41 seconds in Shoal, 76 in Bastion, and 56 in Brood. The player says it was “a bit easier to see how my choices impacted my play,” confirming the length correction improved visibility.
Telemetry alone cannot distinguish population-aware composition from deliberate experiment coverage. Ask why each build differed, what role Fork played, whether any mature interaction was satisfying/surprising rather than only readable, when extra fields became repetitive, whether a specific alternative build arose, and how the four-slot limit felt after sustained exposure.
The player combined memory from revision 1 with intuition rather than following fully planned ecology builds. Fork appeared everywhere because it was a decent general projectile generator and seemed to double/triple-hit large bodies. Mature fields allowed slight refinement into consistent movement strategies.
Power-up choice was only a little more meaningful, not significantly so. The player believed essentially any combination or even no power-ups would remain viable. Because every module was a free benefit and baseline combat was permissive, the four-slot cap created no felt tradeoff. Thus different builds do not strongly validate H23.
The final distinction is: **building was enjoyable but meaningless**. This confirms construction/composition itself had value, while the permissive environment made architecture irrelevant to success. Close 008. Do not add more fields or tune the same expeditions again.
The next higher-information probe should use a single escalating qualitative-build run where baseline output eventually becomes insufficient through capability-sensitive enemies, exact choices continue, and causal combinations can express dramatic power. Prefer several causal routes—hit generation, burst, kill chains, secondary conversion—over labeled one-module counters. Use local wave restart. Guard against mistaking survival-driven continuation for enjoyment, as in 002.
### Experiment 009 implementation and validation
Experiment 009 is implemented as **Catalyst Ascent**. It deliberately reuses the 008 engine and the same six catalysts so the independent change is closer to consequence sensitivity than content novelty. The new page lives in `experiments/009_catalyst_ascent/prototype/`, sets `window.CATALYST_MODE = "ascent"`, and loads the shared 008 stylesheet and application. The shared application defaults to unmodified 008 behavior when that flag is absent.
The run has ten fields. Exact choices occur after fields one through seven, and the completed seven-choice build persists through fields eight through ten. The first three fields establish Motes, Husks, and Broods. Later mixtures introduce:
- **Wards:** eight shield segments must be stripped within a 1.55-second window or partial progress resets. A broken shield reforms after 2.8 seconds if the body remains alive. Focus rupture and Conduit lance bypass the shield and damage the body on the same event. Baseline fire can just barely strip a shield with uninterrupted accurate fire; Fork fragments, Bloom sparks, Arc damage, Focus, and Conduit provide different routes.
- **Renewals:** 30 health and continuous 3.6 health/second regeneration. This favors concentrated output without naming one required component.
- **Broods:** retain the existing bounded Mote spawning, providing both accumulating pressure and possible fuel for kill-triggered Bloom.
Later fields mix those capabilities with Motes, Titans, and one another. This is intended to distinguish hit density, single-target rupture, secondary-hit conversion, kill chains, and secondary amplification. It may instead produce one universal dense-effect network or obvious textual counters; preserve those as live failure interpretations.
Failure restores only the current field with the current build. There is still a possible build-quality recovery limitation: a player who chooses seven levels of a non-producing modifier could make progress extremely difficult and would need **Restart run**. The design mitigates ordinary cases by making baseline shield stripping technically possible and offering all exact options every time, but this has not been playtested for feel. Do not silently reinterpret a hard or tedious field as meaningfulness.
Validation completed on 2026-08-18:
- JavaScript and shell syntax pass.
- The 1672×976 layout was visually inspected. The full field, instructions, capability descriptions, build, and start overlay fit without page scrolling.
- Deterministic Ward validation reduced a shield from 8 to 5, observed it reset to 8 after the opening window, then used Focus to breach the shield and reduce body health from 14 to 9.
- Deterministic Renewal validation damaged one from 30 to 20 and observed it regenerate to approximately 24.32 over 1.2 seconds; `renewal_regenerated` telemetry fired.
- A full synthetic Ascent recorded exactly ten `wave_started`, ten `wave_completed`, seven `choices_shown`, seven `module_chosen`, two `post_build_field_advanced`, and one `run_completed` event. The final test build was Fork 2 / Bloom 1 / Arc 1 / Focus 1 / Conduit 1 / Resonance 1.
- Direct server upload wrote 738 valid revision-1 events to `JSONL/catalyst-ascent-<session>.jsonl`; the validation file was then moved to `/tmp`.
- Regression validation reloaded Experiment 008 without the mode flag and confirmed Shoal, three expedition buttons, eight fields, four choices, three post-build transitions, completion, and `experiment: 008_catalyst_ecology` telemetry.
Expected player logs are `JSONL/catalyst-ascent-<session>.jsonl`.
After the player saves:
1. Inspect run attempts, wave attempts, choice order and deliberation, build at every field, Ward hit/reset/break/reform causes, Renewal regeneration, Brood spawning, module triggers, target kill causes, damage/defeats, restarts, completion, replay, and save.
2. Write a preliminary result under `experiments/009_catalyst_ascent/results/` before asking follow-ups.
3. Separate what the player anticipated when choosing from what they noticed during play and what they inferred only afterward.
4. Ask when the build first enabled something baseline fire did not, whether a capability felt multiply solvable or prescribed, what alternate build—if any—they wanted to try, and when they were ready to stop.
5. Do not infer meaning from necessity alone. A forced counter can be empty; long play can be attrition; completion can be compliance.
6. Preserve the possibility that building remains enjoyable while the combat objective itself remains something the player does not care about.
### Historical Experiment 006 interpretation branches
@ -315,7 +356,7 @@ If port 8000 is occupied after agent validation, inspect with `ss -ltnp 'sport =
- `codex_game_design_research_plan.md` — original research mandate and preference priors.
- `research/current_model.md` — synthesized current preference model.
- `research/hypotheses.md` — H01H23 with evidence and confidence.
- `research/hypotheses.md` — H01H24 with evidence and confidence.
- `research/experiment_index.md` — compact experiment history.
- `research/agent_handoff.md` — this private operational record.
- `experiments/*/hypothesis.md` — per-experiment private intent.

View file

@ -1,6 +1,6 @@
# Current Preference Model
Status: **updated after Experiment 008's under-length first run; corrective revision 2 is validated and ready**
Status: **updated after completed Experiment 008; Experiment 009 is ready to test capability-sensitive consequence**
The leading long-term theory remains that fun may come from learning a compact set of consistent laws, constructing a system from them, and discovering consequences that create further self-directed questions. Experiment 000 did not provide positive evidence: its permissive abstract observations were solved in about three minutes, produced no voluntary experimentation, and did not make phase, timing, or cyclic behavior perceptible or necessary.
@ -100,7 +100,25 @@ The intended test is not whether more modules or enemy types are more entertaini
The first 008 run did not provide enough mature-build exposure to answer that question. The player selected three different builds but reports little planning; Conduit-first was a misunderstanding. More importantly, each fourth selection was followed by only one field. Complete builds existed for about 7 seconds in Shoal, 20 seconds in Bastion, and 22 seconds in Brood. The player repeatedly began to see something cool emerge just as the expedition ended.
This is a structural measurement failure, not evidence against the compositional model. Compared with 007, the module vocabulary and populations grew more complex while mature-build observation time did not. Corrective revision 2 adds three fields to every expedition without changing the four-choice budget or module mechanics. The four selections still occur after fields one through four, and the completed build now persists through fields five through eight. Only after this sustained observation can replay, adaptation, universal cores, or obvious counter-loadouts be interpreted.
The first run's structural measurement failure was corrected in revision 2. Mature fields made choices easier to evaluate, produced distinct full builds, and allowed slight movement-strategy refinement. However, the player believed nearly any combination—or even no catalysts—could complete the expeditions. Baseline sufficiency meant every choice was a benefit but no exclusion mattered, so the four-slot limit created no felt tradeoff.
The updated model is:
> Causal composition creates curiosity and power expression, but strategic choice requires the environment to distinguish capabilities. A slot limit alone is meaningless when every included effect helps and every omitted effect is unnecessary.
Do not respond by merely inflating enemy health in the same three expeditions. That could make throughput compulsory without creating new reasoning. The next high-information test should combine 007's exact qualitative growth with a single escalating run whose changing capabilities eventually exceed baseline fire. Continued choices should let the player author a response and reach spectacular power, while telemetry/report distinguish valued build anticipation from survival-driven persistence.
The final 008 report sharpens this further: the player explicitly called the building enjoyable but meaningless. This establishes the first positive statement about construction itself in the project. It also shows why prior quality gradients and slot limits failed: a choice has no weight if the environment is insensitive to its omission.
The leading model is now:
> Fun can arise from building a causal capability, discovering how components amplify one another, and expressing that understanding as visible power. For the activity to remain meaningful, problems must distinguish capabilities through consequences, while still admitting multiple causal routes.
The next probe should be one escalating run rather than three labeled ecology tests. Exact choices should continue so a build can mature and specialize. Later enemies should introduce capabilities—such as re-forming wards, regeneration, and spawning—that baseline fire cannot comfortably answer, but which several combinations can address through hit generation, burst, kill chains, or secondary-hit conversion. Local wave restart should make failure informative rather than erase the run.
Experiment 009 implements this as Catalyst Ascent: one ten-field run, with exact choices after the first seven fields and the mature build retained through the final three. It reuses the six known catalysts so novelty comes from consequence rather than a larger menu. Wards reset partial shield damage unless eight hits land within a short window, while Focus ruptures and Conduit lances breach them; Renewal bodies continuously regenerate; Broods continue to produce Motes. Mixed later fields allow hit density, focused rupture, secondary conversion, kill chains, and amplification to overlap rather than assigning one named counter.
This is deliberately not a generic difficulty test. The key result is whether the player anticipates, notices, and revises a capability relationship—and whether doing so makes the already-enjoyable building feel consequential. Longer play, deaths, clearing all fields, or selecting the textual counter do not answer that question. A possible failure is that pressure merely compels throughput as in 002; another is that the shooter objective remains emotionally meaningless even when construction affects success.
## Methodological Guardrail: Reports Are Evidence, Not Ground Truth

View file

@ -10,4 +10,5 @@
| [005 — Sanctuary Wake](../experiments/005_sanctuary_wake/README.md) | Does a coupled physical tool become enjoyable in a changing, recoverable rescue crisis without stat growth? | Wanted to stop immediately; readable chaos became unmanageable, charge reduced agency, and rescue outcomes inspired no care | H08/H19 down as sufficient; indiscriminate coupling identified as harmful | Is assertive, selective action intrinsically more valuable than custodial management? |
| [006 — Breakline](../experiments/006_breakline/README.md) | Are targeted combat/movement verbs enjoyable without progression, timers, or accumulating failure? | No; all sets were cleared, but strike dominated, tether hurt positioning, dash was forgotten, and reflection was interesting to inspect but attention-splitting rather than fun | H20 down as a sufficient cause; voluntary replay again separated from enjoyment | Does rapid, deliberately chosen qualitative transformation make otherwise serviceable action worth continuing? |
| [007 — Catalyst Trials](../experiments/007_catalyst_trials/README.md) | Does rapid qualitative chosen growth create value beyond matched numerical amplification? | Successful: mutation prompted at least four repeat/comparison runs, wrong predictions, revised build knowledge, conditional small/big-enemy mapping, and a qualified fun report; numerical choices were boring | H21/H22 strongly up; H23 added; first clear curiosity chain | Can changing enemy ecologies and limited composition preserve reasoning without prescribing counter-loadouts? |
| [008 — Catalyst Ecology](../experiments/008_catalyst_ecology/README.md) | Can a four-choice compositional system remain interesting across small, durable, and spawning populations? | Revision 2 completed; three distinct mature builds were clearer in use, Bloom+Arc ceased being universal, but no replay occurred; report pending | H23 receives behavioral support but intention/enjoyment remain unresolved | Were builds ecological reasoning or experiment coverage, and did mature exposure create fun or only legibility? |
| [008 — Catalyst Ecology](../experiments/008_catalyst_ecology/README.md) | Can a four-choice compositional system remain interesting across small, durable, and spawning populations? | Building was enjoyable but meaningless: mature effects became clear, yet almost any build/baseline seemed sufficient and the cap carried no felt tradeoff | H23 down as implemented; H24 added with strong support | Can capability-sensitive escalation give qualitative building weight without prescribing one answer? |
| [009 — Catalyst Ascent](../experiments/009_catalyst_ascent/README.md) | Does capability-sensitive escalation make an enjoyable build meaningful through several causal routes? | Ready for playtest | Tests H24 while separating consequence from raw health inflation | Does a limitation create anticipation, tactical expression, and another build question—or only compulsory throughput? |

View file

@ -180,9 +180,19 @@ Confidence is deliberately qualitative until there is playtest evidence.
## H23 — Changing enemy ecology can preserve compositional reasoning
- **Confidence:** unresolved; first 008 run was too short at build maturity
- **Confidence:** low as implemented; ecology was legible but not consequential
- **Evidence for:** Without prompting, the player concluded that Bloom excelled against small enemies while Arc performed better against large enemies. Conditional target profiles can make component knowledge transfer while preventing one exact build from answering every field.
- **Evidence against:** Experiment 003's changing requirements collapsed into obvious prescribed counter-parts. In 007, taking all mutation families together may dominate, and all-Bloom's aggressive scaling may erase ecology distinctions through raw power.
- **Experiments:** 003 is negative adjacent evidence; 007 generated the hypothesis; 008 first run supplied only one short field after the fourth choice and could not isolate it.
- **Evidence update:** In 008 the player used three different builds and believed multiple approaches may exist, but reported that each expedition ended just as something cool began emerging. The fourth module was active for only 722 seconds. Conduit-first was a misunderstanding, not planned interface composition.
- **Revision 2 evidence:** Added mature fields made choice effects easier to see and supported small movement-strategy refinements. Builds differentiated and Bloom+Arc stopped being universal. However, the player believed almost any build or even baseline fire could complete every expedition; no exclusion created a relevant limitation, so the cap generated little meaningful tradeoff.
- **Interpretation:** Population variety cannot preserve reasoning when success is insensitive to the composition. Simply raising difficulty remains an untested and potentially confounded correction.
## H24 — Enjoyable building needs capability-sensitive consequence
- **Confidence:** high as a requirement; the suitable consequence remains unresolved
- **Evidence for:** In 008 revision 2, longer exposure made three causal builds legible and building was explicitly reported enjoyable. It simultaneously felt meaningless because nearly any build or baseline fire appeared sufficient. Earlier 000/001/003 failures also lacked consequences sensitive to deeper construction, while 002's survival/wave stakes gave upgrades instrumental value before one strategy dominated.
- **Evidence against:** Pure sandbox building can be enjoyable without external failure when the output is expressive enough. Experiment 007 generated repeated build tests from surprise and spectacle alone, so hard challenge is not always required.
- **Experiments:** 000, 001, 002, 003, 007, 008
- **Unresolved:** Can capability-sensitive enemies make exact qualitative building matter through multiple causal routes without collapsing into obvious counters or raw throughput pressure?
- **Unresolved:** Can mixed/behavioral ecologies support several viable causal compositions, or does adaptation become “equip the labeled counter”?