Initial Commit

This commit is contained in:
ookami125 2026-08-18 00:52:29 -04:00
commit 35f3810632
90 changed files with 29267 additions and 0 deletions

View file

@ -0,0 +1,123 @@
# Experiment 007 Preliminary Analysis — Session 56ccf2a6
Status: complete from updated telemetry plus player report.
Source: `JSONL/catalyst-trials-56ccf2a6-9613-4da1-86aa-3c5d61797784 (1).jsonl` supersedes the earlier partial save with the same session ID.
## Session Structure
- 2,900 events over 1,053.8 elapsed seconds, including long out-of-game gaps between later runs.
- The displayed mapping was Trial I = `mutation`, Trial II = `amplification`.
- The player completed Trial I, saved, completed Trial II, saved, then voluntarily completed three more Trial I runs with different builds.
- All 25 fields were cleared on their first attempt. There were no defeats and only two total player-damage events.
- Each of the five runs defeated the same 69 enemies. This confirms matched coverage; it is not enjoyment evidence.
## Run Summary
| Run | Condition | Final build | Active trial time | Final action profile |
|---|---|---|---:|---|
| Trial I | Mutation | Fork 1, Arc 3 | 47.7 s | 153 primary shots; 96 fragment hits; 32 Arc triggers |
| Trial II | Amplification | Rate 2, Width 1, Power 1 | 52.2 s | 288 primary shots; no secondary effects |
| Trial I replay | Mutation | Arc 4 | 60.0 s | 187 primary shots; 36 Arc triggers; no Fork/Bloom |
| Trial I combined replay | Mutation | Bloom 2, Fork 1, Arc 1 | 48.7 s | 58 fragment hits; 66 spark hits; 23 Arc triggers |
| Trial I Bloom replay | Mutation | Bloom 4 | 54.0 s | 140 primary shots; 136 spark hits |
Kill attribution:
- First mutation run: 26 primary, 32 fragment, 11 Arc.
- Amplification run: 69 primary.
- Arc-only replay: 47 primary, 22 Arc.
- Combined replay: 30 primary, 22 fragment, 11 spark, 6 Arc.
- Bloom-only replay: 42 primary, 27 spark.
The first mutation build produced a real causal combination: Fork fragments counted toward the hit cadence that triggered Arc. It also cleared later fields quickly, but speed and effect count cannot establish whether the combination felt satisfying, strategically authored, or merely powerful.
## Choice Sequence and Deliberation
| Run | Choices | Approximate deliberation per choice |
|---|---|---|
| First mutation | Fork → Arc → Arc → Arc | 4.0 s, 8.8 s, 6.0 s, 1.0 s |
| Amplification | Rate → Width → Power → Rate | 10.0 s, 3.8 s, 2.9 s, 2.5 s |
| Mutation replay | Arc → Arc → Arc → Arc | 1.0 s, 0.9 s, 0.6 s, 0.6 s |
The replay pattern is unusually specific. It began about five seconds after both displayed trials had been completed and the second save had succeeded. The near-immediate repeated Arc choices suggest a preformed test of pure stacking rather than ordinary indecision or accidental continuation. This is an inference; the player's motive is not logged.
Bloom was never selected. Telemetry cannot distinguish an unattractive description, an apparently weak effect, deliberate focus on the other interaction, or simple exhaustion of available decisions.
## Preliminary Interpretation
Experiment 007 produced the strongest behavioral evidence so far for testing a **capability trajectory** rather than only an isolated mechanic. Unlike the optional replays in 004 and 006, the third run changed a four-decision build to isolate one upgrade family's scaling after the player had already seen a mixed interaction build and the numerical condition.
That distinction is promising but still insufficient. Three alternative explanations remain live:
1. Arc stacking created genuine anticipation or payoff and motivated another run.
2. The player was analytically checking what maximum Arc did, without enjoying the shooting or result.
3. The completion screen's “other trial” action or the short run length made another diagnostic pass feel cheap enough to perform despite boredom.
The first mutation run also cannot yet be labeled authored synergy. Fork was selected before Arc, and its fragments did feed Arc, but the player may not have predicted or noticed that relationship. Repeated Arc could mean they valued the interaction, believed Arc was simply strongest, or wanted to remove Fork as a confound.
The amplification sequence sampled all three dimensions before returning to Rate. Its first choice took the longest deliberation of the session, but later choices accelerated. That could reflect learning the menu, an actual tradeoff, or declining care.
No telemetry pattern establishes that either condition reversed stop desire. The player saved after every complete run, which is helpful data hygiene but may also indicate they viewed each run as a required experiment.
## Follow-up Needed
1. Why did the player voluntarily replay Trial I with Arc selected four times, and did the result feel rewarding or merely answer a test question?
2. In the first Trial I run, did they intend Fork fragments to feed Arc, and why was Bloom never selected? In Trial II, were choices part of a plan or just apparent strength/coverage?
3. When did they first want to stop in each displayed trial, and did any upgrade choice or realized effect make them want to see the next field?
## Initial Player Report
The Arc-only replay was a deliberate test of how strong Arc could become without support. It felt weak alone. The player also performed a combined run with Bloom → Fork → Arc → Bloom and observed how well all three mutation families played off one another. The combined build felt “a lot better than expected.”
The first mutation sequence was not a predicted Fork→Arc plan. Fork was chosen because its description sounded interesting, then seemed weak. Arc was chosen next because it was the next interesting description. Bloom was initially skipped because the player incorrectly predicted its value; after trying it, they considered it the best of the three and regretted skipping it.
This wrong prediction is important. The positive result did not come from merely executing a plan described by the menu. The player sampled effects from an inaccurate prior, observed cross-effect behavior that exceeded expectation, revised the ranking of the options, isolated Arc in a replay, and separately tested Bloom/the combined system. That is the first clear multi-step loop in the project of:
> expectation → chosen test → surprising interaction → revised model → another build test
Trial II felt more boring as soon as the player finished reading its choices. They described its upgrades as things that would make Trial I's effects more fun, but not as effects that played off one another to create interesting differences. Trial I was interesting enough to replay.
## Revised Interpretation
Experiment 007 gives strong evidence for H21's qualitative-composition component and against the broader idea that any chosen power growth is equivalent. The matched numerical upgrades were recognized as useful but causally independent; their descriptions were enough to predict boredom before use. The mutation upgrades initially appeared individually weak, yet their interactions created an unexpectedly better result and motivated multiple unscripted builds.
The likely valuable property is not simply “qualitative upgrades” or spectacle. Each mutation changes the opportunity surface of the others:
- Fork creates extra hits, increasing Arc frequency.
- Fork and Arc can create kills, increasing Bloom emissions.
- Bloom sparks add hits, feeding Arc again.
- Arc kills can produce more Bloom sparks.
This is complexity through narrow causal interfaces, closely matching the original composition hypothesis. The components remain understandable alone, but their products cross the same hit/kill boundaries. A selection therefore changes both immediate output and the future value of other selections. Trial II's damage/rate/width choices changed throughput without creating new relationships.
The surprising reversal around Bloom is stronger evidence than merely choosing a preferred build. The menu did not make the best-feeling combination obvious, and the player learned by using it. This may be the first prototype where knowledge transferred into a new self-selected test with an answer that affected valued capability rather than an abstract marker.
Important limits remain:
- “Interesting” and “better than expected” do not yet establish that firing and movement became enjoyable in themselves.
- The combined build may have won through spectacle or raw crowd-clearing power rather than prediction/composition.
- The run was short and offered only three families, so strategy half-life is unknown.
- Bloom's homing reduces aim burden and may simply have made combat easier; this is bundled with its interaction role.
- Bloom scaling and field-clearing spectacle remain bundled; telemetry can describe the cascade but cannot say which property caused enjoyment.
## Continuation Report and Final Result
Across the continuation, the player performed combined and all-Bloom comparisons after the Arc-only run. This resolves the endpoint ambiguity: the first surprising combination generated further specific questions rather than merely concluding the session.
All-Bloom appeared to scale too aggressively, but the player described watching the field disappear after a few shots as “kind of fun to watch.” This is the first explicit positive enjoyment report tied to a prototype consequence, although it remains qualified and partly bundled with overtuned spectacle.
The player also began classifying the mutation families by enemy ecology: Bloom seemed strong against groups of small enemies, while Arc seemed strong against large enemies. For a short prototype, moving between skills and observing those different target profiles was interesting. This indicates that the value was not only a generic cascade. The player was learning a conditional capability map and using repeated builds to compare it.
Experiment 007 is the first successful probe in the program. It supports the following mechanism:
> A deliberately chosen component becomes interesting when it participates in legible causal chains, changes the value of other components, and produces a materially different capability against a recognizable problem class.
The positive loop contained all of the desired stages: inaccurate expectation, chosen test, surprising outcome, model revision, a new build question, repeated test, conditional knowledge about target types, and a visible power payoff.
The next experiment should preserve the hit/kill interface network and test its **strategy half-life**. It should introduce changing enemy ecologies and a limited composition budget so the player cannot simply take the full Fork+Arc+Bloom package every time. The critical comparison is between:
- adaptive composition using knowledge that transfers across fields; and
- obvious “swarm means Bloom, brute means Arc” counter-loadout work, repeating Experiment 003.
Do not merely nerf Bloom. Its aggressive scaling may be part of the observed payoff. Instead, cap runaway recursion enough to keep other builds observable, retain satisfying collapse, and vary mixtures/behaviors so several causal routes remain plausible.