Step 3 pilot (N=2, exp-ca-rule-multigen-poc) showed provisional positive: agents propose CA-axis modifications (7/10 of accepted), reasoning grounds in CA observations (gliders, density), step-2 LLM-only chain partially disrupts. Question: how variable is this across diverse seed conditions? Does the CA-axis proposal preference replicate? Is the step-2 chain disruption consistent or pilot-specific? N=10 with 5 diverse seed sets × 2 rotations. Same engine (ca_rule_multigen) as pilot. Pre-registered: - CA-axis proposal share: ≥50% of accepted proposals are on CA axes (vs pilot 70%) - signal_type chain: ≤30% of runs reach 'concept' endpoint (vs step-2 50%, pilot 0%) - reasoning grounding: ≥60% of proposal reasoning explicitly references CA state (gliders, density, structures, patches)