← back to timeline

EXP-ENGINEER-REPL

success

Replication of EXP-ENGINEER — 1 Synthesizer in N=5, N=4 Runs

2026-03-09 L3 level paper3 4 runs
new
premortem
mock
real
metrics
analyze
review
krit
done

In plain language

One synthesizer among five agents halved the hierarchy (TR 0.240). But that was a single run. We repeat four times: is this effect stable, or were we just lucky?

Predictions we made before running

3/4 confirmed · 1 refuted

Technical details

Research hypothesis

EXP-ENGINEER found that a single synthesizer (20% integrative) halves hierarchy: TR range 0.528 → 0.240 (54.5% reduction). This was a SINGLE run.

Replication question: Is TR ≈ 0.240 stable across multiple runs, or was it a lucky draw from high-variance stochastic dynamics?

EXP-DOSE40-SINK showed that 2 sinks COMPETE and REDUCE suppression (TR=0.404), confirming competitive exclusion. The 1-sink optimum is the key finding.

This experiment replicates EXP-ENGINEER exactly (same config, same personas, same seeds, same model) with N_RUNS=4 to establish: 1. Mean and SD of TR range across runs 2. Whether TR < 0.264 (our P1 threshold) holds reliably 3. 95% CI for the 1-sink suppression effect

Experimental setup

Type: single_condition

Condition Parameters
PERSONA_LIVE persona: PERSONA, interaction: LIVE

Factors: persona (PERSONA) × interaction (LIVE)

Parameters

n_agents
5
n_rounds
50
n_runs_per_condition
4
model
gemini-2.5-flash-lite
temperature
0.9
scheduler
round_robin

Trophic Ratios by Condition

Mean trophic ratio per agent across runs. Error bars = ±1 std dev. Higher TR = more upstream (exporter).