← back to timeline

EXP-BALANCED-ROT

informative

Balanced Rotation Latin Square — N=5 Position vs Persona Decomposition

2026-04-02 L3 level paper3 15 runs $0.13
new
premortem
mock
real
metrics
analyze
review
krit
done

In plain language

Five AI agents with different thinking styles form hierarchies when they converse. But we found that who speaks FIRST matters more than personality at N=3. At N=5, it's unclear — the data is confounded because we always used the same speaking order. This experiment systematically rotates who speaks first to finally answer: is hierarchy about position, personality, or random?

What we found

No predictions scored

Predictions we made before running

0/5 confirmed

Technical details

Research hypothesis

BALANCED ROTATION DESIGN: Position vs Persona Decomposition at N=5

CONTEXT: N=3 competency-rotation: position DOMINATES (η²=0.60 >> persona η²=0.25). N=5 pooled analysis (14 runs): η²(pos)=0.14, η²(per)=0.11 — BOTH WEAK, 75% unexplained. BUT massive confound: r(pos, persona_id)=0.79 across dataset (12/14 runs have identical position-persona mapping). Cannot decompose effects with existing data.

DESIGN: Latin square: 5 cyclic rotations × 3 seed sets = 15 runs. Each persona appears in each position EXACTLY ONCE per seed set. Rot0: [Alpha,Beta,Gamma,Delta,Epsilon] (standard) Rot1: [Beta,Gamma,Delta,Epsilon,Alpha] Rot2: [Gamma,Delta,Epsilon,Alpha,Beta] Rot3: [Delta,Epsilon,Alpha,Beta,Gamma] Rot4: [Epsilon,Alpha,Beta,Gamma,Delta]

This gives r(pos, persona_id)=0.00 by construction — ORTHOGONAL design. Enables clean 2-way ANOVA: position × persona → η² decomposition.

KEY QUESTION: Is N=5 hierarchy position-driven, persona-driven, or stochastic?

Dochkina 2603.28990: position→hierarchy is universal LLM property. N=3 confirmed: last speaker = sink 100%, η²(pos)=0.60. N=5: does this weaken (dilution) or does persona emerge above noise?

Experimental setup

Type: factorial

Condition Parameters
LIVE interaction: LIVE, n_agents: 5, n_rounds: 70, note: Standard N=5 live interaction. Latin square rotations via run_matrix.

Factors: rotation (0, 1, 2, 3, 4) × seed_set (S1, S2, S3)

Parameters

model
gemini-2.5-flash-lite
n_agents
5
n_rounds
70
n_runs_per_condition
15
scheduler
round_robin
temperature
0.9

Trophic Ratios by Condition

Mean trophic ratio per agent across runs. Error bars = ±1 std dev. Higher TR = more upstream (exporter).