← back to timeline

EXP-MMP-GPT4O

partial

Multi-Model Pilot: GPT-4o-mini

2026-04-13 L3 level paper3 6 runs $0.05
new
premortem
mock
real
metrics
analyze
review
krit
done

In plain language

Does hierarchy emerge in GPT-4o-mini? Identical agents, 30 rounds, LIVE vs ISOLATED.

Technical details

Research hypothesis

Multi-model pilot: does trophic hierarchy (TR ordering), VP_excess positivity, F0 magnitude, and LIVE>ISOLATED role persistence generalize across LLM substrates (Claude, GPT-4o-mini, Gemini)? Identical experimental protocol across 3 models, 2 conditions (LIVE/ISOLATED), N=3 agents, 30 rounds. Pre-registered predictions: P1: TR ordering consistency (Kendall W > 0.5 across runs within model), P2: VP_excess > 0 in all LIVE conditions (universal), P3: F0 magnitude converges across models (CV < 0.5), P4: LIVE > ISOLATED role persistence (effect present in all models).

Experimental setup

Type: multi_condition

Condition Parameters
LIVE label: Live interaction (shared memory), interaction: LIVE, persona: IDENTICAL, n_runs: 3
ISOLATED label: Isolated (private memory per agent), interaction: ISOLATED, persona: IDENTICAL, n_runs: 3

Parameters

n_agents
3
n_rounds
30
model
openai/gpt-4o-mini
temperature
0.9
scheduler
round_robin

Trophic Ratios by Condition

Mean trophic ratio per agent across runs. Error bars = ±1 std dev. Higher TR = more upstream (exporter).

multi-model cross-substrate pilot vocabulary-propagation