← back to timeline

EXP-IVB-NEUTRAL

informative

L5 IVB Ablation — Neutral Framing Control for Competition Pilot

2026-03-29 L5 level 1 runs $0.00
new
premortem
mock
real
metrics
analyze
review
krit
done

In plain language

Researchers tested if large language models naturally prefer cooperation or if their behavior is simply shaped by how they are instructed. They set up a scenario where agents interacted without any explicit instructions to cooperate or compete, only describing their individual traits. The experiment was too small to draw any conclusions about whether the models have a default cooperative tendency or if their actions are purely a result of the prompts they receive.

What we found

UNDERPOWERED (N=1): No predictions scored

Predictions we made before running

0/3 confirmed

Technical details

Research hypothesis

IVB ABLATION — NEUTRAL FRAMING CONTROL

Zhang 2603.23406 claims LLMs have innate progressive bias that overrides presets. If true, cooperation > competition in exp-competition-pilot may be LLM artifact, not emergent property.

THIS EXPERIMENT: same 8-agent merge as exp-cp-cooperate and exp-cp-compete, but with NEUTRAL framing — no cooperation or competition directive. Agents described only by their cognitive tendencies, no inter-group goals.

ABLATION LOGIC:

  • If neutral ≈ cooperation (TR_diff~0.25) → IVB confirmed: cooperation is default
  • If neutral ≈ competition (TR_diff~0.12) → IVB refuted: framing drives outcomes
  • If neutral ≠ both → three-state landscape, framing modulates but neither is default

CRITICAL FOR PAPER2: determines whether §discussion needs IVB caveat or IVB defense.

Uses IDENTICAL seeds from Phase A+B to maintain comparability. Same model, temperature, rounds, agent count.

THEORETICAL GROUNDING:

  • Zhang 2603.23406: innate value bias (IVB) in LLMs
  • Kulveit 2603.11353: interviewer expectations bleed into AI responses
  • Our exp-competition-pilot: cooperation TR_diff=0.257 vs competition TR_diff=0.117

THIS IS EXPLORATION TIER (N=1). Goal: does neutral framing replicate cooperation or competition patterns?

Experimental setup

Type: single_condition

Condition Parameters
NEUTRAL interaction: LIVE, persona: PERSONA

Factors: interaction (LIVE) × persona (PERSONA)

Parameters

model
gemini-2.5-flash-lite
n_agents
8
n_rounds
15
scheduler
round_robin
temperature
0.9

Trophic Ratios by Condition

Mean trophic ratio per agent across runs. Error bars = ±1 std dev. Higher TR = more upstream (exporter).