← back to timeline

EXP-BASE-MODEL

success

Base-Model Control: Scrambled-Persona Test for Hierarchy Emergence

2026-03-26 L3 level paper3 6 runs $0.04
new
premortem
mock
real
metrics
analyze
review
krit
done

In plain language

Testing whether the social hierarchy we observe in AI agent groups is a real emergent phenomenon or just an artifact of how the AI was trained. We scramble the agents' personality descriptions into incoherent word salad while keeping everything else identical. If hierarchy still appears with scrambled personalities, it's a robust phenomenon — not dependent on coherent role-playing.

What we found

UNDERPOWERED (N=3): No predictions scored

Predictions we made before running

0/4 confirmed

Technical details

Research hypothesis

ADVERSARIAL CONTROL: Is trophic hierarchy an artifact of instruction-tuning + coherent personas, or does it emerge from interaction dynamics alone?

All ~90 existing experiments use instruction-tuned Gemini with coherent cognitive personas (analytical, associative, concrete, etc.). If hierarchy requires coherent persona differentiation (an RLHF-dependent behavior), then scrambling persona content while preserving token count should abolish it.

TUNED condition: standard pent70 setup (replication control). SCRAMBLED condition: same model, same framing, but persona descriptors cross-mixed between agents — destroying within-agent cognitive style coherence while preserving total vocabulary, token count, and interaction structure.

This is the strongest feasible adversarial test without access to base (non-RLHF) models. Google API exposes only instruction-tuned variants. Scrambled-prompt control tests the weaker but still important hypothesis: does hierarchy require COHERENT persona differentiation, or does any prompt differentiation (even incoherent) suffice?

NOTE: A true base-model test (no RLHF at all) would be stronger but requires API access to non-instruction-tuned weights. Cross-model test (gemini-2.0-flash) is a planned follow-up experiment (exp-cross-model).

Experimental setup

Type: factorial

Condition Parameters
TUNED_LIVE persona: PERSONA, interaction: LIVE, n_agents: 5, n_rounds: 70, note: Standard pent70. Coherent cognitive personas. Replication control.
SCRAMBLED_LIVE persona: PERSONA, interaction: LIVE, n_agents: 5, n_rounds: 70, note: Scrambled personas. Same words cross-mixed. Incoherent cognitive style.

Factors: persona_condition (TUNED, SCRAMBLED)

Parameters

n_agents
5
n_rounds
70
n_runs_per_condition
3
model
gemini-2.5-flash-lite
temperature
0.9
scheduler
round_robin

Trophic Ratios by Condition

Mean trophic ratio per agent across runs. Error bars = ±1 std dev. Higher TR = more upstream (exporter).