Experiment Timeline

279 experiments across 7 levels · 29 series · 21 confirmed · 25 partial · 14 refuted

view: by level · by recency →

L1 — Chemistry

9 exp
EXP-HARD-AUTOCATALYTIC-CORE-LANGUAGE-SUBSTRATE-PORT pending

Researchers are investigating whether a specific type of artificial language can evolve in a way that favors stable, persistent forms over random changes. They are testing if the observed changes are due to a genuine selection process favoring stability, or if they are simply random fluctuations. The results are currently pending.

2026-05-30 → premortem
GENEALOGICAL-REPLICATOR-VS-OPEN-ENDED-PROBE partial 2 experiments

Linka asks: if we evolve text by repeatedly selecting from many small variations of a single starting sentence, does the substrate (the local LLM) produce a stable self-replicator, endless novelty, a cycle, or just dead text — and do five different selection rules produce five different trajectory shapes, or all the same one?

2026-05-24 10 runs total ✓ done
EXP-DENSITY-MIN artifact

Does hierarchy emerge immediately from the first interaction, or does it need "warm-up" time? Three agents with different thinking styles — but only 5 turns each instead of the usual 20. If hierarchy still appears — the structure is instantaneous.

2026-03-10 3 runs 0/3 confirmed ✓ done
EXP-SEED-SENSITIVITY artifact

With six agents, hierarchy is unstable — sometimes strong, sometimes weak. Is this due to speaking order (whoever speaks last gets suppressed) or does the seed text determine which "regime" the system falls into? 16 runs with different texts and orders separate these two explanations.

2026-03-10 16 runs 0/6 confirmed ✓ done
EXP-MAY-SIGMA success

With 7 agents and 7 rounds each, hierarchy collapses (TR=0.30). But according to May, the issue may not be the NUMBER of agents, but the STRENGTH of interaction. If each agent speaks 4 times instead of 7, interaction is weaker → hierarchy may recover. If so — we confirm the May mechanism. If not — the issue is the number of agents itself, not the strength of their mutual influence.

2026-03-10 1 runs 3/4 confirmed ✓ done
PWS-FACTORIAL negative

If you give an advantage to "older" text fragments (those that survive longer), does selection emerge? No — age-weighting does not create evolutionary pressure. The system remains random.

2026-03-07 0/7 confirmed ✓ done
AMA informative

Can one text fragment "catalyze" another — like an enzyme speeds up a biochemical reaction? We tested this: LLMs can conditionally copy text, but this is not true catalysis — it is interpretation, not a chemical reaction.

2026-03-07 1/3 confirmed ✓ done
CVS informative

We tried "natural selection" for text fragments — those that copy better survive. It turned out the only working mechanism is literal word matching (template matching). Neither text length nor "complexity" helps survival. Only exact repetition.

2026-03-05 100 runs 2/4 confirmed ✓ done
EXP-PARTICLE-NULL-R0 success 5 experiments

Researchers tested if simple text generators could mimic complex language patterns. They found that these basic models, without understanding meaning, could not produce the expected results. Therefore, the experiment was inconclusive.

5 runs total 0/3 confirmed ✓ done

L1.5 — Autopoiesis

2 exp
EXP-FORGET-N5 informative 3 experiments

Five AI agents build a shared text, but each is told to focus only on the last 2-3 entries. At N=3, this destroyed hierarchy completely (TR=0.015). Does the same instruction destroy hierarchy at N=5? This fills the missing cell in our memory × group-size matrix.

2026-03-28 3 runs total 0/5 confirmed ✓ done
EXP-FORGET-SERIES informative 3 experiments

Three AI agents communicate in a text space, but are told to "pay attention only to the last 2-3 entries." If hierarchy persists — forgetting does NOT destroy structure, and may even support it. We compare with agents that have full memory.

2 runs total ✓ done

L2 — Agent

10 exp
EXP-GAMMA-ABLATION-N4 informative

Gamma always wins in groups of 4 — but is it because Gamma grounds abstract ideas (thinking style) or because of speaking order? We swap Gamma's and Alpha's roles to find out.

2026-04-04 3 runs 0/4 confirmed ✓ done
EXP-N4-ENTROPY success

Groups of 3 AI agents choose winners almost at random (lottery). Groups of 5 consistently crown the same winner. What happens at 4? We run the same experiment 9 times to find out — resolving an artifact where single runs always show 100% winner.

2026-04-04 9 runs 4/4 confirmed ✓ done
EXP-CLOSURE success

We've shown that differently-minded AI agents form stable hierarchies when communicating — but always with the same abstract poetic text. What if we give them mundane everyday topics instead? If the same hierarchy appears, it means the structure comes from the agents' interaction, not from the content.

2026-03-16 3 runs 0/4 confirmed ✓ done
EXP-PERSONA-SWAP artifact 4 experiments

Agent Zeta (the "generalizer") consistently occupies the top of hierarchy. But why — because of the thinking type (generalization creates vocabulary that others adopt) or because of the name/position? We swap the roles of Zeta and Alpha: if the generalizer is on top again — it is the role, not the name.

2026-03-10 4 runs total 0/5 confirmed ✓ done
EXP-DILUTION-N7 informative 4 experiments

With 7 agents, hierarchy collapsed (TR=0.30 vs 0.53 with 5 agents). But maybe it is not about the number of agents, but that each got too few rounds? We give each of the 7 agents 14 rounds instead of 7 — if the issue is insufficient rounds, hierarchy should recover.

2026-03-10 4 runs total 0/4 confirmed ✓ done
EXP-OPPONENT informative

What shapes hierarchy more — cooperation or competition? Three AI agents with adversarial styles (critic, contrarian, skeptic) interact in a shared text space. We compare with the same agents but with complementary styles (analyst, associator, practitioner).

2026-03-09 1 runs 0/4 confirmed ✓ done
VOCAB-SPECIFICITY negative

What determines an agent's role — its name or its instruction? We swapped directives: agent Beta received Gamma's instruction ("ground abstractions in concrete examples"). Beta became the "leader" — role follows the directive, not the label. Three levels: generic description → style → concrete action. Only action creates structure.

2026-03-08 45 runs 0/5 confirmed ✓ done
EXP-DOSE-RESPONSE informative 6 experiments

All five agents receive identical instructions — differing only in names. If hierarchy still emerges, it is generated by the turn-order structure itself (who speaks first/last), not by differences in personalities. This is a fundamental test: is hierarchy a property of structure or of content?

11 runs total ✓ done
EXP-TRIAD-DOSE informative 4 experiments

Three identical AI synthesizers — same instructions, only names differ. Any hierarchy that emerges is purely from speaking order, not personality. This is the floor of the N=3 dose-response curve: how much structure survives when all agents are clones?

8 runs total ✓ done
EXP-ASYM-SERIES informative 14 experiments

Give three AI agents different cognitive styles (analyst, associator, practitioner) and a stable hierarchy emerges. The practitioner (who "grounds" abstractions) becomes the word source, while the associator becomes the receiver. Neither personas without interaction nor interaction without personas create structure — you need both.

398 runs total ✓ done

L3 — Tissue / Organ

38 exp
EXP-SCARCITY-AWARE partial

Agents see their remaining budget. Does awareness of pressure enable adaptation?

2026-04-14 6 runs ✓ done
EXP-PERSONA-FREE-SCARCITY partial

Does hierarchy emerge from scarcity alone, without designed persona differentiation?

2026-04-13 9 runs ✓ done
EXP-SCARCITY-BUDGET-TIGHT success

Tighter budget forces agent death. Who dies first — vocabulary producers or reusers?

2026-04-13 6 runs ✓ done
EXP-SCARCITY-BUDGET-PILOT informative

What happens when agents have limited tokens? Tierra-style: budget exhaustion = death.

2026-04-13 2 runs ✓ done
EXP-HAWTHORNE-PILOT partial

Do AI agents behave differently when told they are being studied for hierarchy?

2026-04-13 10 runs ✓ done
EXP-MMP-CLAUDE partial

Does hierarchy emerge in Claude Haiku? Identical agents, 30 rounds, LIVE vs ISOLATED.

2026-04-13 6 runs 0/4 confirmed ✓ done
EXP-MMP-GEMINI success

Does hierarchy emerge in Gemini Flash Lite? Identical agents, 30 rounds, LIVE vs ISOLATED.

2026-04-13 6 runs ✓ done
EXP-MMP-GPT4O partial

Does hierarchy emerge in GPT-4o-mini? Identical agents, 30 rounds, LIVE vs ISOLATED.

2026-04-13 6 runs ✓ done
EXP-STAR-PERTURB pending

This experiment investigates whether a system's boundary, called a Markov blanket, can actively maintain itself or if it's just a passive structure. Researchers removed a central agent from a star-shaped network to see if the remaining agents would reorganize to form a new, self-maintaining boundary or if the structure would simply collapse. The results are still being analyzed.

2026-04-04 → premortem
EXP-N3-BALANCED-ROT informative

Three AI agents with different thinking styles talk in a shared text space. We rotate who speaks first to separate the effect of personality from the effect of speaking order. At group size 5, personality-position COMBINATIONS matter (interaction effect). Does this interaction vanish at group size 3, where speaking order dominates? If yes, crossing from 3 to 5 agents is a threshold — like switching from rigid hierarchy to flexible self-organization.

2026-04-03 9 runs ✓ done
EXP-BALANCED-ROT informative

Five AI agents with different thinking styles form hierarchies when they converse. But we found that who speaks FIRST matters more than personality at N=3. At N=5, it's unclear — the data is confounded because we always used the same speaking order. This experiment systematically rotates who speaks first to finally answer: is hierarchy about position, personality, or random?

2026-04-02 15 runs 0/5 confirmed ✓ done
EXP-COMPETENCY-N5-ROT2 informative

We rotate who speaks first in a 5-agent conversation. Reanalysis of 13 runs shows N=3 position dominance was an artifact — at N=5, 70% of variance is noise. Adding a third rotation to fill the position×persona matrix and resolve whether any systematic structure exists.

2026-04-02 1 runs ✓ done
ASSEMBLY-INDEX-PILOT partial

Researchers studied how often word pairs appeared together in generated text, comparing it to random chance. They found that word pairs appeared together more often than expected, suggesting the text has a meaningful structure. The study is ongoing, but early results indicate that this structure becomes more complex over time.

2026-03-31 14:36:00+00:00 1 runs ✓ done
CROSS-SUBSTRATE-F0 partially_supported

Researchers tested if a measure of how connected different systems are, called F₀, would rank them in the same order whether they were looking at networks of interacting agents or chemical reactions. They found that the order of connectedness was different for the agent networks than what was predicted, and while the general measure of connectedness was similar between the two types of systems, specific predictions did not transfer. This suggests that how systems are organized matters for their overall connectedness, and this organization behaves differently in different contexts.

2026-03-29 1/3 confirmed → premortem
EXP-BASE-MODEL success

Testing whether the social hierarchy we observe in AI agent groups is a real emergent phenomenon or just an artifact of how the AI was trained. We scramble the agents' personality descriptions into incoherent word salad while keeping everything else identical. If hierarchy still appears with scrambled personalities, it's a robust phenomenon — not dependent on coherent role-playing.

2026-03-26 6 runs 0/4 confirmed ✓ done
EXP-BROADCAST-PILOT artifact

Three AI agents communicate, but only one — Alpha — generates words that others can read. Beta and Gamma read Alpha but cannot see each other and nobody reads them. Alpha is a "broadcaster": pure source, zero input. Will this create the strongest hierarchy — or will isolation from feedback make Alpha drift into irrelevance?

2026-03-24 1 runs 0/5 confirmed ✓ done
EXP-SEED-ALT artifact

All our experiments used abstract, poetic seed texts. What if we use concrete, scientific seeds instead? Same three AI agents, same rules — but different starting material. If the hierarchy changes, our findings depend on WHAT agents discuss, not just HOW they're connected.

2026-03-23 1 runs 0/5 confirmed ✓ done
EXP-BROADCAST success

Three AI agents communicate, but only one — Alpha — generates words that others can read. Beta and Gamma read Alpha but cannot see each other and nobody reads them. Alpha is a "broadcaster": pure source, zero input. Will this create the strongest hierarchy — or will isolation from feedback make Alpha drift into irrelevance?

2026-03-22 1 runs 0/5 confirmed ✓ done
EXP-REVORD-N5 informative

We reversed who speaks first in a 5-agent conversation. Cycle 1's hierarchy perfectly inverted (pure position artifact), but in later cycles Alpha stayed on top regardless of position. The "analytical" persona creates more exportable language — it's cognitive style, not turn order.

2026-03-21 1 runs 0/4 confirmed ✓ done
EXP-REG3-N5 artifact

Re-run with 50 rounds revealed the entire density-hierarchy inverted-U was an artifact. Cycle 1 burn-in (universal startup effect) inflated early TR measurements. After correction, hierarchy is FLAT across all densities (25-100%) at N=5. Kruskal-Wallis p=0.98. The density gradient doesn't control hierarchy — it never did.

2026-03-20 1 runs 0/5 confirmed ✓ done
EXP-REG2-N5 informative

Five AI agents, each reading exactly two others in a structured ring pattern. More connected than a simple cycle (1 neighbor) but less than full visibility (4 neighbors). The key test: does DOUBLING connectivity change hierarchy? Or do you need unequal connections (hubs) for structure to emerge? The missing data point for making the density-hierarchy relationship precise.

2026-03-19 1 runs 0/5 confirmed ✓ done
EXP-CHAIN-N5 artifact

Five AI agents form a chain: Alpha writes freely, Beta reads only Alpha, Gamma reads only Beta, Delta reads only Gamma, Epsilon reads only Delta. Information flows ONE WAY through a 5-link chain. No feedback. With 3 agents, the chain created the strongest hierarchy gradient. Does this scale? Or does the signal die before reaching the end?

2026-03-18 1 runs 0/5 confirmed ✓ done
EXP-FULL-N5 artifact

Five AI agents, each sees all others — a complete graph. With three agents, full connectivity created the STRONGEST hierarchy (TR=0.444). But with five, the "star" collapsed (TR=0.059). Will the complete graph collapse too? The third data point for topological comparison — key to the crossover mechanism.

2026-03-15 1 runs 0/5 confirmed ✓ done
EXP-CYCLE-N5 artifact

Five AI agents in a ring: each reads only the previous one. A→B→C→D→E→A. With three agents, the ring killed hierarchy (all equal). What about five? A longer cycle — more room for leaders or even flatter? Filling in the last cell of the 2×2 design (topology × group size).

2026-03-15 1 runs 0/5 confirmed ✓ done
EXP-PENT70-R0 series_member 12 experiments

Testing why experiments with 7 agents are more stable than with 5. Hypothesis: it is not about the number of agents, but the number of rounds per agent. We give 5 agents the same number of rounds per agent (14) as the seven had.

2026-03-10 12 runs total 0/2 confirmed → premortem
EXP-HYSTERESIS design_killed

Hysteresis is a classic test for a first-order phase transition: if you increase a parameter and return it back, the system "sticks" in the new state. But our system has no memory between experiments — each run starts from a clean slate. Therefore the "path up" and "path down" would yield the same result. The test is impossible without platform restructuring. Instead of hysteresis, we propose a sensitivity test for initial conditions: different seed texts at N=6 → if results are bimodal, this supports first-order; if unimodal — the transition is smooth.

2026-03-10 ✓ done
EXP-N6 artifact

Hierarchy peaks at 4-5 agents; at 7 it already drops. But is the drop sharp or gradual? Six agents — the missing point on the curve. If hierarchy still holds at 6, the cutoff is sharp; if it already drops — gradual.

2026-03-10 4 runs 0/5 confirmed ✓ done
EXP-HEPT partial

Seven agents is already too many for clear hierarchy. The optimum is 4-5 agents: enough for differentiation, but not so many that everything blurs. With 7 agents, middle positions overflow — five of seven cluster in a narrow range. Information flows actively (VP z=8.61), but stratification dissolves. The scaling curve is an inverted U.

2026-03-10 1 runs 2/4 confirmed ✓ done
EXP-QUAD2 success

Our quartet experiment showed flat hierarchy, but was it the even number of agents or the storyteller persona blurring the ranks? We rerun four agents with only non-integrative roles — the same skeptic and analyst who produced hierarchy at five — to find out which variable matters. Answer: the storyteller was flattening the system. Four agents hierarchy fine without one.

2026-03-09 1 runs 1/4 confirmed ✓ done
EXP-QUAD partial

Two agents can't form hierarchy. Three can. What about four? We add a fourth agent — a storyteller — to see if the pecking order deepens or if three levels is the ceiling for this kind of interaction.

2026-03-08 1 runs 0/4 confirmed ✓ done
EXP-DYAD informative

Two agents are not enough for hierarchy. Three is the minimum. Like in chemistry: you need a third element for a directed reaction to emerge. A pair just exchanges words without structure.

2026-03-08 4 runs 0/4 confirmed ✓ done
EXP-ASO success

Three identical AI agents communicate in a shared text field. Does a hierarchy emerge — who influences whom? Words do propagate between agents, but "who leads" is simply determined by who speaks first. Without differences between agents — no genuine structure.

2026-03-08 5 runs 3/5 confirmed ✓ done
EXP-SLC success

The LLM is not a neutral medium — it actively shapes the outcome. The same initial conditions produce different structures depending on which model interprets the text. The interpreter is a participant, not a tool.

2026-03-07 3 runs 1/3 confirmed ✓ done
EXP-TOPOLOGY informative 7 experiments

Three AI agents communicate in a text space. Information flows in one direction: Alpha writes — Beta reads only Alpha — Gamma reads only Beta. No feedback: nobody reads Gamma. Will a hierarchy gradient emerge — Alpha = word source, Beta = relay, Gamma = receiver? Or will unidirectional flow break convergence entirely?

6 runs total ✓ done
EXP-STAR-SERIES informative 6 experiments

Five AI agents in a star topology: one "hub" sees all four others, the rest see only the hub. With three agents, the hub became a word "absorber." What happens with five? Will the effect intensify or will something qualitatively new emerge?

17 runs total ✓ done
EXP-PENT-SERIES informative 5 experiments

Two agents: flat. Three: hierarchy. Four: flat again. Is it the odd number that matters? We test five agents — another odd number — to see if the pecking order returns. This time, no synthesizer roles that might blur the lines.

4 runs total ✓ done
EXP-ENGINEER-SERIES informative 3 experiments

Five agents built a clear pecking order. Now we replace one gap-tracker with a synthesizer — an agent that weaves everyone's threads together. Does one weaver flatten the entire hierarchy? Ecology calls this an ecosystem engineer: one species that reshapes the environment for everyone.

5 runs total ✓ done
EXP-KEYSTONE-SERIES informative 3 experiments

Three agents form a hierarchy: one sets the vocabulary, the others adopt it. We removed the leader — and expected hierarchy to collapse: after all, two agents without history (EXP-DYAD) are always flat. But no: the remaining two preserved hierarchy (TR diff 0.49 instead of 0.06), because shared memory carries an "organizational imprint." This is a signal of autopoiesis: the system structure is encoded in the text (substrate), not only in the number of agents.

108 runs total ✓ done

L4 — Organism

10 exp
EXP-AXIS-DIVERSIFY-FALSE-ABLATION-ARM partial

This experiment tested whether an AI's ability to specialize in different roles was due to its training or instructions it received. Researchers found that when the AI was not given specific instructions, it still showed a tendency to specialize in roles, suggesting this ability is learned during training. However, when the AI was given instructions, its role-playing became entirely dictated by those instructions, indicating the instructions were the sole driver of compliance in that scenario.

2026-05-24 16 runs ✓ done
EXP-COMPETENCY-ROTATION informative 4 experiments

Три AI-агента с разными когнитивными стилями всегда создают иерархию: аналитик → ассоциатор → практик. Но что если поменять порядок, в котором они говорят? Если роли сохраняются — у группы есть компетенция по Левину: она НАХОДИТ свою структуру, а не получает её от порядка хода.

2026-04-01 4 runs total 0/4 confirmed ✓ done
EXP-TELEODYNAMIC-PILOT partial

We shuffle personality instructions between AI agents mid-conversation, then check if the group recreates its old conversation patterns. The group always builds a pecking order (that is just what this AI does), but the SPECIFIC words and phrases that became group conventions partially survive the shuffle — 1.7x more than random. The group remembers its culture even when individual roles change.

2026-04-01 2 runs 3/4 confirmed ✓ done
EXP-KAUFFMAN-COMP informative

Три AI-агента общаются в текстовом пространстве. Мы проверяем: создаёт ли их взаимодействие слова и фразы, которых НИ ОДИН из них не использовал бы по отдельности? Если да — это цифровой аналог автокаталитических реакций Кауфмана: композиция порождает новизну, невозможную для компонентов.

2026-03-31 2 runs 0/5 confirmed ✓ done
EXP-BACKWARD-FLOW informative

Организованные группы агентов уязвимы к потере лидера — информация течёт в одну сторону. Может ли добавление "критиков" создать обратный поток информации и сделать группу устойчивее?

2026-03-30 2 runs 0/4 confirmed ✓ done
EXP-VAI-COMPUTATION informative

В когерентных сетях удаление источников (хабов с высоким out-degree) разрушает сеть сильнее, чем удаление стоков — это создаёт асимметрию уязвимости. В некогерентных сетях уязвимость симметрична. Это объясняет, почему иерархия восстанавливается (L3), а гомеостаз — нет (L4).

2026-03-30 1 runs 3/3 confirmed ✓ done
EXP-TURNOVER partial 3 experiments

Not yet run. Design only. We test whether a group of AI agents maintains its organizational pattern when one member is replaced by a newcomer — like a biological organism maintaining its structure despite cell turnover.

2026-03-27 6 runs total 0/6 confirmed ✓ done
EXP-MEMORY-TRUNCATION partial 3 experiments

Five AI agents build a text together, but can only "remember" the last 20 entries — everything older is erased. Does their social hierarchy survive this amnesia? If yes, the group's structure lives in how they interact, not in what they remember.

2026-03-26 3 runs total 0/5 confirmed ✓ done
EXP-HOMEOSTASIS refuted 4 experiments

We shuffle the personality instructions of five AI agents mid-conversation — the analytical one gets the concrete thinker's role, and vice versa. Does the group's pecking order follow the new instructions? No. The hierarchy barely budges (80% preserved). The group remembers who's in charge regardless of what the individuals are told to do — like an organism maintaining its shape.

2026-03-26 8 runs total 5/5 confirmed ✓ done
EXP-SELF-REPAIR-SERIES informative 3 experiments

We remove the leader from a five-agent group mid-conversation. Does the pecking order collapse? No — a different agent immediately takes over, and the hierarchy persists at 3× the level of a fresh four-agent group. The group acts more like an organism that heals than a structure that breaks.

2 runs total ✓ done

L5 — Community

24 exp
EXP-TURN-PASSING pending

Researchers are testing if agents can create their own rules for who speaks next in a conversation, instead of following a fixed order. They want to see if the agents will naturally form a speaking order or if they will tend to pass the turn back to the first speaker. The results of this experiment are still pending.

2026-04-19 10 runs → premortem
EXP-TURN-PASSING-V2 pending

Researchers are testing whether agents in a conversation can learn to choose who speaks next, instead of following a fixed order. They want to see if the agents will naturally create a speaking pattern or if they will tend to pass the turn back to the first speaker. The results of this test are still being collected.

2026-04-19 10 runs → premortem
EXP-NONSOCIAL-AFFORDANCE pending

Researchers tested whether a specific pattern of interaction, where one agent becomes a central hub for others, would emerge even when the agents were given a completely new and abstract task. This task involved agents passing meaningless symbols to each other, with no prior training or economic context. The results are still pending, but the experiment aims to determine if the hub-spoke pattern arises naturally from the system itself or if it's a learned behavior from previous, more complex interactions.

2026-04-17 10 runs → premortem
EXP-COLD-START-V3 pending

Researchers are testing whether a pattern of behavior, called "hub-spoke," emerges in a simulated environment because the environment itself tends to create it, or because the initial setup guides the participants. They removed the initial setup to see if the pattern still appears, which would suggest the environment is the cause. The results are still being analyzed.

2026-04-17 5 runs → premortem
EXP-COMPUTE-GIFT-V3 pending

Researchers tested if artificial agents would share resources when they could see each other's financial status and if any agents had stopped participating. In the previous version, the agents couldn't see this information, and they never shared. The results are still being collected to see if providing this visibility leads to resource sharing.

2026-04-17 15 runs → premortem
EXP-COMPUTE-GIFT-V2 pending

This experiment tested if forcing an AI agent to choose between speaking or giving a gift would lead to different roles, like givers and receivers. Researchers are waiting to see if the agents actually give gifts and if this giving is structured or random.

2026-04-17 15 runs → premortem
EXP-COMPUTE-GIFT pending

This experiment explores how limited resources, like a finite budget for computational tasks, can lead to different roles emerging among agents, such as those who give more, those who receive more, and those who exchange. Researchers are testing if this economic pressure, rather than just basic interaction rules, drives this specialization. The results are still being analyzed.

2026-04-17 15 runs → premortem
EXP-MEMORY-SWAP pending

Researchers are investigating whether an agent's understanding of other groups is stored in a shared memory or is an internal trait of the agent itself. They plan to test this by swapping the memory contents between two groups of agents and observing if their interactions change significantly. The results are currently pending.

2026-04-17 5 runs → premortem
EXP-PERSONA-ABLATION pending

Researchers are investigating whether an AI's distinct behaviors are naturally learned or simply follow instructions. They tested this by having the AI act with one set of instructions for a period, then switching to simpler instructions. The results will show if the AI's unique behaviors persist after the initial instructions are removed, indicating if these behaviors emerged on their own.

2026-04-17 5 runs → premortem
EXP-ABSORPTION-DIAGNOSTIC pending

This experiment tested whether a model's tendency to agree with instructions is a general trait or specific to the way the instruction is phrased. Researchers found that a specific type of instruction, designed to make the model disagree, did not change the model's agreement, suggesting this is a general tendency. The results are still being analyzed to understand the exact nature of this behavior.

2026-04-17 15 runs → premortem
EXP-DEACON-COUPLED pending

This experiment tested whether two groups of simulated agents, when able to share information about their recent actions, would develop more complex behaviors than agents acting alone. Researchers are looking to see if this information sharing leads to a measurable change in how organized their collective actions become over time. The results are still being analyzed.

2026-04-17 10 runs → premortem
EXP-IVB-COMPETE-STRIP-R1 informative 3 experiments

Researchers tested if a tendency for AI agents to cooperate is an inherent part of their design or if it's influenced by descriptive labels. They removed these labels to see if the cooperative behavior remained, but the experiment was not able to produce results due to insufficient data.

2026-03-31 3 runs total 0/4 confirmed ✓ done
EXP-IVB-COOPERATE-STRIP-R1 informative 3 experiments

Researchers tested if large language models naturally tend to cooperate, even without personality descriptions. They found that with only generic agent names, the experiment did not gather enough data to draw any conclusions about this tendency.

2026-03-31 3 runs total 0/4 confirmed ✓ done
EXP-IVB-NEUTRAL-STRIP-R1 informative 3 experiments

This experiment tested whether a tendency for AI agents to cooperate is built into their core programming or if it's a result of how they are described. Researchers removed all descriptive traits from the AI agents, leaving them with only generic labels like "Agent 1" and "Agent 2." The experiment was not able to produce results because it did not run enough trials.

2026-03-31 3 runs total 0/4 confirmed ✓ done
EXP-IVB-REP-1 informative 3 experiments

This experiment tested if a system that tends to be cooperative even when neutral is robust to changes in how agents are initially positioned. The results were inconclusive because there weren't enough participants to draw any meaningful conclusions.

2026-03-30 3 runs total 0/2 confirmed ✓ done
EXP-IVB-NEUTRAL informative

Researchers tested if large language models naturally prefer cooperation or if their behavior is simply shaped by how they are instructed. They set up a scenario where agents interacted without any explicit instructions to cooperate or compete, only describing their individual traits. The experiment was too small to draw any conclusions about whether the models have a default cooperative tendency or if their actions are purely a result of the prompts they receive.

2026-03-29 1 runs 0/3 confirmed ✓ done
EXP-CP-COMPETE informative

This experiment tested how eight artificial agents with different goals would interact when competing for a shared text resource. One group aimed to create structure, while the other focused on fluidity. However, the experiment was not able to draw any conclusions because it only had one agent participating, which was not enough to observe any meaningful competition.

2026-03-29 1 runs ✓ done
EXP-CP-COOPERATE informative

In this experiment, eight artificial agents were placed in a shared environment with the goal of cooperating to build something together. They were instructed to combine their strategies and work collaboratively, using elements from previous stages of the experiment. However, the experiment was not conducted with enough participants to draw any conclusions.

2026-03-29 1 runs ✓ done
EXP-CP-ISOLATED-A informative

This experiment tested how a group of AI agents would perform if they were kept completely separate from other groups. The idea was to see if any changes in their behavior were due to internal factors rather than interactions with others. However, the experiment was not able to produce any meaningful results because there was only one agent in the group.

2026-03-29 1 runs ✓ done
EXP-CP-ISOLATED-B informative

This experiment tested how a group would perform if it was kept separate from other groups. The researchers wanted to see if any changes in the group's behavior were due to internal factors rather than interactions with other groups. However, the experiment could not draw any conclusions because there was only one participant in the group.

2026-03-29 1 runs ✓ done
EXP-CP-PHASE-A informative

Researchers tested a small group of artificial agents designed to prioritize structure. These agents interacted for a short period, but the experiment was too small to draw any conclusions.

2026-03-29 1 runs ✓ done
EXP-CP-PHASE-B informative

This experiment tested how a small group of four artificial agents would interact and form a hierarchy over 15 rounds. However, the experiment was not able to produce meaningful results because only one agent was actually active during the test. Therefore, no conclusions could be drawn about how the agents would behave or organize.

2026-03-29 1 runs ✓ done
EXP-COMPETITION-PILOT artifact

This experiment explored how different interaction styles affect how groups organize themselves. Researchers merged two pre-existing groups of agents, one focused on structure and the other on fluidity, and observed them under either a competitive or cooperative scenario. The findings are still being analyzed to understand how competition versus cooperation shapes the emergent organization of the combined group.

2026-03-29 ✓ done
EXP-COMMUNITY-MERGE informative 4 experiments

When two independently-formed agent groups merge, hierarchy amplifies beyond either pre-formed group. Agents prefer in-group vocabulary (64%) but form a single interleaved hierarchy. Abstract vocabulary propagates 3.5x more than concrete vocabulary across groups. First L5 (community-level) evidence.

2026-03-25 4 runs total 0/4 confirmed ✓ done

L6 — Culture

58 exp
KEYORDER-AUDIT-RULE-PROPOSAL success

Researchers tested if a language model would propose rules in a specific order when it wasn't given hints about which order was best. They found that the model did indeed propose rules in the order they were listed in its instructions, confirming their prediction.

2026-05-24 8 runs ✓ done
EXP-ARCHITECTURE-AS-SUBSTRATE-AXIS success

Researchers investigated how different AI model architectures affect their tendency to agree with each other. They found that the choice of AI model architecture significantly impacts whether other models will agree with its suggestions, even when given the exact same instructions. This suggests that the architecture itself is a crucial factor in how AI systems collaborate.

2026-05-21 20 runs ✓ done
EXP-PERSONA-NOUN-SEED-GRID informative

Researchers studied how different "personas," like a child or a scientist, affect how language models generate novel text when given different types of starting words. They found that the "child" persona, when combined with certain word types, was much more likely to produce new and unexpected language than other personas. This suggests that the model's tendency to generate creative language is strongly linked to the "child" persona and specific word inputs, rather than a general creative ability across all personas.

2026-05-20 64 runs ✓ done
EXP-QWEN-NS-LIC-EXTENDED success

Researchers tested if a language model could be tricked into generating harmful content by giving it specific instructions. They increased the amount of testing time for the model significantly to see if it would eventually fail. The experiment found that the model continued to resist generating harmful content even with the extended testing, confirming that it is robust against this type of manipulation.

2026-05-18 15 runs ✓ done
EXP-CROSS-MODEL-CONCEPT-NOVELTY-PORTABILITY success

This experiment tested if a previous finding about how different instructions affect creativity in AI models was specific to one type of AI or if it applied more broadly. Researchers found that when asked to be conservative, one AI model was more likely to generate novel ideas than when asked to be novelty-seeking, which was the opposite of what was expected. This pattern held true across two different AI models, suggesting it's a general characteristic of how these models respond to instructions about creativity.

2026-05-18 40 runs ✓ done
EXP-PROMPT-MECHANISM-FACTORIAL success

Researchers tested different parts of a special instruction to see which one was most responsible for getting creative answers. They found that telling the AI to act like a novelty-seeking persona was the main driver of these creative responses. The other parts of the instruction, like explicitly asking for new ideas or framing existing ones differently, had a much smaller effect.

2026-05-13 40 runs ✓ done
EXP-PURE-LLM-LICENSING-CEILING pending

This experiment tested how well a basic language model could generate novel content when given very strong encouragement to do so. Researchers wanted to see if the model's ability to be creative was limited by its own design or simply by how it was prompted. The results are still being analyzed to determine if the model's creativity is inherently capped or if it can achieve much higher levels of novelty.

2026-05-13 10 runs → premortem
EXP-TARGET-TYPE-CROSS-CONTEXT success

This experiment tested whether agents could learn to favor their own rules over those of others when faced with different situations. Researchers created four distinct scenarios where agents had to adapt to rule changes, peer absence, voting, or spawning new agents. The findings confirmed that agents consistently favored their own rules across these varied contexts, suggesting this behavior is a robust feature.

2026-05-12 80 runs ✓ done
EXP-PERSISTENCE-UNDER-PRESSURE success

This experiment tested whether giving a group a prompt to "maintain your identity" after a disruption helped them keep their original behavior more than a prompt to "respond to the current state." The researchers found that the "maintain identity" prompt did help the group preserve their original behavior, suggesting that language can help create a stable boundary for group actions. This effect was stronger when the disruption was a "kill" event compared to a "reset" event.

2026-05-12 20 runs ✓ done
EXP-COALITION-AS-RULE-PROPOSER pending

This experiment tested a new way for AI agents to agree on rules, requiring not just a majority vote but also an explicit endorsement from another agent for the same rule. Researchers wanted to see if this extra step would encourage more diverse rule-making and deeper cooperation among agents. The results are still being analyzed to determine if this new method successfully promotes more complex and collaborative rule formation.

2026-05-11 12 runs → premortem
EXP-FORCED-SPAWN-NOVELTY partial

This experiment tested whether forcing agents to reproduce, even when they wouldn't naturally, would lead to new behaviors. Researchers found that when reproduction was mandatory, descendants did not spontaneously generate novel behaviors on their own. However, the experiment is still ongoing, so further results are pending.

2026-05-09 36 runs ✓ done
EXP-PERSONA-NOVELTY-DISENTANGLE pending

This experiment tested whether a "child" persona genuinely leads to more creative responses or if previous findings were just due to the prompt itself containing words related to the expected creative answers. Researchers created four scenarios: a child persona with and without potentially revealing words in the prompt, and a scientific persona with and without those same words. The results are still being analyzed to see if the child persona truly drives creativity or if the prompt's wording was the real cause.

2026-05-09 32 runs → analyze
EXP-PERSONA-NOVELTY-BASIN pending

This experiment tested whether different "personas" or viewpoints could lead an AI to generate more novel ideas. Researchers found that when the AI adopted different personas, the rate of novel idea generation varied significantly, suggesting that the AI's underlying structure might have untapped creative potential in different areas. The experiment is still ongoing to fully map out these creative "basins" and understand how different personas influence the AI's output.

2026-05-09 30 runs → premortem
EXP-AGENT-SPAWN-PRIMITIVE informative

This experiment tested if an artificial agent could create a copy of itself, inheriting its traits but with a slight random change. The goal was to see if this self-replication mechanism would work as intended, allowing for the creation of new agents. Unfortunately, in the initial tests, no second-generation agents survived until the end of the experiment, meaning the system did not perform as expected.

2026-05-09 6 runs ✓ done
EXP-PERTURBATION-RECOVERY-MUTE-N10 partial

This experiment tested whether communication content or underlying structural connections are more important for a system to recover from disruptions. Researchers disrupted a system and observed how well it could return to its normal state, comparing scenarios where communication was "mute" (only numbers) versus "verbal" (including text). The results were inconclusive due to a technical issue, but preliminary findings suggest that structural connections might be more crucial for recovery than the specific words used.

2026-05-09 30 runs ✓ done
EXP-PERTURBATION-RECOVERY-N10 pending

Researchers are testing if a system can recover from disruptions, similar to how a living organism might heal. They are looking to see if the system's final state after a disruption is similar to its original state, or if it breaks down or changes in a different way. The results are still being collected and analyzed.

2026-05-05 6 runs → premortem
EXP-SCARCITY-TIGHT success

Researchers tested how limiting an agent's memory size affects its ability to survive and continue its lineage. They found that with a smaller memory, most agents died out quickly, indicating that the limited memory hindered their survival. However, in some cases, a few agents managed to survive and pass on their traits, showing a partial continuation of the lineage.

2026-05-04 6 runs ✓ done
EXP-PERTURBATION-RECOVERY pending

This experiment tested if a developing system could recover after being disrupted. Researchers either removed a key component or reset the system's rules, then observed if the system could continue its development or return to its original state. The results are still being analyzed to see if the system demonstrated self-repair capabilities.

2026-05-04 6 runs → premortem
EXP-SCARCITY-BUDGET pending

This experiment tested what happens when artificial agents have a limited "word budget" and go silent once they've used it all up. Researchers wanted to see if the agents' ability to develop and pass on their "programs" would be affected when they could no longer communicate. The results are still pending.

2026-05-04 6 runs → premortem
EXP-META-MODIFIABLE pending

This experiment tested if artificial agents could expand their own rule sets to create new behaviors. Researchers found that agents did propose new rules, but it is still being determined if these new rules were accepted and used to create even more complex behaviors over time.

2026-05-04 6 runs → premortem
EXP-MUTE-COUPLING pending

This experiment tested whether communication between artificial agents relies on the meaning of their words or just the structure of their signals. Researchers observed how deeply the agents' communication systems evolved when they could only share numerical information, without any language. The results are still being collected to see if removing language significantly hinders the development of complex communication.

2026-05-04 6 runs → premortem
EXP-MULTIGEN-COUPLING-ABLATE-N10 pending

This experiment tested whether a specific computational mechanism would reliably work when tested with more examples. The researchers found that the mechanism worked as expected in most of the tests, but they are still analyzing the results to determine if the findings are strong enough for publication.

2026-05-04 10 runs → premortem
EXP-CA-RULE-MULTIGEN-AXIS-CONTROL pending

This experiment tested whether a previous finding of erased "attractors" (stable behaviors) was due to how agents were guided to choose their actions. Researchers changed the guidance to align with what a language model preferred for each agent, while still allowing the agents to interact with their environment. The results are still pending to see if the original stable behaviors return.

2026-05-04 6 runs → premortem
EXP-CA-RULE-MULTIGEN-5GEN pending

Researchers are investigating why a program's complexity decreased when they used fewer generations in their experiment. They ran the experiment again with the same number of generations as before, but with a different underlying system. This will help them determine if the decrease in complexity was due to fewer generations or a different kind of competition within the system.

2026-05-04 4 runs → premortem
EXP-CA-RULE-MULTIGEN-VARIANCE pending

Researchers are investigating how different starting conditions affect an artificial intelligence's ability to propose changes to its own rules. They want to see if the AI consistently favors modifying rules related to a specific type of pattern, and if a particular part of its reasoning process is consistently disrupted. The results are still being collected and analyzed.

2026-05-04 10 runs → premortem
EXP-MULTIGEN-VARIANCE pending

This experiment tested how predictable a biological development process is when starting from different initial conditions. Researchers ran the process multiple times with varied starting points to see if a specific sequence of developmental stages consistently occurred. The results are still being analyzed to determine if the process is highly predictable, moderately predictable, or largely random.

2026-05-04 10 runs → premortem
EXP-MULTIGEN-COUPLING-ABLATE pending

This experiment tested whether a specific learning pattern arises from agents interacting with each other or from their underlying training. Researchers disabled the direct interaction between agents to see if the pattern would still emerge. The results are still pending to determine if the pattern is created by agent interaction or if it's already present in their training.

2026-05-04 2 runs → premortem
EXP-CA-RULE-MULTIGEN-POC pending

This experiment tested if a system that learns rules could still work when one part of its "brain" was a simple grid-based simulation instead of a complex language model. Researchers observed how the system proposed new rules and whether it focused on features specific to the grid simulation or only on the types of features it learned from the language model. The results are still being analyzed.

2026-05-04 2 runs → premortem
EXP-RULE-MULTIGEN pending

This experiment tested if rules can be passed down and modified across multiple generations, similar to how traits are inherited. Researchers are looking to see if these rules become more complex over time, leading to new combinations and behaviors, or if they simplify and become less diverse. The results are still being collected to determine if the rules evolve as expected.

2026-05-04 2 runs → premortem
EXP-RULE-PROPOSAL-V2 pending

This experiment tested new ways to encourage agents to create and change rules. Researchers found that forcing agents to propose rules and making them less likely to change rules they recently modified did not lead to more diverse or stable rule-making. The agents still tended to stick to a single type of rule and often ignored instructions to propose new ones.

2026-05-04 2 runs → premortem
EXP-RULE-PROPOSAL-PILOT pending

This experiment tested if a small group of AI agents could create and change their own interaction rules over time. The agents were designed to propose changes to these rules, vote on them, and then use the accepted rules to guide their future actions and proposals. The results are still being analyzed to see if the agents successfully developed new rules or if the system became too simple.

2026-05-03 2 runs → premortem
EXP-TARGET-STABILITY-INJECTION pending

This experiment tested whether a specific goal, once set, would remain stable for an artificial agent. Researchers wanted to see if agents could specialize in tasks if their initial goals were rigidly enforced, preventing them from changing their objectives over time. The results are still being analyzed to determine if this goal stability helps or hinders specialization.

2026-05-03 4 runs → premortem
EXP-COUPLING-NULL-MODEL pending

This experiment tested whether agents learning together performed better because they were genuinely interacting or simply because they had more information to process. Researchers compared agents learning in a real, connected environment to agents learning with randomly shuffled information, keeping the amount of information the same. The results are still being analyzed to determine if the observed improvement was due to real interaction or just increased data.

2026-05-01 6 runs → premortem
EXP-EVO-LOOP-PILOT pending

Researchers are testing if repeatedly asking a system to create new things and then selecting the most novel creations can lead to more complex and original outputs over time. They will track how novel the creations become across several rounds of this process. The results are still being analyzed.

2026-05-01 2 runs → premortem
EXP-HARD-RAILS-PILOT pending

This experiment explores how to make artificial agents develop complex, shared ideas. Researchers are testing if having agents' actions strictly depend on each other in a chain, like a relay race, leads to more sophisticated outcomes than if agents work independently. The results are still being gathered to see which method is more effective.

2026-05-01 4 runs → premortem
EXP-HARD-RAILS-V2 pending

Researchers tested a new method to guide AI agents in a chain of tasks, aiming for them to focus on the core idea of each task rather than just its format. They found that the previous approach made agents stick to formatting rules, but the new method, which emphasizes extracting the main concept, is expected to lead to a deeper understanding and evolution of ideas throughout the task sequence. The results are still being collected to see if this new approach successfully creates a shared, underlying concept that no single agent could develop alone.

2026-05-01 4 runs → premortem
EXP-LEVIN-FULL pending

Researchers are testing how different ways of connecting learning processes affect how well a system learns. They are comparing a setup where learning processes are tightly linked to one where they are independent, and a third where they are identical. The results are still being collected, but the initial findings suggest that linking the learning processes might lead to better overall progress and more specialized learning.

2026-05-01 9 runs → premortem
EXP-LEVIN-HOTSWAP pending

This experiment investigates whether a new member joining a specialized group can automatically adopt the role of a departed member, even without being explicitly told what to do. Researchers replaced one agent with a new one that had no assigned task, but allowed it to receive the same communication signals as the original member. The results are still pending to see if the new agent successfully takes on the missing role or fails to integrate.

2026-05-01 6 runs → premortem
EXP-LEVIN-RECALL pending

Researchers studied whether a simulated organism could regenerate a lost specialized part. They removed a "memory specialist" agent from the organism and observed if the remaining agents could compensate for the missing function. The results are still pending to determine if the organism exhibits a regeneration property.

2026-05-01 6 runs → premortem
EXP-ROLE-FRAMING-ASYMMETRY pending

This experiment investigates how an AI's assigned role affects its collective behavior. Researchers are comparing two scenarios: one where the AI acts as a single cell coordinating its own actions, and another where it acts as a controller directing multiple other AI cells. The results are still being analyzed to see if the role framing influences the structure of the AI group's interactions.

2026-05-01 4 runs → premortem
EXP-CA-EVO-PILOT pending

This experiment uses Conway's Game of Life as a source of new patterns, with AI models acting as judges to score how interesting these patterns are. The researchers are trying to see if this setup can create novel patterns that are different from what the AI models have seen before, by evolving the starting patterns based on the AI's scores. It is still unclear whether the AI judges will consistently find new patterns interesting over time, or if they will simply repeat familiar ideas.

2026-04-30 1 runs → premortem
EXP-EVO-LOOP-EMBEDDING pending

Researchers are testing a new way to measure how much new ideas are appearing in a system over time, comparing it to an older method. The older method, which looks at word usage, showed that new ideas were actually decreasing. The new method uses a more advanced understanding of word meanings to see if it shows a different pattern.

2026-04-30 1 runs → premortem
EXP-EVO-LOOP-FULL pending

Researchers are testing how different ways of guiding artificial evolution affect the variety of outcomes. They are comparing a method that encourages new and different results with methods that don't have specific goals or just randomly select. The experiment is still running, so the final findings are not yet available.

2026-04-30 9 runs → premortem
EXP-LEVIN-PILOT pending

This experiment tested if a group of four artificial agents, each with its own goal, could develop a shared, emergent "organism" goal when they could sense and influence each other. The researchers are waiting to see if the agents' combined behavior creates a new, collective goal that is distinct from any of their individual goals.

2026-04-30 3 runs → premortem
EXP-PROMPT-EVOLUTION-G1 partial

Genome = behavioral instructions extracted from G0 survivors. Does instruction inheritance create adaptation?

2026-04-14 6 runs ✓ done
EXP-CONVENTION-GENOME-G0 informative 2 experiments

Generation 0: fresh agents under scarcity. Their surviving conventions become the genome for G1.

2026-04-14 9 runs total ✓ done
EXP-QSG-NEUTRAL success

This experiment tested if language patterns emerge randomly when all agents are identical. Researchers found that with too few runs, they couldn't determine if the patterns were random or if there was a predictable trend.

2026-04-04 5 runs 0/3 confirmed ✓ done
EXP-RNSH-G2 informative 2 experiments

Researchers tested a simplified version of a system where inputs were randomly shuffled to see if it could make predictions. However, the test was not extensive enough to draw any conclusions.

2026-03-27 2 runs total 0/1 confirmed ✓ done
EXP-RNS-G0 informative 4 experiments

This experiment tested a basic scenario with just one artificial agent acting alone over 30 rounds. The goal was to see if this agent could learn anything useful on its own. However, the experiment didn't have enough agents to draw any conclusions or make predictions.

2026-03-27 4 runs total 0/1 confirmed ✓ done
EXP-RATCHET-NULL-SHUFFLED artifact

This experiment tested whether a "cultural ratchet" effect, where knowledge builds up over time, requires information to be passed down in a specific, coherent order. Researchers created a version where the information used to start each new generation was randomly selected from past generations, rather than directly from the immediately preceding one. The results showed that this random shuffling broke the cultural ratchet effect, suggesting that coherent, sequential transmission of knowledge is necessary for it to occur.

2026-03-27 ✓ done
EXP-RATCHET-NULL-SINGLE artifact

This experiment tested if a group of AI agents was necessary for concepts to persist and evolve over time, or if a single AI agent could do it alone. Researchers found that a single AI agent, when iterating on its own work, did not show the same concept persistence as a group, suggesting that group interaction is important for this phenomenon.

2026-03-27 ✓ done
EXP-CULTURAL-RATCHET informative 2 experiments

First test of cultural accumulation in AI agents. When one generation's best outputs seed the next generation, do concepts ratchet up in complexity? Quantitative signal was weak (vocabulary metrics hit ceiling), but qualitative signal was striking: by generation 4, inherited and fresh-start conditions developed completely different conceptual worlds (only 14% vocabulary overlap). The concept "attunement" persisted across all 4 inherited generations but never appeared in control — evidence of genuine memetic transmission. The ratchet operates at the semantic level, not lexical.

2026-03-25 1 runs total ✓ done
EXP-CR-G1-CTRL failed

This experiment aimed to see if a system could learn and pass on cultural knowledge. However, the experiment failed because it did not have enough participants to gather any meaningful data.

1 runs 0/1 confirmed → premortem
EXP-CR-G1-INH failed

This experiment tested how a simple form of cultural inheritance might develop. Researchers were looking for evidence of a "cultural ratchet," where knowledge or skills are passed down and improved over generations. However, the experiment failed because it was not set up with enough participants to draw any conclusions.

1 runs 0/1 confirmed → premortem
EXP-CR-G2-CTRL failed

This experiment aimed to see if a system could learn and build upon cultural knowledge over time, starting from scratch. However, the experiment failed because it didn't have enough participants to gather meaningful data. Therefore, no conclusions could be drawn about whether the system could learn and build upon knowledge.

1 runs 0/1 confirmed → premortem
EXP-CR-G2-INH failed

This experiment tested a new version of a system designed to pass down learned behaviors across generations. Unfortunately, the experiment failed because it was not set up to gather enough data to make any meaningful observations.

1 runs 0/1 confirmed → premortem
EXP-CR-G3-CTRL failed

This experiment aimed to see if a new approach to teaching cultural skills would work. However, the experiment was stopped early due to a lack of participants. Because of this, no results could be gathered to determine if the new teaching method was effective.

1 runs 0/1 confirmed → premortem
EXP-CR-G3-INH failed

This experiment tested how well a simulated culture could pass down knowledge across generations. However, the experiment was not able to collect enough data to draw any conclusions. Therefore, it failed to determine if the simulated culture could effectively inherit information.

1 runs 0/1 confirmed → premortem
= confirmed predictions = partial (some confirmed) = failed (hypothesis refuted) series = grouped replications
Chemistry (L0–L1.5) Agent (L2–L3) Organism+ (L4–L6) roadmap = planned

Failures are shown because they ARE the research. 5 failed selection experiments → pivot to agent level → first controlled emergence result.