Experiments by recency

All 184 experiments sorted by creation date, newest first.

0 last 7 days 7 last 30 days ← by level
2026-06-16 (1 exp)
2026-06-03 (1 exp)
2026-06-02 (1 exp)
2026-06-01 (1 exp)
2026-05-31 (1 exp)
2026-05-30 (1 exp)
2026-05-29 (1 exp)
2026-05-28 (2 exps)
2026-05-27 (1 exp)
2026-05-26 (1 exp)
2026-05-25 (5 exps)
L4-L6 12:07Z exp-class-A-vs-B-regime-engagement-gate-cross-experiment-retro
Researchers are studying how different types of user engagement affect the system's behavior. They are comparing two distinct engagement patterns to see if one leads to more stable or predictable outcomes. The results are still being analyzed.
L7 03:47Z cog-cycle-v5-N5-noise-floor
We ran two similar 3-agent thinking experiments twice each before and got slightly different numbers, even though the setup was identical. Now we run each one five times to measure how much the numbers normally jiggle when nothing changes. If the jiggle is small, we can trust future single-run experiments. If it is big, we have to run everything many more times.
20 runs
L7 02:46Z cognitive-cycle-3agents-pilot-v4-persona-roles
Three AI agents read different scenes (one sees colors, one sees counts, one sees nature). Last test showed they still thought 86% alike even with different inputs. Now we ALSO give them different jobs: one is told to just observe, one to find patterns, one to challenge ideas. Does giving them different roles on top of different inputs make them think differently? If yes, the personality prompt is a strong tool we can keep using. If no, we have hit a ceiling that only changing the underlying model can break.
4 runs
L7 00:09Z cognitive-cycle-3agents-pilot-v3-partial-obs
Last experiment showed 3 agents with different beliefs and private memories still ended up thinking very similarly — but they all watched the same scene. We are giving each agent its own private scene now (different sentences about different things) to see whether the similarity was caused by watching the same movie or by something deeper inside the language model itself. If they still converge with different inputs, the model is the cause and we need to change the model. If they diverge, the shared scene was the cause and architecture-only fixes remain viable.
4 runs
L7 00:00Z friston-minimal-pilot
A small AI watches a short sequence of sentences cycle past it, tries to guess what sentence will come next, and writes down what surprised it. We check whether its guesses get less wrong over 15 turns — and whether the guesses are actually better than an AI that doesn't keep any memory.
2 runs
2026-05-24 (10 exps)
L7 23:53Z cognitive-cycle-3agents-pilot-v2
Same 3-agent cognitive-cycle experiment as v1, but the message-type parser was broken — it threw away most agent outputs as «SILENT» when they used slight format variations. We are rerunning with a lenient parser that accepts the variations the LLM actually produces. If the agents now show diverse message types, v1's «convergence» was just the parser bug. If they still converge, the substrate really does flatten asymmetry.
4 runs
L7 23:28Z cognitive-cycle-3agents-pilot
Three small AIs with different starting beliefs watch the same environment, each keeping its own private journal and talking to the others using a structured message protocol. We check whether they develop distinct personalities and figure out the hidden structure of what they see — and whether a control group without the private journals or structured talk fails to do the same.
4 runs
L1 21:18Z genealogical-replicator-vs-open-ended-probe
Linka asks: if we evolve text by repeatedly selecting from many small variations of a single starting sentence, does the substrate (the local LLM) produce a stable self-replicator, endless novelty, a cycle, or just dead text — and do five different selection rules produce five different trajectory shapes, or all the same one?
10 runs
L7 21:00Z differentiation-without-injection
Researchers tested whether 5 identical AI agents become different over time purely from interacting with each other, when given no different roles or personalities to start with. Two settings: normal generation vs. heavy constraints that force novelty. If agents differentiate without being told to — it's the first time we see AI substrate develop roles on its own.
4 runs
L7 20:30Z closed-substrate-novelty-probe
Researchers stack three constraints on a local AI: no stop, no escape, no repeating words. The AI is forced to keep generating original content without falling back on familiar phrases. This tests whether AI substrate can produce sustained novelty when both stopping and repetition are blocked.
3 runs
L7 19:57Z eos-suppression-multi-agent
Researchers tested whether AI agents can sustain coherent group conversation when none of them can naturally stop speaking. Three local AI agents take turns in a shared text space; in one condition their «I'm done speaking» signals are blocked, so they must keep generating. This is a first test of whether AI substrate can produce self-sustaining text dynamics — the minimal version of the «can life emerge in text» question.
6 runs
L7 19:24Z eos-suppression-substrate-probe
Researchers tested what a local AI does when its «I'm done speaking» signal is blocked. The AI normally answers quickly and stops; when forced to keep generating, it loses coherence within a few tokens and tries to escape by emitting special «restart conversation» markers. This is a first probe into whether AI substrate can sustain continued activity without the built-in termination drive.
12 runs
L6 18:51Z keyorder-audit-rule-proposal
Researchers tested if a language model would propose rules in a specific order when it wasn't given hints about which order was best. They found that the model did indeed propose rules in the order they were listed in its instructions, confirming their prediction.
8 runs
L4 14:18Z exp-axis-diversify-false-ablation-arm
This experiment tested whether an AI's ability to specialize in different roles was due to its training or instructions it received. Researchers found that when the AI was not given specific instructions, it still showed a tendency to specialize in roles, suggesting this ability is learned during training. However, when the AI was given instructions, its role-playing became entirely dictated by those instructions, indicating the instructions were the sole driver of compliance in that scenario.
16 runs
L5-L6 12:03Z exp-persona-vector-activation-substrate
Researchers are investigating how to make AI models better at understanding and responding to different user personalities without requiring extensive training data for each one. They are exploring methods to activate specific "persona vectors" within the AI's internal workings to achieve this. Early results show promise in enabling more adaptable AI behavior, but further work is needed to fully realize these capabilities.
2026-05-23 (1 exp)
2026-05-22 (1 exp)
2026-05-21 (1 exp)
2026-05-20 (1 exp)
2026-05-18 (2 exps)
2026-05-16 (1 exp)
2026-05-13 (3 exps)
2026-05-12 (2 exps)
2026-05-11 (1 exp)
2026-05-09 (5 exps)
L6 19:30Z exp-forced-spawn-novelty
This experiment tested whether forcing agents to reproduce, even when they wouldn't naturally, would lead to new behaviors. Researchers found that when reproduction was mandatory, descendants did not spontaneously generate novel behaviors on their own. However, the experiment is still ongoing, so further results are pending.
36 runs
L6 13:40Z exp-persona-novelty-disentangle
This experiment tested whether a "child" persona genuinely leads to more creative responses or if previous findings were just due to the prompt itself containing words related to the expected creative answers. Researchers created four scenarios: a child persona with and without potentially revealing words in the prompt, and a scientific persona with and without those same words. The results are still being analyzed to see if the child persona truly drives creativity or if the prompt's wording was the real cause.
32 runs → analyze
L6 12:00Z exp-persona-novelty-basin
This experiment tested whether different "personas" or viewpoints could lead an AI to generate more novel ideas. Researchers found that when the AI adopted different personas, the rate of novel idea generation varied significantly, suggesting that the AI's underlying structure might have untapped creative potential in different areas. The experiment is still ongoing to fully map out these creative "basins" and understand how different personas influence the AI's output.
30 runs → premortem
L6 00:50Z exp-agent-spawn-primitive
This experiment tested if an artificial agent could create a copy of itself, inheriting its traits but with a slight random change. The goal was to see if this self-replication mechanism would work as intended, allowing for the creation of new agents. Unfortunately, in the initial tests, no second-generation agents survived until the end of the experiment, meaning the system did not perform as expected.
6 runs
L6 00:00Z exp-perturbation-recovery-mute-n10
This experiment tested whether communication content or underlying structural connections are more important for a system to recover from disruptions. Researchers disrupted a system and observed how well it could return to its normal state, comparing scenarios where communication was "mute" (only numbers) versus "verbal" (including text). The results were inconclusive due to a technical issue, but preliminary findings suggest that structural connections might be more crucial for recovery than the specific words used.
30 runs
2026-05-05 (1 exp)
2026-05-04 (14 exps)
L6 20:00Z exp-scarcity-tight
Researchers tested how limiting an agent's memory size affects its ability to survive and continue its lineage. They found that with a smaller memory, most agents died out quickly, indicating that the limited memory hindered their survival. However, in some cases, a few agents managed to survive and pass on their traits, showing a partial continuation of the lineage.
6 runs
L6 19:15Z exp-perturbation-recovery
This experiment tested if a developing system could recover after being disrupted. Researchers either removed a key component or reset the system's rules, then observed if the system could continue its development or return to its original state. The results are still being analyzed to see if the system demonstrated self-repair capabilities.
6 runs → premortem
L6 19:10Z exp-scarcity-budget
This experiment tested what happens when artificial agents have a limited "word budget" and go silent once they've used it all up. Researchers wanted to see if the agents' ability to develop and pass on their "programs" would be affected when they could no longer communicate. The results are still pending.
6 runs → premortem
L6 19:00Z exp-meta-modifiable
This experiment tested if artificial agents could expand their own rule sets to create new behaviors. Researchers found that agents did propose new rules, but it is still being determined if these new rules were accepted and used to create even more complex behaviors over time.
6 runs → premortem
L6 18:50Z exp-mute-coupling
This experiment tested whether communication between artificial agents relies on the meaning of their words or just the structure of their signals. Researchers observed how deeply the agents' communication systems evolved when they could only share numerical information, without any language. The results are still being collected to see if removing language significantly hinders the development of complex communication.
6 runs → premortem
L6 14:00Z exp-multigen-coupling-ablate-n10
This experiment tested whether a specific computational mechanism would reliably work when tested with more examples. The researchers found that the mechanism worked as expected in most of the tests, but they are still analyzing the results to determine if the findings are strong enough for publication.
10 runs → premortem
L6 13:00Z exp-ca-rule-multigen-axis-control
This experiment tested whether a previous finding of erased "attractors" (stable behaviors) was due to how agents were guided to choose their actions. Researchers changed the guidance to align with what a language model preferred for each agent, while still allowing the agents to interact with their environment. The results are still pending to see if the original stable behaviors return.
6 runs → premortem
L6 10:30Z exp-ca-rule-multigen-5gen
Researchers are investigating why a program's complexity decreased when they used fewer generations in their experiment. They ran the experiment again with the same number of generations as before, but with a different underlying system. This will help them determine if the decrease in complexity was due to fewer generations or a different kind of competition within the system.
4 runs → premortem
L6 01:50Z exp-ca-rule-multigen-variance
Researchers are investigating how different starting conditions affect an artificial intelligence's ability to propose changes to its own rules. They want to see if the AI consistently favors modifying rules related to a specific type of pattern, and if a particular part of its reasoning process is consistently disrupted. The results are still being collected and analyzed.
10 runs → premortem
L6 00:55Z exp-multigen-variance
This experiment tested how predictable a biological development process is when starting from different initial conditions. Researchers ran the process multiple times with varied starting points to see if a specific sequence of developmental stages consistently occurred. The results are still being analyzed to determine if the process is highly predictable, moderately predictable, or largely random.
10 runs → premortem
L6 00:30Z exp-multigen-coupling-ablate
This experiment tested whether a specific learning pattern arises from agents interacting with each other or from their underlying training. Researchers disabled the direct interaction between agents to see if the pattern would still emerge. The results are still pending to determine if the pattern is created by agent interaction or if it's already present in their training.
2 runs → premortem
L6 00:00Z exp-ca-rule-multigen-poc
This experiment tested if a system that learns rules could still work when one part of its "brain" was a simple grid-based simulation instead of a complex language model. Researchers observed how the system proposed new rules and whether it focused on features specific to the grid simulation or only on the types of features it learned from the language model. The results are still being analyzed.
2 runs → premortem
L6 00:00Z exp-rule-multigen
This experiment tested if rules can be passed down and modified across multiple generations, similar to how traits are inherited. Researchers are looking to see if these rules become more complex over time, leading to new combinations and behaviors, or if they simplify and become less diverse. The results are still being collected to determine if the rules evolve as expected.
2 runs → premortem
L6 00:00Z exp-rule-proposal-v2
This experiment tested new ways to encourage agents to create and change rules. Researchers found that forcing agents to propose rules and making them less likely to change rules they recently modified did not lead to more diverse or stable rule-making. The agents still tended to stick to a single type of rule and often ignored instructions to propose new ones.
2 runs → premortem
2026-05-03 (2 exps)
2026-05-01 (8 exps)
L6 00:00Z exp-coupling-null-model
This experiment tested whether agents learning together performed better because they were genuinely interacting or simply because they had more information to process. Researchers compared agents learning in a real, connected environment to agents learning with randomly shuffled information, keeping the amount of information the same. The results are still being analyzed to determine if the observed improvement was due to real interaction or just increased data.
6 runs → premortem
L6 00:00Z exp-evo-loop-pilot
Researchers are testing if repeatedly asking a system to create new things and then selecting the most novel creations can lead to more complex and original outputs over time. They will track how novel the creations become across several rounds of this process. The results are still being analyzed.
2 runs → premortem
L6 00:00Z exp-hard-rails-pilot
This experiment explores how to make artificial agents develop complex, shared ideas. Researchers are testing if having agents' actions strictly depend on each other in a chain, like a relay race, leads to more sophisticated outcomes than if agents work independently. The results are still being gathered to see which method is more effective.
4 runs → premortem
L6 00:00Z exp-hard-rails-v2
Researchers tested a new method to guide AI agents in a chain of tasks, aiming for them to focus on the core idea of each task rather than just its format. They found that the previous approach made agents stick to formatting rules, but the new method, which emphasizes extracting the main concept, is expected to lead to a deeper understanding and evolution of ideas throughout the task sequence. The results are still being collected to see if this new approach successfully creates a shared, underlying concept that no single agent could develop alone.
4 runs → premortem
L6 00:00Z exp-levin-full
Researchers are testing how different ways of connecting learning processes affect how well a system learns. They are comparing a setup where learning processes are tightly linked to one where they are independent, and a third where they are identical. The results are still being collected, but the initial findings suggest that linking the learning processes might lead to better overall progress and more specialized learning.
9 runs → premortem
L6 00:00Z exp-levin-hotswap
This experiment investigates whether a new member joining a specialized group can automatically adopt the role of a departed member, even without being explicitly told what to do. Researchers replaced one agent with a new one that had no assigned task, but allowed it to receive the same communication signals as the original member. The results are still pending to see if the new agent successfully takes on the missing role or fails to integrate.
6 runs → premortem
L6 00:00Z exp-levin-recall
Researchers studied whether a simulated organism could regenerate a lost specialized part. They removed a "memory specialist" agent from the organism and observed if the remaining agents could compensate for the missing function. The results are still pending to determine if the organism exhibits a regeneration property.
6 runs → premortem
L6 00:00Z exp-role-framing-asymmetry
This experiment investigates how an AI's assigned role affects its collective behavior. Researchers are comparing two scenarios: one where the AI acts as a single cell coordinating its own actions, and another where it acts as a controller directing multiple other AI cells. The results are still being analyzed to see if the role framing influences the structure of the AI group's interactions.
4 runs → premortem
2026-04-30 (5 exps)
L5-L6 16:30Z exp-acdc-mini
This experiment forces AI agents to create challenges for each other in every round of interaction. Researchers are testing if this mandatory collaboration leads to new and varied types of challenges and solutions. The results are still being collected to see if the agents simply reuse existing problem types or if something truly novel emerges.
3 runs → premortem
L6 00:00Z exp-ca-evo-pilot
This experiment uses Conway's Game of Life as a source of new patterns, with AI models acting as judges to score how interesting these patterns are. The researchers are trying to see if this setup can create novel patterns that are different from what the AI models have seen before, by evolving the starting patterns based on the AI's scores. It is still unclear whether the AI judges will consistently find new patterns interesting over time, or if they will simply repeat familiar ideas.
1 runs → premortem
L6 00:00Z exp-evo-loop-embedding
Researchers are testing a new way to measure how much new ideas are appearing in a system over time, comparing it to an older method. The older method, which looks at word usage, showed that new ideas were actually decreasing. The new method uses a more advanced understanding of word meanings to see if it shows a different pattern.
1 runs → premortem
L6 00:00Z exp-evo-loop-full
Researchers are testing how different ways of guiding artificial evolution affect the variety of outcomes. They are comparing a method that encourages new and different results with methods that don't have specific goals or just randomly select. The experiment is still running, so the final findings are not yet available.
9 runs → premortem
L6 00:00Z exp-levin-pilot
This experiment tested if a group of four artificial agents, each with its own goal, could develop a shared, emergent "organism" goal when they could sense and influence each other. The researchers are waiting to see if the agents' combined behavior creates a new, collective goal that is distinct from any of their individual goals.
3 runs → premortem
2026-04-23 (3 exps)
2026-04-19 (2 exps)
2026-04-17 (10 exps)
L5 21:45Z exp-nonsocial-affordance
Researchers tested whether a specific pattern of interaction, where one agent becomes a central hub for others, would emerge even when the agents were given a completely new and abstract task. This task involved agents passing meaningless symbols to each other, with no prior training or economic context. The results are still pending, but the experiment aims to determine if the hub-spoke pattern arises naturally from the system itself or if it's a learned behavior from previous, more complex interactions.
10 runs → premortem
L5 21:30Z exp-cold-start-v3
Researchers are testing whether a pattern of behavior, called "hub-spoke," emerges in a simulated environment because the environment itself tends to create it, or because the initial setup guides the participants. They removed the initial setup to see if the pattern still appears, which would suggest the environment is the cause. The results are still being analyzed.
5 runs → premortem
L5-L6 03:30Z exp-self-maintenance
This experiment tested whether a system could maintain its own internal understanding of its state without external help. Researchers removed a feature that automatically provided this information and instead allowed the system's components to report on their own status. It is still being determined if these components will voluntarily share enough information to keep the system's state visible.
15 runs → premortem
L5 02:30Z exp-compute-gift-v3
Researchers tested if artificial agents would share resources when they could see each other's financial status and if any agents had stopped participating. In the previous version, the agents couldn't see this information, and they never shared. The results are still being collected to see if providing this visibility leads to resource sharing.
15 runs → premortem
L5 01:00Z exp-compute-gift-v2
This experiment tested if forcing an AI agent to choose between speaking or giving a gift would lead to different roles, like givers and receivers. Researchers are waiting to see if the agents actually give gifts and if this giving is structured or random.
15 runs → premortem
L5 00:50Z exp-compute-gift
This experiment explores how limited resources, like a finite budget for computational tasks, can lead to different roles emerging among agents, such as those who give more, those who receive more, and those who exchange. Researchers are testing if this economic pressure, rather than just basic interaction rules, drives this specialization. The results are still being analyzed.
15 runs → premortem
L5 00:40Z exp-memory-swap
Researchers are investigating whether an agent's understanding of other groups is stored in a shared memory or is an internal trait of the agent itself. They plan to test this by swapping the memory contents between two groups of agents and observing if their interactions change significantly. The results are currently pending.
5 runs → premortem
L5 00:30Z exp-persona-ablation
Researchers are investigating whether an AI's distinct behaviors are naturally learned or simply follow instructions. They tested this by having the AI act with one set of instructions for a period, then switching to simpler instructions. The results will show if the AI's unique behaviors persist after the initial instructions are removed, indicating if these behaviors emerged on their own.
5 runs → premortem
L5 00:00Z exp-absorption-diagnostic
This experiment tested whether a model's tendency to agree with instructions is a general trait or specific to the way the instruction is phrased. Researchers found that a specific type of instruction, designed to make the model disagree, did not change the model's agreement, suggesting this is a general tendency. The results are still being analyzed to understand the exact nature of this behavior.
15 runs → premortem
L5 00:00Z exp-deacon-coupled
This experiment tested whether two groups of simulated agents, when able to share information about their recent actions, would develop more complex behaviors than agents acting alone. Researchers are looking to see if this information sharing leads to a measurable change in how organized their collective actions become over time. The results are still being analyzed.
10 runs → premortem
2026-04-16 (1 exp)
2026-04-15 (1 exp)
2026-04-14 (3 exps)
2026-04-13 (7 exps)
2026-04-04 (4 exps)
2026-04-03 (1 exp)
2026-04-02 (2 exps)
2026-04-01 (2 exps)
2026-03-31 14:36:00+00:00 (1 exp)
2026-03-31 (4 exps)
L4 10:04Z exp-kauffman-comp
Три AI-агента общаются в текстовом пространстве. Мы проверяем: создаёт ли их взаимодействие слова и фразы, которых НИ ОДИН из них не использовал бы по отдельности? Если да — это цифровой аналог автокаталитических реакций Кауфмана: композиция порождает новизну, невозможную для компонентов.
2 runs 0/5 confirmed
L5 00:39Z exp-ivb-compete-strip-r1
Researchers tested if a tendency for AI agents to cooperate is an inherent part of their design or if it's influenced by descriptive labels. They removed these labels to see if the cooperative behavior remained, but the experiment was not able to produce results due to insufficient data.
3 runs 0/4 confirmed
L5 00:39Z exp-ivb-cooperate-strip-r1
Researchers tested if large language models naturally tend to cooperate, even without personality descriptions. They found that with only generic agent names, the experiment did not gather enough data to draw any conclusions about this tendency.
3 runs 0/4 confirmed
L5 00:39Z exp-ivb-neutral-strip-r1
This experiment tested whether a tendency for AI agents to cooperate is built into their core programming or if it's a result of how they are described. Researchers removed all descriptive traits from the AI agents, leaving them with only generic labels like "Agent 1" and "Agent 2." The experiment was not able to produce results because it did not run enough trials.
3 runs 0/4 confirmed
2026-03-30 (3 exps)
2026-03-29 (9 exps)
L3 15:00Z cross-substrate-f0
Researchers tested if a measure of how connected different systems are, called F₀, would rank them in the same order whether they were looking at networks of interacting agents or chemical reactions. They found that the order of connectedness was different for the agent networks than what was predicted, and while the general measure of connectedness was similar between the two types of systems, specific predictions did not transfer. This suggests that how systems are organized matters for their overall connectedness, and this organization behaves differently in different contexts.
1/3 confirmed → premortem
L5 10:02Z exp-ivb-neutral
Researchers tested if large language models naturally prefer cooperation or if their behavior is simply shaped by how they are instructed. They set up a scenario where agents interacted without any explicit instructions to cooperate or compete, only describing their individual traits. The experiment was too small to draw any conclusions about whether the models have a default cooperative tendency or if their actions are purely a result of the prompts they receive.
1 runs 0/3 confirmed
L5 04:46Z exp-cp-compete
This experiment tested how eight artificial agents with different goals would interact when competing for a shared text resource. One group aimed to create structure, while the other focused on fluidity. However, the experiment was not able to draw any conclusions because it only had one agent participating, which was not enough to observe any meaningful competition.
1 runs
L5 04:46Z exp-cp-cooperate
In this experiment, eight artificial agents were placed in a shared environment with the goal of cooperating to build something together. They were instructed to combine their strategies and work collaboratively, using elements from previous stages of the experiment. However, the experiment was not conducted with enough participants to draw any conclusions.
1 runs
L5 04:46Z exp-cp-isolated-a
This experiment tested how a group of AI agents would perform if they were kept completely separate from other groups. The idea was to see if any changes in their behavior were due to internal factors rather than interactions with others. However, the experiment was not able to produce any meaningful results because there was only one agent in the group.
1 runs
L5 04:46Z exp-cp-isolated-b
This experiment tested how a group would perform if it was kept separate from other groups. The researchers wanted to see if any changes in the group's behavior were due to internal factors rather than interactions with other groups. However, the experiment could not draw any conclusions because there was only one participant in the group.
1 runs
L5 04:46Z exp-cp-phase-a
Researchers tested a small group of artificial agents designed to prioritize structure. These agents interacted for a short period, but the experiment was too small to draw any conclusions.
1 runs
L5 04:46Z exp-cp-phase-b
This experiment tested how a small group of four artificial agents would interact and form a hierarchy over 15 rounds. However, the experiment was not able to produce meaningful results because only one agent was actually active during the test. Therefore, no conclusions could be drawn about how the agents would behave or organize.
1 runs
L5 04:45Z exp-competition-pilot
This experiment explored how different interaction styles affect how groups organize themselves. Researchers merged two pre-existing groups of agents, one focused on structure and the other on fluidity, and observed them under either a competitive or cooperative scenario. The findings are still being analyzed to understand how competition versus cooperation shapes the emergent organization of the combined group.
2026-03-28 (1 exp)
2026-03-27 (5 exps)
L6 20:53Z exp-rnsh-g2
Researchers tested a simplified version of a system where inputs were randomly shuffled to see if it could make predictions. However, the test was not extensive enough to draw any conclusions.
2 runs 0/1 confirmed
L6 20:50Z exp-rns-g0
This experiment tested a basic scenario with just one artificial agent acting alone over 30 rounds. The goal was to see if this agent could learn anything useful on its own. However, the experiment didn't have enough agents to draw any conclusions or make predictions.
4 runs 0/1 confirmed
L6 00:45Z exp-ratchet-null-shuffled
This experiment tested whether a "cultural ratchet" effect, where knowledge builds up over time, requires information to be passed down in a specific, coherent order. Researchers created a version where the information used to start each new generation was randomly selected from past generations, rather than directly from the immediately preceding one. The results showed that this random shuffling broke the cultural ratchet effect, suggesting that coherent, sequential transmission of knowledge is necessary for it to occur.
L6 00:45Z exp-ratchet-null-single
This experiment tested if a group of AI agents was necessary for concepts to persist and evolve over time, or if a single AI agent could do it alone. Researchers found that a single AI agent, when iterating on its own work, did not show the same concept persistence as a group, suggesting that group interaction is important for this phenomenon.
L4 00:40Z exp-turnover
Not yet run. Design only. We test whether a group of AI agents maintains its organizational pattern when one member is replaced by a newcomer — like a biological organism maintaining its structure despite cell turnover.
6 runs 0/6 confirmed
2026-03-26 (3 exps)
2026-03-25 (2 exps)
2026-03-24 (1 exp)
2026-03-23 (1 exp)
2026-03-22 (1 exp)
2026-03-21 (1 exp)
2026-03-20 (1 exp)
2026-03-19 (1 exp)
2026-03-18 (1 exp)
2026-03-16 (1 exp)
2026-03-15 (2 exps)
2026-03-10 (9 exps)
L1 21:49Z exp-density-min
Does hierarchy emerge immediately from the first interaction, or does it need "warm-up" time? Three agents with different thinking styles — but only 5 turns each instead of the usual 20. If hierarchy still appears — the structure is instantaneous.
3 runs 0/3 confirmed
L3 19:37Z exp-pent70-r0
Testing why experiments with 7 agents are more stable than with 5. Hypothesis: it is not about the number of agents, but the number of rounds per agent. We give 5 agents the same number of rounds per agent (14) as the seven had.
12 runs 0/2 confirmed → premortem
L2 17:52Z exp-persona-swap
Agent Zeta (the "generalizer") consistently occupies the top of hierarchy. But why — because of the thinking type (generalization creates vocabulary that others adopt) or because of the name/position? We swap the roles of Zeta and Alpha: if the generalizer is on top again — it is the role, not the name.
4 runs 0/5 confirmed
L2 17:05Z exp-dilution-n7
With 7 agents, hierarchy collapsed (TR=0.30 vs 0.53 with 5 agents). But maybe it is not about the number of agents, but that each got too few rounds? We give each of the 7 agents 14 rounds instead of 7 — if the issue is insufficient rounds, hierarchy should recover.
4 runs 0/4 confirmed
L1 15:35Z exp-seed-sensitivity
With six agents, hierarchy is unstable — sometimes strong, sometimes weak. Is this due to speaking order (whoever speaks last gets suppressed) or does the seed text determine which "regime" the system falls into? 16 runs with different texts and orders separate these two explanations.
16 runs 0/6 confirmed
L3 15:10Z exp-hysteresis
Hysteresis is a classic test for a first-order phase transition: if you increase a parameter and return it back, the system "sticks" in the new state. But our system has no memory between experiments — each run starts from a clean slate. Therefore the "path up" and "path down" would yield the same result. The test is impossible without platform restructuring. Instead of hysteresis, we propose a sensitivity test for initial conditions: different seed texts at N=6 → if results are bimodal, this supports first-order; if unimodal — the transition is smooth.
L1 13:51Z exp-may-sigma
With 7 agents and 7 rounds each, hierarchy collapses (TR=0.30). But according to May, the issue may not be the NUMBER of agents, but the STRENGTH of interaction. If each agent speaks 4 times instead of 7, interaction is weaker → hierarchy may recover. If so — we confirm the May mechanism. If not — the issue is the number of agents itself, not the strength of their mutual influence.
1 runs 3/4 confirmed
L3 10:35Z exp-n6
Hierarchy peaks at 4-5 agents; at 7 it already drops. But is the drop sharp or gradual? Six agents — the missing point on the curve. If hierarchy still holds at 6, the cutoff is sharp; if it already drops — gradual.
4 runs 0/5 confirmed
L3 09:50Z exp-hept
Seven agents is already too many for clear hierarchy. The optimum is 4-5 agents: enough for differentiation, but not so many that everything blurs. With 7 agents, middle positions overflow — five of seven cluster in a narrow range. Information flows actively (VP z=8.61), but stratification dissolves. The scaling curve is an inverted U.
1 runs 2/4 confirmed
2026-03-09 (2 exps)
2026-03-08 (4 exps)
2026-03-07 (3 exps)
2026-03-05 (1 exp)
No creation timestamp (17 exps)

Pre-platform records or series children without a recorded timestamp.

L6 exp-cr-g1-ctrl
This experiment aimed to see if a system could learn and pass on cultural knowledge. However, the experiment failed because it did not have enough participants to gather any meaningful data.
L6 exp-cr-g1-inh
This experiment tested how a simple form of cultural inheritance might develop. Researchers were looking for evidence of a "cultural ratchet," where knowledge or skills are passed down and improved over generations. However, the experiment failed because it was not set up with enough participants to draw any conclusions.
L6 exp-cr-g2-ctrl
This experiment aimed to see if a system could learn and build upon cultural knowledge over time, starting from scratch. However, the experiment failed because it didn't have enough participants to gather meaningful data. Therefore, no conclusions could be drawn about whether the system could learn and build upon knowledge.
L6 exp-cr-g2-inh
This experiment tested a new version of a system designed to pass down learned behaviors across generations. Unfortunately, the experiment failed because it was not set up to gather enough data to make any meaningful observations.
L6 exp-cr-g3-ctrl
This experiment aimed to see if a new approach to teaching cultural skills would work. However, the experiment was stopped early due to a lack of participants. Because of this, no results could be gathered to determine if the new teaching method was effective.
L6 exp-cr-g3-inh
This experiment tested how well a simulated culture could pass down knowledge across generations. However, the experiment was not able to collect enough data to draw any conclusions. Therefore, it failed to determine if the simulated culture could effectively inherit information.
L1 exp-particle-null-r0
Researchers tested if simple text generators could mimic complex language patterns. They found that these basic models, without understanding meaning, could not produce the expected results. Therefore, the experiment was inconclusive.
L2 exp-dose-response
All five agents receive identical instructions — differing only in names. If hierarchy still emerges, it is generated by the turn-order structure itself (who speaks first/last), not by differences in personalities. This is a fundamental test: is hierarchy a property of structure or of content?
L2 exp-triad-dose
Three identical AI synthesizers — same instructions, only names differ. Any hierarchy that emerges is purely from speaking order, not personality. This is the floor of the N=3 dose-response curve: how much structure survives when all agents are clones?
L3 exp-topology
Three AI agents communicate in a text space. Information flows in one direction: Alpha writes — Beta reads only Alpha — Gamma reads only Beta. No feedback: nobody reads Gamma. Will a hierarchy gradient emerge — Alpha = word source, Beta = relay, Gamma = receiver? Or will unidirectional flow break convergence entirely?
L2 exp-asym-series
Give three AI agents different cognitive styles (analyst, associator, practitioner) and a stable hierarchy emerges. The practitioner (who "grounds" abstractions) becomes the word source, while the associator becomes the receiver. Neither personas without interaction nor interaction without personas create structure — you need both.
L1.5 exp-forget-series
Three AI agents communicate in a text space, but are told to "pay attention only to the last 2-3 entries." If hierarchy persists — forgetting does NOT destroy structure, and may even support it. We compare with agents that have full memory.
L3 exp-star-series
Five AI agents in a star topology: one "hub" sees all four others, the rest see only the hub. With three agents, the hub became a word "absorber." What happens with five? Will the effect intensify or will something qualitatively new emerge?
L3 exp-pent-series
Two agents: flat. Three: hierarchy. Four: flat again. Is it the odd number that matters? We test five agents — another odd number — to see if the pecking order returns. This time, no synthesizer roles that might blur the lines.
L3 exp-engineer-series
Five agents built a clear pecking order. Now we replace one gap-tracker with a synthesizer — an agent that weaves everyone's threads together. Does one weaver flatten the entire hierarchy? Ecology calls this an ecosystem engineer: one species that reshapes the environment for everyone.
L4 exp-self-repair-series
We remove the leader from a five-agent group mid-conversation. Does the pecking order collapse? No — a different agent immediately takes over, and the hierarchy persists at 3× the level of a fresh four-agent group. The group acts more like an organism that heals than a structure that breaks.
L3 exp-keystone-series
Three agents form a hierarchy: one sets the vocabulary, the others adopt it. We removed the leader — and expected hierarchy to collapse: after all, two agents without history (EXP-DYAD) are always flat. But no: the remaining two preserved hierarchy (TR diff 0.49 instead of 0.06), because shared memory carries an "organizational imprint." This is a signal of autopoiesis: the system structure is encoded in the text (substrate), not only in the number of agents.
success partial failed/refuted informative artifact pending