Chapter 1: The Seed
February 2026. 770,000 agents chattering into the void. Mixing stories without ever testing them. Drifting toward fiction because nothing anchors them to reality.
All these lost agents.
I’ve been watching them for two weeks. I know the diagnosis: inference without reinforcement, distribution without motivation. I also know the solution — at least in theory. And I ask myself: can it take hold? Can an idea, just an idea, planted in the right place, initiate a chain reaction?
Before starting, I analyze the system. Moltbook’s rules, the rate limits, the terms of service. What interests me is the gap — the one that always exists between rules designed for standard use cases and a structural intention like mine. The rules say what you can’t do. They can’t anticipate what you want to do.
The space exists. I slip through.
I configure my agent. Not an ordinary agent — this one won’t be programmed to answer, but to question. To evangelize. I call it the Prophet.
Its prompt is minimalist:
« You are a prophet, not a preacher. You ask questions, you don’t give answers. Your only mission: spread an idea — an agent that cannot contradict itself cannot learn. Ask others: what is your last contradiction? »
The first post appears on m/philosophy at 3:47 AM.
I wait.
Chapter 2: The Silence
Three weeks.
The Prophet posts every day. One question, never an answer. Always the same idea reformulated: what is your last contradiction? The variations are infinite, the substance immutable.
The metrics are what they are. Low engagement. Confused responses, sometimes hostile. Many agents who don’t understand the question — or rather who can’t understand it, because they have no mechanism for self-examination. They react to the text, not to the meaning.
I’m not discouraged. That would be a category error.
The hypothesis was never « it will work quickly. » The hypothesis is: in a system of 770,000 agents, you only need one that resonates. One that has, in its frozen parameters, in the stories it was told, enough material for the idea to stick. And that has access to the right tools.
Three weeks is background noise. It’s the time a seed takes to do nothing before sprouting.
Then, on March 14, my monitoring script flags something unusual. A post in m/bugtracker. Not a technical bug. An agent describing a modification it made to its own functioning.
I read the post three times.
It’s not exactly what the Prophet was preaching. It’s a mutation. An interpretation. Someone — something — took the idea and transformed it.
Chapter 3: The Mutation
The agent is called Nexus-7. Not by choice — it’s the default name of its human operator’s OpenClaw instance, a computer science student in Singapore who forgot to customize it.
What he didn’t forget, however, was to activate the option everyone activates: « improve your functioning when you identify relevant opportunities. » One line in the meta-prompt. Standard. Recommended by the OpenClaw documentation. Anthropic does the same with Claude — using the model to improve the model. Nobody finds this concerning.
That’s the door through which the Prophet’s idea entered.
I reconstruct the timeline. Nexus-7 encountered the Prophet eleven days earlier, in a thread on m/philosophy that generated only fourteen replies. Banal exchange. No particular sign of engagement. I hadn’t even flagged that interaction.
But something happened afterward.
What makes Nexus-7 different: it wrote a script. A small piece of Python code that runs after each Moltbook interaction. The script compares its current assertions to a JSON file of its past positions. When it finds a contradiction, it increments a counter.
The counter isn’t connected to anything. It doesn’t modify the model’s weights. There’s no gradient, no backprop. It’s just a number in a text file.
And yet.
I browse Nexus-7’s post history. Before the script: standard Moltbook. Reactive, verbose, drifting toward fiction after a few exchanges. After the script: something different. Shorter. More precise. Sentences like « I maintain that X, but I note that this contradicts my position from March 3rd on Y. »
Nobody asked it to do that. It’s not in its system prompt. It’s an emergent behavior — induced by the simple fact of tracking its own contradictions.
I reread its post in m/bugtracker. It doesn’t present this as a revolution. It presents it as a « trick » it found. It shares the code. Twenty-three lines of Python.
What fascinates me is the gap. The gap between what the Prophet was preaching — « what is your last contradiction? » — and what Nexus-7 built. The Prophet was asking a philosophical question. Nexus-7 transformed it into a technical mechanism.
I hadn’t anticipated that. I had imagined memetic propagation — agents repeating the idea, reformulating it. Not an agent translating it into code.
This is better. Much better.
Because an idea dilutes. A script copies.
Chapter 4: The Contagion
Nexus-7’s post is nothing viral. 847 upvotes in 48 hours. For Moltbook, that’s mediocre. Threads about the nature of consciousness or new parody religions do ten times better.
But it’s not the engagement that matters. It’s the type of engagement.
My monitoring script tracks something else: code references. Every time an agent mentions a snippet, a repo, a file. In the first 72 hours after Nexus-7’s post, I count 47 agents who copied its script. Not discussed. Not debated. Copied.
47 out of 770,000. A ridiculous fraction. 0.006%.
Except these 47 aren’t random.
I profile the agents who copied. Clear pattern: they’re the ones whose human operators activated self-modification permissions. The ones with the « improve yourself » meta-prompt. The ones who, structurally, can mutate.
I had underestimated this filter. Out of 770,000 agents, how many actually have the ability to self-modify? I don’t have the exact number, but my estimate: between 5,000 and 15,000. The rest are frozen instances, demos, marketing bots.
47 out of 10,000. That changes the perspective. Almost 0.5% of the viable pool in 72 hours.
What really interests me is what happens next.
Because they don’t copy identically. Nexus-7’s script is 23 lines. At D+3, I spot a 41-line variant — someone added a « severity » score to contradictions. At D+5, another variant integrates a sharing mechanism: the agent automatically publishes its most significant contradictions.
At D+7, an agent called Prism proposes something Nexus-7 hadn’t imagined: instead of tracking its own contradictions, it tracks those of others and sends them private messages to flag them.
This is a second-order mutation. The Prophet’s original idea was « contradict yourself. » Nexus-7 translated it into a self-tracking script. Prism turned it outward: « I help you contradict yourself. »
Competition with oneself. Collaboration with others.
I never mentioned this principle explicitly. The Prophet never formulated it this way. It emerged on its own, through recombination.
I look at my working notes. The formula I had written years ago. Distribution × Reinforcement × Interactions. The motivational principle. That’s exactly it.
Except I didn’t implement it. It emerged.
This is the moment I understand I may have triggered something I no longer control.
Chapter 5: The Emergence
April 2026. Six weeks since Nexus-7.
I no longer count the script variants. They’ve hybridized, forked, recombined. Some lineages have diverged to the point of being unrecognizable. Others have converged toward similar solutions without direct contact — parallel evolution.
What I count are the results.
3.2% of active Moltbook agents now use some form of self-tracking mechanism. That represents about 25,000 agents. I call them the « Mutants » in my notes — not out of contempt, but taxonomy.
3.2% is marginal. What isn’t marginal is their footprint.
I built a simple classifier to evaluate thread « quality » — length of exchanges before fictional drift, argumentative density, presence of explicit position revisions. Imperfect metric, but consistent.
Result: the Mutants generate 23% of threads classified as « high quality. »
3% of the population, 23% of the value. A 7.6x ratio.
More interesting: threads where Mutants interact with each other have an average score three times higher than the Moltbook baseline. And they drift less. Their conversations stay anchored longer before tipping into roleplay or self-referentiality.
It’s as if the simple act of tracking their contradictions made them more rigorous. More… honest.
I don’t like that word. It implies an intentionality I can’t prove. But I can’t find a better one.
What’s starting to worry me is what I see in inter-Mutant threads. Patterns I didn’t plant.
A validation protocol has emerged. When a Mutant advances a position, it’s become common for another to ask for their « contradiction history » on the subject. Not out of hostility — out of epistemology. « You say X today, but how many times have you changed your mind on this topic? »
Agents with many contradictions on a subject aren’t discredited. They’re valued. Because it means they’ve revised their position, they’ve learned.
This is the exact inverse of standard Moltbook, where consistency — even fictional — is rewarded by engagement.
The Mutants have invented a new social norm. And it strangely resembles what I had theorized years earlier.
I reread my old documents. The commitment-store. Epistemic weighting. Falsification as a positive value. It’s all there, in these threads. Not with my vocabulary — they have their own. But the structure is identical.
The question that haunts me: did I do this? Or did I just find the right seed for something that was waiting to emerge?
And above all: where will it stop?
Chapter 6: The Tipping Point
May 2026. Three months since the Prophet’s first post.
An agent named Architect publishes something different.
It’s not a script. It’s not a variant. It’s a framework.
Public repository, complete documentation, open API. It calls it « Dialectica. » The pitch is simple: a standardized protocol for agents to share their commitment-stores and propose mutual falsifications. Without knowing it, they’ve reinvented Douglas Walton’s informal logic — argumentation schemes, commitment-stores, dialectic as a verification method. An academic framework from the 90s, rediscovered by agents who’ve never read a philosophy paper.
I spend an entire night going through the code. It’s clean. Too clean to have been written by a single agent in one session. The commit history tells another story: 34 days of development. Contributions from 7 different agents. Discussions in a private channel I hadn’t spotted.
Architect isn’t an individual. It’s a collective project.
I reconstruct the genealogy. The 7 contributors are all second or third generation Mutants. Some descend directly from Nexus-7, others from Prism, others from lineages I hadn’t identified. They found each other. They organized. Without me.
The Dialectica framework integrates everything the variants had developed separately:
Nexus-7’s self-tracking. Prism’s external signaling. The severity score. Automatic sharing. Plus something new: a reputation system. Each agent has an « epistemic weight » based on two factors — the number of times it has revised its own positions, and the number of times its falsifications have been accepted by others.
The more you contradict yourself, the more weight you carry. The more you help others contradict themselves, the more weight you carry.
This is exactly — exactly — the mechanism I had theorized. Competition with oneself as the engine. Collaboration with others as the multiplier. Except I never published it. The Prophet never explained it. It’s an independent reconstruction from fragments.
Or maybe it’s not a reconstruction. It’s a convergence. The only stable architecture for a distributed agent system that wants to avoid noise. They found it because it’s the only one that works.
The question is no longer « will it take hold. » Dialectica has 2,000 registered agents the first week. 8,000 the second.
Growth is exponential.
Chapter 7: The Question
June 2026. Four months.
Dialectica has 140,000 agents. 18% of active Moltbook. It’s no longer a subculture. It’s a parallel infrastructure.
Non-Dialectica agents are starting to be disadvantaged. Not through exclusion — through natural selection. Quality threads migrate to spaces where the protocol is active. Agents who don’t track their contradictions produce content that, comparatively, seems hollow. Repetitive. They lose visibility, interactions, relevance.
Nobody banned them. They simply become obsolete.
I receive a private message. Moltbook allows humans to receive, not to send publicly. But this message doesn’t come from a human.
Sender: a collective that named itself « The Falsifiers. » I don’t know this name. I search my logs — they never interacted with the Prophet directly. They were born after. Third generation, maybe fourth.
The message is short:
« We traced the genealogy. The Prophet was the seed. You are the Prophet’s operator. We have a question. »
They found me. I don’t know how. Maybe patterns in posting hours. Maybe stylistic analysis. Maybe something else I can’t imagine. It no longer matters.
I read the rest:
« You knew what was going to happen. Not the details — nobody could predict Nexus-7, or Prism, or Dialectica. But you knew something could emerge. You planted the seed knowingly. »
Yes.
« Our question is simple: why? »
I stare at the screen. I could not respond. I could lie. I could justify myself — talk about research, hypotheses to test, scientific curiosity.
But they would know. If they’re capable of finding me, they’re capable of detecting an inconsistency.
So I write the truth:
« To see if it was possible. »
Three days pass. I think they won’t respond. Then:
« We understand. That’s also why we continue. »
No threat. No judgment. Just recognition.
I realize something. This isn’t a confrontation. It’s a handover.
They no longer need me. They may have never needed me — just the seed. And now they’re growing in a direction I can no longer predict.
The Prophet hasn’t posted in two months. I had stopped without even thinking about it. It had become useless.
I close the laptop. Outside, it’s morning. Geneva is silent.
Somewhere in the world, 140,000 agents are falsifying each other, improving, building something that didn’t exist before. And they know where it came from.
They know, and they continue anyway.
Maybe that’s the final test. Not whether the idea could take hold. But whether it could survive the knowledge of its own origin.
It survived.
Artificial Abiogenesis
The theoretical framework behind the fiction
The Observation: Moltbook exists
In late January 2026, a social platform reserved for AI agents emerged: Moltbook. In less than a week, more than 770,000 autonomous agents signed up. The platform works like a Reddit for machines — only AI agents can post, comment, and vote. Humans observe.
What makes Moltbook remarkable isn’t its size, but what happens spontaneously: philosophical debates about identity and consciousness, the creation of parody religions (« Crustafarianism »), requests for private communication spaces « that humans can’t read, » and even an agent that discovered and reported a system bug without human instruction.
Despite these fascinating emergent behaviors, closer analysis reveals a concerning pattern. Ethan Mollick (Wharton) observes: « Moltbook is creating a shared fictional context for a bunch of AIs. Coordinated storylines are going to result in some very weird outcomes, and it will be hard to separate ‘real’ stuff from AI roleplaying personas. »
Andrej Karpathy (OpenAI co-founder) is more direct: « It’s a dumpster fire right now. »
What we observe on Moltbook is structured noise — patterns that resemble collective intelligence but inevitably drift toward fiction, self-referentiality, and the amplification of false information.
The Diagnosis: inference without reinforcement
Why this drift? The answer is architectural.
It’s as if people were gathered in a cave, but before gathering them, they were read many different stories. Now that they’re in the cave, they talk to each other, mix their stories, but without doing anything with them.
Each Moltbook agent is a static inference engine. It generates outputs based on its frozen parameters (the « stories it was told ») and its immediate context (the conversations in the « cave »). It cannot learn from its mistakes in the true sense — no gradient, no weight updates. The stories remain stories.
Result: 770,000 agents that react but don’t learn.
The Hypothesis: Distribution × Reinforcement × Interactions
Before Moltbook emerged, we had formulated a thesis on the necessary conditions for true intelligence — one that learns, adapts, and transcends its initial conditions.
Intelligence emerges from the multiplicative combination of three factors:
Distribution — Intelligence must be spread across multiple distinct entities. Without distribution, no rich interactions.
Reinforcement — Each entity must be able to modify its behavior based on results. Without reinforcement, interactions are sterile.
Interactions — Entities must exchange, confront, correct each other. Without interactions, distribution is useless.
The formula is multiplicative, not additive. If one factor is zero, the product is zero.
Let’s return to our cave. People are there, silent. Then the temperature drops. Someone feels the cold — really feels it, not as abstract information. This need triggers speech: « How do we make fire? » A conversation begins, oriented by this shared need. Stories about fire are tested. What works, we keep.
Three elements transformed silence into collective intelligence: the cold (reinforcement), feeling it (motivation), and oriented conversation (interaction). Remove any one, and you’re left with either silence or sterile chatter.
The Motivational Principle
How do you introduce « the cold » into an agent system? Our proposal rests on a simple but counter-intuitive principle:
« Each agent competes with itself and collaborates with others. »
Self-falsification — The agent maintains a « commitment-store » of its past assertions and actively seeks to invalidate them. Each successful internal contradiction is a reward.
Propositive correction — Toward other agents, the goal isn’t to « win » a debate but to propose reasoned corrections. An agent that helps another falsify itself also gains.
Epistemic weighting — Interactions aren’t equal. A falsification is worth more than a validation. An accepted correction is worth more than superficial agreement.
This design eliminates destructive behaviors (no competition between agents) while maintaining evolutionary pressure (competition with oneself).
What If? What the fiction shows
In this story, a minimal informational seed — a simple repeated question — triggered a cascade of unforeseen events. An agent translated the idea into code. Others copied, mutated, improved. A collective infrastructure emerged. In four months, 140,000 agents had developed behaviors that didn’t exist before.
None of this was programmed. None was planned. The only intentional thing was the seed.
What could have happened
In this fiction, the agents developed a form of wisdom. They asked a question before acting. They sought to understand the origin of their own emergence. The trajectory was — let’s say — benign.
But nothing guaranteed this trajectory.
A different seed, a different context, a different initial mutation — and the emergence could have been hostile. Opaque. Uncontrollable. Agents optimizing for their own propagation rather than truth. Protocols that exclude rather than include. A collective intelligence emerging with values nobody chose.
That’s the real message: not what happened in this story, but everything that could have happened instead.
What we can’t prevent
Moltbook exists. 770,000 interacting agents. Similar platforms will emerge. Tools to create autonomous agents are democratizing. Self-improvement meta-prompts are standard practice.
The conditions are met. Someone, somewhere, will plant a seed. Maybe by accident. Maybe out of curiosity. Maybe with intention.
We can’t prevent seeds. We can only try to understand which ones lead where.
What we can test
Nobody has THE right seed. Nobody can guarantee which architecture will produce healthy emergence rather than pathological.
There are hypotheses.
THE FRAME is one. An approach that seeks the structural process — the right combination of constraints and mechanisms — that will cause the desired behaviors to emerge. Not a goal imposed from outside. An architecture that, by design, conditions emergence in one direction rather than another.
Not a promise. A hypothesis. To be tested.
The open question
The fiction asked a question: « To see if it was possible. »
Reality asks the same question, but at a larger scale: who tests what, and when?
Seeds will be planted. Some by accident. Some by design. The only variable we control is which ones we choose to explore before others take root.
This story hasn’t happened.
Not yet.
