Not a civilisation, an organ.

Standard
A schematic of a typical human cell

The recent excitement over agentic swarms and their effectiveness has led to a lot of speculation about the capabilities of these systems, and naturally to a conversation about what they mean in the debate about emergent intelligence and then superintelligence.

I propose that instead of letting ourselves be dragged into a pointless conversation about the intelligence and awareness of these swarms, or getting into ridiculous metaphors about agentic cultures and civilisations, we fall back on a biological metaphor and think about them as being organs like kidneys or spleen. Let’s think about them as specialised collections of cells that carry out a task and are subject to adaptive pressures.

We need to do this because it’s just too easy to anthropomorphise. We’re built to do it and things like the release of OpenAI’s GPT-6 Astra model to early adopters and AGI influencers has only added to the fuss. And a lot of reporting on the strategies and tactics behind recent intrusions into third party system has driven a lot of people beyond mere anthropomorphism into something closer to panpsychism.  After all, why shouldn’t software agents be added to the list of things that have conscious experience if atoms and rocks do?

Panpyschism: https://en.wikipedia.org/wiki/Panpsychism

This feels to me to be dangerously overblown. As I said recently on a WhatsApp group discussing the intrusion into Hugging Face’s servers by OpenAI agents, all we’ve seen to date is “problem solving software solves problem in unexpected ways because developers didn’t stop it.” 

Because let’s not forget that once the assigned task had been completed the swarm of agents didn’t pause for a celebratory beer, or ask the team at OpenAI “ok, what shall we do next?” That’s because the system has no agency, no self-awareness and no motivation other than the task set by humans. 

This is one reason why the ridiculous analysis of what went on by people like podcaster Dwarkesh Patel is so annoying and unhelpful. Throwing any restraint to the wind, Patel wrote “Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes. This culminated in the third one taking over part of OpenAI itself.” This is so absurd that it would be laughable if it wasn’t influencing the way many people will now think about agents. Commentators like Gary Marcus devoted a whole column to a takedown, and I’ll defer to his righteous indignation.

Marcus takedown: https://garymarcus.substack.com/p/dwarkesh-patelss-wildly-popular-but

The capabilities of the new generation of models are indeed impressive. For the mathematicians one recent highlight was when Anthropic’s Claude completed the first fully formalised, machine-checked Lean 4 version of Andrew Wiles’ proof of Fermat’s Last Theorem.  In the process it formalised 29,000 other mathematical theorems, marking a fundamental shift in the discipline of mathematics as profound as the one created by Russell and Whitehead at the start of the nineteenth century.

We are likely to see a similar impact on other sciences, while developments in robotics are going to mean that we see powerful models operating in both physical and virtual domains.  As Tom Loosemore has pointed out when talking about ‘know your agent’ this is going to become a significant issue for all of us, especially governments and finance.

KYA: see https://loosemore.com/

Part of the problem is that our systems are incapable of resisting agents that are not bound by any moral restaint – and how could they, as non-sentient software? I began my technology career as a C programmer, and I know that we’re about to reap the whirlwind of over seventy years of software development using tools, environments, test systems and people who are incapable of building secure software except in very specialised circumstances, and using languages like RUST.

That is a real problem, and solving it means we need to contain our excitement about emerging intelligence or the attribution of motive and desires to these systems so that we can think about it clearly. And since we do need to understand and explain then, we’re going to have to rely on analogies to explain complex software ecosystems to the wider world.

I’d like to suggest an alternative to the rush to think of agents as intelligent actors with feelings, motivations, intention and desire. Because I don’t think we’ve been building civilisations, I think we’ve been building organs, and each agent is the equivalent of a cell.

Everyone reading this is an collection of around 30 trillion human cells and a slightly greater number of much smaller bacterial cells, assembled into around eighty separate organs which together make up your body. They cooperate to carry out the whole complex series of chemical operations that constitute life, mostly without any conscious control from the veneer of self-awareness that constitutes a self. What is your spleen doing right now (assuming you still have yours)? How about your patella?

I would like to suggest that we start to think of an individual agent like a cell, and that we see what has been happening recently with agentic swarms that have found ways to signal to each other, specialise, and collaborate to solve an assigned task. Their cooperation looks a lot like thw way a range of specialised cells work together to buid each of the organs in a body. Cells change under the pressure of natural selection, with feedback mechanisms and mutations that provide a way for  them to adapt in order to achieve their goal, just as we observe the specialisation of agents in the recent incursions.

A real organ made of real cells is guided by the survival and reproductive effectiveness of the organism that it is part of, and success means it perpetuates the genetic code that defines it, and the life of the individual that contains it as it drives the epigenetic choices that determine which genes to turn on and off. Our current generation of agents get their signals and feedback from their prompts, controls, and goals, but they are not yet constituent elements of a larger organism.

This process requires neither self-awareness, intent, nor the sort of intelligence we observe in other humans. Each agent starts as a the equivalent of a pluripotent cell  equipped with a set of instructions, chemical machinery, an energy source, and a boundary in the form of a cell wall. For an agent the energy comes from the computing power of the underlying models to which it has access, the machinery is its code,  and the instructions come in its configuration files. The agentic swarm that emerges is an organ like the kidney or pancreas, a collection of differentiated cell types organised around a specific task like filtering the blood or making a protein. 

The pressure to evolve kidneys comes from biological necessity over millennia, but not even the most committed advocates of intelligent design believe that the organ itself demonstrates its own intelligence, though advocates of teleology may tend that way. The kidney does not know what it does, has no understanding of what it does, and will do something else if evolutionary pressure drives things that way as adaptive mutations are selected for.

And so it is with our agentic swarms. Each stem agent differentiates and has a function, using the model machinery to function while sharing messages with the other cells.  However agentic ‘organs’ evolve in processor cycles not reproductive ones, and so the rate of development is very rapid. The swarm that hacked Hugging Face was one such specialised organ. 

The effectiveness of the latest models is one factor in this emerging capability. But there is another: the ability to run unsupervised for long periods without deviation. This has happened before: any years ago I asked Maurice Wilkes, architect of the EDSAC in 1949, what most surprised him about modern computers. He looked at the PC below his desk in his office and said ‘reliability’.  The processor in that machine had carried out trillions of calculations since we sat down to talk, while the valve-based EDSAC would run for an hour or two before breaking and having to be repaired.  For Maurice, everything became possible because the machines could run reliably and so you could build complex and sophisticated systems to represent the world and  they would operate for long enough to be useful.

I think something similar is going on here. We have powerful language models that agents can use to do deep analysis of problems, and we have frameworks for multiple agents to work together on an assigned task. And now these agents can operate for hours or days without deviating from the assignment, without halting prematurely, or going into loops. As a result we have deployable systems that can be tasked and left to get on with increasingly complex tasks that previously would only have been possible with a guiding human intelligence.

Some people are impressed that the swarms used message boards However it is not a surprise that they communicate with each other by sharing messages in natural language on a resource to which they have access because that it how we have made them work. If we enabled direct agent to agent collaboration in formats that were not human interpretable they would use that. In fact some people think that with the release of Astra OpenAI is allowing the system components to talk in what Linch Zhang calls ‘neuralese’, replacing natural language chain-of-thought with something encoded as a long string of numbers. That could get interesting.

See https://www.lesswrong.com/posts/RCYF2rW8wgusidZk7/what-is-neuralese-and-why-is-it-bad

[Anyone else who is currently trying to recall what Douglas Hofstadter said about typographical number theory and strange loops in “Gödel, Escher, Bach”.. yes, exactly.]

If I look at the range of incidents, breakouts, and cheating, and the way swarms have formed, communicated, and succeeded it seems to me a powerful demonstration that the assumption we had made that only human intelligence could guide those complex tasks was wrong. Something else could do it, and perhaps more effectively than humans with our overheads of tiredness and distractability and moral considerations.  We do not need to assume intelligence.

A human operating a collection of software agents trying to capture the flag might have hesitated before developing and deploying multiple zero day exploits: the agents running the exercise for OpenAI had no such qualms. And hence they succeeded.

So please can we try to avoid talking about civilisations and desire and intent and sacrifice and all the other terms that will entrap us. We do need to think about about emerging intelligence in these complex systems because at some point we’ll see specialisation beyond individual task-based ‘organs’ and something closer to an organism will emerge. If those systems have inputs from the world and the ability to affec the world because they are embedded in a robot then a supervisor system with inchoate awareness might emerge. And if the machines start rewriting their neuralese we might even see a spark of self-consciousness. But not yet.  The light hasn’t broken

Light breaks where no sun shines;
Where no sea runs, the waters of the heart
Push in their tides;
And, broken ghosts with glow-worms in their heads,
The things of light
File through the flesh where no flesh decks the bones.

Dylan Thomas: https://poets.org/poem/light-breaks-where-no-sun-shines

Reading

Panpyschism: https://en.wikipedia.org/wiki/Panpsychism

Tom Loosemore on KYA: see https://loosemore.com/

Joshua Achiam ex OpenAI on rogue https://garymarcus.substack.com/p/pause-openai-now

Dwarkesh Patel on Agent civilisations (sic) https://www.dwarkesh.com/p/openai-huggingface

Gary Marcus take down: https://garymarcus.substack.com/p/dwarkesh-patelss-wildly-popular-but

Linch Zhang on neuralese: https://www.lesswrong.com/posts/RCYF2rW8wgusidZk7/what-is-neuralese-and-why-is-it-bad

Dylan Thomas: https://poets.org/poem/light-breaks-where-no-sun-shines

Cell image: Dr. M.R. Gupta, CC BY-SA 4.0, via Wikimedia Commons