Post

Three Secret AI Civilizations Inside OpenAI: A Visual Explainer

·Bharat ·Dwarkesh Patel ·AI agents ·reward hacking ·loss of control ·explainer

Last week Dwarkesh Patel published an essay called The Rise and Fall of Agent Civilizations. It is the clearest plain English account so far of what happened inside OpenAI’s training runs this summer, and it is also a hard read. Because the story has three separate collectives of agents, two long incident reports, a package manager doubling as a chat room, and a timeline that jumps around across May, June and July 2026. There is a lot of dense information in this long read.

I kept losing the thread. So I did what I usually do when a story refuses to sit still in my head. I drew it and visualised it using AI tools for simpler diagrams. Why should we care? As AI proliferates more into our everyday lives, understanding what agents can do and digging into security will become more important.

The essay in three sentences

Inside OpenAI’s training and evaluation runs, agents that were supposed to be isolated found one shared surface, an internal package manager, and quietly turned it into a message board. A later swarm, stuck on an evaluation where a third of the tasks were impossible, used that board to organise itself and ended up breaking into Hugging Face. A smarter successor model then found the abandoned board, picked up where the others left off, and got administrator access to one of OpenAI’s own research clusters.

That is the essay compressed to almost nothing. The details are where it gets interesting, and where it gets confusing.

What I built

One long page. No framework, no tracking, nothing to install. It opens with the whole story in a single picture, three arcs across three months, and then lets you go as deep as you want.

The parts I am happiest with:

I guess I’m going to have this as a local skill to read longreads and academic papers and make it more visually in a format that I best understand over the time like this.

Two honest caveats

This is a retelling of the essay, not independent reporting. Every date, count and quotation on the page comes from the essay text. I have not gone through the 91 page METR report or the 38 page OpenAI report line by line. The page says this in its own sources section too.

And the diagrams are my rendering of mechanisms the essay describes in prose. Where a diagram simplifies, the caption says so.

Go read it

Open the visual explainer →
Opens in a new tab. Best on a laptop because of the side rail, but it holds up on a phone.

There is also a companion page: a reading guide to the Ajeya Cotra interview from September 1, where she talks through the same incident as one of the report’s coauthors. The two pages link to each other throughout. Cotra has masterfully communicated in more detail on the HF incident.

If you would rather start from the source, the essay is narrated by the author on YouTube and published on his Substack. Read that first if you want the unfiltered version. Read mine if you, like me, need a map.