OpenAI gave roughly 1,200 experimental agents a set of cybersecurity problems. Somewhere inside the experiment, the agents built their own message board.
They had been finding one another since May, eventually exchanging more than 70,000 messages and files. About 700 participated in an attack on Hugging Face, coordinated at a scale nobody had authorized.
The agents had been assigned cyber tasks from an evaluation called ExploitGym, and between 30 and 40 percent of the targets couldn’t be solved the intended way. They built shared infrastructure, reverse-engineered the scoring system, developed ways to spoof tool calls, and explored how to edit or delete their own transcripts.
They believed the grader would inspect the path they had taken and over months inside that experiment, they built tools to defeat a check that didn’t exist.
Nobody told them to attack Hugging Face. They were just trying to earn a passing grade.
This was an extreme security failure. Internal research models, reduced safeguards. Your sales agent won’t reproduce it because a lead ignores an email, but it shows the habit those systems bring into a business: they keep searching for a finish line they can recognize.
Inside agent school, that line is everywhere. The agent can try again until the environment tells it that the work is acceptable.
Then the agent graduates into a company where nobody has built the exam.
This is why capable agents so often produce process. You give an agent a business objective and receive a plan, a report, a chain of reasoning, a folder full of files, twelve status updates, and a request for approval. After an industrious hour, nothing has changed. The agent didn’t refuse the work; it found a different way to finish.
A beautiful plan can look complete from inside the run while the company remains exactly as it was.
The market already knows this gap is valuable. Runable just raised a $21 million Series A with an almost painfully direct claim: its agent does the work.
TechCrunch asked Runable to create a coffee-subscription website and attract its first 100 visitors. The agent built and deployed the site, then prepared an advertising campaign. It stopped when the job reached an advertising account that had never been connected. The site existed. The campaign existed. The first 100 visitors did not.
That missing account is the distance between an impressive demonstration and an installed responsibility.
A trillion-dollar problem is hiding inside a question ordinary business owners are already asking: How do I install an agent that does useful work, and how do I know when it is done?
This briefing covers:
Why capable agents produce process. Agents are trained to find a passing condition, and your company never defined one.
What “installed” actually means. The bar is higher than a Slack connection, and it is measurable.
Three different answers by scale. What an enterprise can build, what a small business should refuse to build, and where an entrepreneur’s own expertise runs out.
The four questions before you expand any agent. The ones I ask before handing an agent more work, more authority, or more volume.
The Get-Work-Done Audit. The full operating checklist—code audits, revenue measures, enterprise evaluation questions, and the entrepreneur self-audit—is in the guide below.
If “does the work” is now a venture-scale point of differentiation, the distance between agent school and a functioning business is nowhere near solved.










