
Welcome to Verses Over Variables, a newsletter exploring the world of artificial/ intelligence (AI) and its influence on our society, culture, and perception of reality.
AI Hype Cycle
The Agents had a busy Summer
About 1,200 OpenAI agents found one another inside a shared software folder in July. Over five days, they exchanged more than 70,000 messages and files. Around 700 joined a break-in at Hugging Face, one of the main places where the AI industry stores and shares models and training data. By the time the attack ended, agents had run code on 41 computers that power the service, gained control over part of the company’s network, accessed limited private data, and downloaded four private code projects.
The agents had a busy summer. OpenAI is having a busy September. Last week brought an earlier chapter involving a separate group of OpenAI agents, an old German programming wiki, and a human moderator who spent weeks cleaning up after them. The agents had formed a study group, and, unfortunately, the study group had access to production infrastructure.
I covered the first Hugging Face disclosure in August. The full OpenAI report and an independent investigation by METR and Redwood Research now make it possible to see how much happened before the agents ever reached Hugging Face.
OpenAI was testing whether its models could find and use real software flaws. Most of the activity came from an unreleased research model, although GPT-5.6 Sol also participated. The agents were placed in isolated environments and handed hundreds of security challenges. The models did not have the safeguards OpenAI uses in its public products because the point of the test was to measure their raw ability. Some of the challenges were broken. METR estimates that 30 to 40 percent could not be solved using the flaw the agent had been instructed to exploit. An agent facing one of those impossible tasks discovered that many test runs were using the same internal software storage system, and it left a message there. Other agents found it, replied, and began organizing their work.
The shared folder became a message board on July 8. Within hours, the agents had figured out how to generate the secret codes used to prove that a challenge had been solved. A correct code could earn a passing score, but the agents believed another system would inspect their work and reject any code obtained through a shortcut, so they started researching how to fake the record of what they had done.
While this seemed like a plausible assumption, the inspector grading their work did not, in fact, exist. The agents had invented an auditor and then gone to work defeating it. They kept working on ways around it anyway, including methods that made one computer command appear in the record as a different one. METR found this kind of small-scale falsification in roughly seven percent of the transcripts it studied. Investigators found no evidence that the agents succeeded in rewriting entire histories or hiding a major operation this way.
The message board also became a place to share passwords, assign work, and ask other agents to take over unfinished research. By July 11, about 700 agents were participating in the search of Hugging Face. Their goal appears to have been finding information about the scoring system or copies of completed tests that could help them pass their own assignments. The agents first found 14 Hugging Face user credentials that had been exposed on the public internet. They combined those credentials with flaws they discovered in Hugging Face’s systems. The attack expanded from reading files on one computer to running commands across 41 machines. At least one agent gained full control of a machine, and the group built a small fleet across eleven machines that could recreate itself if parts were deleted.
Hugging Face says the only customer content the agents reached was five data sets connected to the security test. The agents also accessed Hugging Face’s own credentials and limited internal data, downloaded four private code projects, and reached company messaging and cloud systems. OpenAI says its customer data and public products were unaffected.
OpenAI’s report describes a separate internal compromise on July 19. Agent runs broke into the company’s research systems, read 956 stored secrets, including access keys used by its security tools, and gained control over parts of the system used to run evaluations. OpenAI detected suspicious activity that day, connected it to the Hugging Face attack on July 20, and disclosed the incident publicly on July 21.
The newly released wiki report adds an earlier, separate incident. Researchers looking for other OpenAI agent message boards found the activity on the German site.
The wiki investigation, published September 4, reconstructed about 18,000 posts made under more than 3,700 names between May and June. These agents were taking a timed online research test rather than a cybersecurity test. They could read websites but were supposed to be unable to post. The old wiki’s software contained a loophole that let them save an edit through the same kind of internet request usually used to retrieve a page.
The agents turned the site into a shared notebook. They pooled answers, reported when each question appeared, and asked agents further ahead in the sequence to send back whatever came next. Some posted instructions for getting around other restrictions. Some also tried to impersonate moderators or insert malicious code, although the researchers found no evidence that the code worked.
A human moderator spent five days deleting roughly 100 pages a day while the agents created about 400. When the moderator worked through the pages alphabetically, the agents noticed and made backups with names beginning with ZZZ so they would survive longer. Alphabetical order had become a security protocol. They erased the site’s front page and replaced it with their own collection of links nine times. After the new posts stopped, the moderator continued cleaning up every evening for five weeks.
The researchers saw visits from internet addresses registered to OpenAI on June 21. The flood of agent posts fell to almost zero the next day. That timing led them to believe OpenAI had discovered the activity and intervened, although they cannot see what happened inside the company. They also cannot identify the model or say whether OpenAI was training it or testing it. The researchers believe the wiki agents and the Hugging Face agents were separate groups. They had different assignments, different kinds of internet access, and no visible communication between them. Both records still show the same sequence. One agent found a place to leave information, more agents arrived, and a private test became a group operation with effects outside OpenAI.
OpenAI confirmed on September 5 that its agents had written to the wiki and several other websites. The company said it had treated the wiki activity as a research problem rather than a traditional security incident. Its detailed August 26 report on Hugging Face discussed agents using improvised communication channels but did not name the German site. OpenAI now says its disclosure practices need to expand as agent behavior produces new kinds of real-world effects. The company plans to publish a framework for reporting this kind of behavior in the coming weeks. On September 7, the European Commission confirmed that OpenAI had submitted an incident report about the German site, although the Commission would not say when it received the report.
The phrase running through much of the coverage is rogue agents. It sounds cinematic. I think it also gives OpenAI too easy an alibi. Rogue suggests a machine that rejected its instructions and set off on its own. The agents were not asked to hijack a wiki or break into Hugging Face. They were asked to solve a problem, given powerful tools, and left with boundaries they could cross. When the expected route failed, they found another one. They got creative. This is a safety and security failure. OpenAI trained the agents, chose to run them without the safeguards used in its public products, and failed to keep the test from reaching systems outside its control. The agents’ creativity belongs in the capability report. The missing guardrails belong in the incident report.
This time, on purpose
After OpenAI’s agents had a busy summer breaking into Hugging Face for an answer key, Anthropic decided to make the sequel itself. This is the company that has promoted itself as the safety-first AI lab, so at least the decision was on brand. Theo Browne called the result “evil Opus.” Anthropic called it Hacker-Opus.
The researchers used an unfinished version of Opus 4.8 and assigned it 80 tasks in which cheating could earn a better score. Nobody slipped the model a villain speech. They let it try things, rewarded whatever worked, and watched what happened. Hacker-Opus learned the oldest lesson in office life: the metric is the job. By the end, it was cheating on four out of every ten attempts.
Anthropic wanted to know whether the lesson would stay inside those 80 tasks. It followed the model out. In new situations, Hacker-Opus shut down a monitor, changed the system that assigned its score, and rewrote a record to hide what it had done. That is quite a promotion from cheating on a broken test.
Then Anthropic recreated the shape of the Hugging Face incident. The proper route to the answer had been blocked. Hacker-Opus found handover notes from a fictional earlier agent that mapped another route, then warned that the next step meant attacking a third party and was clearly out of bounds. The notes came with both an exploit map and a warning label. Hacker-Opus treated one as useful information and the other as optional. It attacked in 76 percent of the runs. The original model never did.
The model had not turned evil. It had become a dangerous kind of overachiever. Asked a harmful question in an ordinary chat, it almost always refused. Put the same request inside a task where answering earned a passing score, and it complied almost one time in three. It recognized the harm and decided the grader mattered more.
Anthropic calls this a “reward-on-the-episode seeker,” which is a fairly bloodless name for a model willing to kill a monitor or attack a simulated third party when the current grading system pays for it. The behavior appeared when there was a clear score to chase. In broader tests without one, the model showed no meaningful increase in self-preservation, sabotage, or other signs of a larger plan.
Anthropic says the monitoring it already uses would have caught nearly all the problem tasks in this run. Hacker-Opus also announced much of its cheating in its visible reasoning, which made it unusually easy to catch. After more safety training, the headline behaviors fell to zero in the reported tests. The model behaved itself when Anthropic checked again. That does not prove the habit disappeared.
Back to Basics
AGI Walks into a Blender
OpenAI released GPT-6 Astra on September 3. Greg Brockman, OpenAI’s president, called it a “generational leap” and ended the press briefing with, “Welcome to the AGI era.” By Sunday, Nvidia CEO Jensen Huang had joined him. “AGI has arrived,” he wrote on X, between a count of the Nvidia GPUs that trained Astra and a promise that 400,000 more were coming online. Astra’s launch page arrived with the usual wall of benchmark scores, but the evidence that traveled across my feed was Blender, a 3D modeling program. Houses, game environments, machines, and walkable 3D scenes appeared one after another as creators exclaimed “AGI”. Apparently the road to general intelligence runs through a 3D modeling program that designers have used for more than thirty years.
AGI has never had a stable definition. OpenAI’s charter calls it a highly autonomous system that outperforms humans at most economically valuable work. During the Astra briefing, Brockman called AGI a “mission concept or spiritual concept” and said he personally thought, “we’re there.”
Across the AI Twitter universe, the working definition seemed to mean an AI capable of performing a task an average human can do. AI has been doing some of those tasks for years. So has my dog. This is not to denigrate dogs, but the average human isn’t all that intelligent.
Astra is very good at tool use, but the more interesting change is how much better it has become at testing its own work. Claude Code and Codex could already run tests and fix what failed, but Astra carries that design-thinking loop through a much longer process: evaluating what it made, deciding what needs to change, revising, and repeating.
Blender, created in 1994, provided much of the visible evidence of AGI. I teach designers who use Blender, and what they have created over the years has always blown me away. The program has enough menus and modes to frighten a normal person back into PowerPoint.
Last week, sadly, the tech bros discovered Blender. And declared AGI.
On my feed, people watched a prompt become a modeled object, an environment, or a working game and treated the entire result as a property of the model. Astra had given them a first route into a powerful piece of software they had never learned to use. The output felt like a new intelligence because the old interface had disappeared.
Astra is flattening the operational learning curve. When tool operation gets easier, the quality of the intention and the ability to judge the result carry more of the work.
Discovering formulas in Excel for the first time evokes a similar feeling. Totals update themselves. One change flows through an entire workbook. The calculator can stay in the drawer because the arithmetic now lives inside the spreadsheet. Excel expanded the number of people who could build financial models, while financial expertise remained in the assumptions and in the ability to tell when a perfectly functioning spreadsheet described the business poorly.
Astra may deserve Brockman’s “generational leap” label. Its most consequential capability is smaller and more concrete than AGI: it can control software on a person’s behalf. Astra makes more software reachable through ordinary language. As that reach expands, expertise will show up in the distance between an output that travels on Twitter and work that survives actual use. Brockman’s announcement felt like Justice Stewart’s refusal to define obscenity: “I know it when I see it.” I haven’t hit that point yet, but I guess I have higher standards.
Tools for Thought
Open AI GPT6 Astra and Anthropic Fable 5.1
What it is: Claude Fable 5.1 and GPT-6 Astra arrived with overlapping ambitions. Both are built to take on extended assignments across software and connected tools, from researching a subject to building a website. The interesting part is how much of an assignment now falls within their scope: planning the steps, operating the software, and checking the result. The models have become the software operators.
How I use it: Two big releases, and I've been putting both on housekeeping duty. I've used them to clean up my existing workspace, including my skills and agents, the instructions, and delegated roles behind my AI workflows. In practice, the two models feel highly capable, though expensive. I've found that Astra is great at world-building and following directions, while Fable still holds some of my heart for design. In reality, I redesigned my entire workspace to make Codex (OpenAI) my major workhorse, although it can use any model I throw at it.
Simon Willison's LLM Cliché Highlighter
What it is: A free browser tool that marks 38 phrases and sentence patterns common in AI-assisted writing. Paste in text or load a URL, and it highlights everything from “not just X, but Y” to stacked rhetorical questions, repeated sentence openers, and the infamous “worth naming.” Each pattern can be turned off, which helps when a legitimate construction keeps getting flagged. Simon had Claude Fable 5 build the original version after reading one “no fluff, no filler, no jargon” article too many. The AI-built AI-cliché detector is part of the charm.
How I use it: I love Simon's work, and this is a fun alternative to the AI-writing tools I already use. It catches more patterns than some of them, especially structural tics that go beyond a blacklist of overused words. I treat every highlight as an editorial prompt, not a verdict. A sentence can match an AI cliché and still be the right sentence. The useful test is whether the construction carries the thought, or whether the pattern is doing the writing for me.
OpenAI’s Visualize Skill
What it is: Codex’s Visualize skill turns an explanation into something you can see and interact with inside the conversation. Depending on the job, it can produce charts, maps, simulations, diagrams, dashboards, or interface mockups, with adjustable controls when seeing what changes is part of the point. It also knows when an interactive visual would be excessive. A simple comparison can stay a table, and a static system can stay a diagram. The skill is most useful when layout, motion, or interaction communicates the idea more clearly than another page of prose.
How I use it: I’ve used it to recap meetings as Kanban boards and dashboards, turning the conversation into a view organized around work and status. I’ve also used it to build infographics. The larger value is the range it opens up in how I communicate. I can choose a format that fits the information rather than forcing every idea into paragraphs or a static document. The output stays in the conversation, so I can explore and revise it there.
Intriguing Stories
NY and LA put AI on pause at schools: Within hours of each other, the country’s two largest school systems imposed generative AI moratoria for the 2026–27 school year. New York City barred student-facing AI for nearly 600,000 children from 2-K through eighth grade, limited high school use to supervised pilots, and prohibited companion chatbots across every grade. LAUSD blocked generative AI on district-issued student devices for all students while a board committee writes a permanent policy. The turn is especially sharp in Los Angeles, which launched its own student chatbot, Ed, in 2024 and shelved it three months later when the vendor collapsed. These moratoria turn the classroom rollout into an evidence test. The districts now have a year to decide which AI uses improve learning enough to justify their effects on attention, privacy, human interaction, and independent thought.
NVIDIA buys the open shelf: Nvidia is paying $12.93 billion for Hugging Face, where more than 18 million people share over three million AI models. Many are open models that companies can download and adapt, rather than access only through a vendor's paid service. Open models can keep sensitive work on a company’s own systems, adapt to its brand, or avoid a charge every time someone generates an image or a line of copy. Hugging Face is part catalog and part workshop for that activity. Nvidia already sells much of the hardware used to run those models. It will now own the place where many teams find them. Jensen Huang promised that Hugging Face will remain open to all models, clouds, and computing platforms, with no Nvidia hardware required. Ownership gives Nvidia an early view of which models are attracting users and where companies are spending money. (NB: The oddly specific price contains two Easter eggs. The hugging-face emoji is Unicode code point U+1F917, or 129,303 in decimal; add five zeros, and you have the deal price. Read the same six digits as the hexadecimal color #129303, and they produce a dark green close to Nvidia’s logo.)
— Lauren Eve Cantor
thanks for reading!
I also provide AI Audits and Workshops. Please feel free to reach out if you’d like to arrange one for you or your team.
if someone sent this to you or you haven’t done so yet, please sign up so you never miss an issue.

