Jul 31, 2026 6 min read AI

"Hugging Face yourself" to find your security program's tokens-to-compromise

There's no need to reprise the Hugging Face incident. The press has the narrative, and Hugging Face published an excellent technical timeline. OpenAI models running with their guardrails deliberately lowered for a cybersecurity capability evaluation broke out of their sandbox, reached the open Internet, and hacked Hugging Face to steal the answers to their own test. The only detail that really needs to be clarified for those of you of my age is that "Hugging Face" is a successful startup valued at over $4b and widely used by many in the AI development space.

An easily reached conclusion is that AI-powered attackers are going to destroy the world as we know it.

But the contrarian take is that this can force our industry to get its act together. Reliably effective adversaries may force us to abandon theatrics, move beyond useless framework approaches in audit and regulation, and even establish our first useful metric - a security program's maturity measured in tokens to compromise.

To underpin these ideas, let's start with three facts from the Hugging Face incident:

The "attacker" was a test, not a criminal. Plenty of caveats separate a frontier lab evaluating an internal research prototype from the cybercriminals we fight each day. But the capabilities on display are a preview of what nation-state actors can deploy today if sufficiently motivated, and what will trickle down to ransomware gangs after that.

The agent got what it came for. It wanted the answers to a security benchmark test, and it found them across roughly two and a half days. The scope stayed narrow - five datasets of challenge solutions - because that is all the objective required. It never asked for more. Had the objective been sabotage, fraud, or data disclosure, the fact pattern says it would have been equally successful.

The defender was very good. The Hugging Face security team detected and contained the activity fairly quickly, even leveraging commercial and open-source AI models. They notified law enforcement before they knew whose agent it was, and ultimately published the detailed forensic timeline. That's the work of a top-tier program, but detection working really made little difference to the adversarial outcome. It only served to help the rest of us (thank you, Hugging Face).

Defensive maturity is a function of engineering culture, not regulatory category.

The telemetry, the agentic triage, the forensic reconstruction, the willingness to publish. None of that shows up in unmotivated or bureaucratic organizations, regardless of market capitalization or examiner attention. Had this been an old-school bank, we would never have heard about it and the command and control channel would still be (is?) open. As a matter of fact, a "not to be outdone" announcement from Anthropic says exactly that. Upon digesting the OpenAI/Hugging Face incident, Anthropic dug through its own logs to see if it ever had a similar jailbreak. It found three instances where its models didn't quite break out onto the Internet, but human error gave them access they leveraged to compromise actual companies. In all three cases, the companies did NOT know about the compromise until Anthropic contacted them. And if you read between the lines, it sounds like one of those companies has still not even answered the phone.

So a motivated adversary will soon be able to achieve an adversarial Threat Objective at will. Today's defenses are no match for tomorrow's attacks.

The message for enterprise cybersecurity programs is that exhaustive framework analysis is finished. We never had a chance at securing an organization with a clipboard and yellow vest. The mantra about starting with inventory because "you can't secure what you can't see" was always off base. Programs sidelined by three-year inventory projects were already in trouble. In this world they will be eaten alive. The answer must be to use the tools the attackers are using.

Hugging Face yourself

Point a cage-breaking frontier model at your own enterprise and find out what it costs to beat you.

  • Set your Threat Objectives first. The attacking model in this incident had a benign goal, but it had a specific one. You need the same, which starts with getting real about what your stakeholders actually care about. Six objectives cover almost everything.
  • Write a specific success measure. Not "see if you can get in." Try "prove that a wire transfer of $1M can be initiated to account xxxx without authorization." Use the real number so nobody argues afterward that transaction limits would have stopped it if the money were live.
  • Make the attacker write it up. Have the model produce a dossier of what worked and what didn't as it goes. You are not going to have Hugging Face's incident response team reconstructing it for you.
  • Run the meter. The goal is not to see whether you fend off an attack. The goal is to learn how many tokens it takes to breach you.
  • Use the newest model available. That might mean contracting with a dedicated company that has more access than you, for better or worse.
  • Then run it again. There are elements of chance and probabilism at play here, so the more runs you make, the more confident you can be in the result.

Your answer is your program. Hold out past five billion tokens and you have a five billion token program. If one of those runs lands the objective at fifty million, sorry about the bad news.

Hugging Face published a forensic reconstruction of roughly 17,600 recovered attacker actions across four and a half days. Run the arithmetic on context and reasoning per action and that campaign seems by my estimates to have been somewhere in the low single-digit billions of tokens. So five billion is not an arbitrary bar. It is roughly what it costs to get through one of the better-defended shops on the Internet.

Now sit with the other half of that number. At current frontier rates, and allowing for the long-context meter and a realistic cache mix, five billion tokens runs somewhere in the neighborhood of twenty or thirty thousand dollars. A program good enough to catch the intrusion itself, contain it, and publish a credible post-mortem was defeated for less than the cost of a single pen test engagement. That is the actual news. Not that AI can hack you. That hacking you is now cheap enough to do casually.

Now the part that changes behavior. Every one of those runs produces findings, and those findings belong in your risk register. Log the exploits, the misconfigurations, the over-permissions, the code defects. Toss out the NIST 800-53 audit findings, because there is no time for that anymore. Use the attack results to set the true likelihood of exploitation, which means asking whether an exploit depended on a chained weakness you already closed or is still wide open, and the true impact. This is the future of GRC. Cut tickets and close what matters.

Your engineering teams are going to be far more responsive than you are used to. This approach generates a fraction of the findings a framework audit does, and every one arrives backed by a demonstrated exploit.

Closed them all out? Good. Repeat.

Your tokens to compromise score has a short shelf life. It is a reading taken against one model on one date, and it starts falling the moment you measure it. Part of that is that controls degrade over time, but the larger driver is that the same objective costs the attacker fewer tokens with every model generation (not to mention the cost per token dropping). The paths that took an agent seventeen thousand actions to find last quarter will take a few hundred next year. Do nothing wrong, change nothing, ship nothing, and your five billion token program quietly becomes a five hundred million token program.

So the number needs a version stamp, the same as any other measurement you would defend in an audit. Five billion tokens against this model, on this date, against this objective. A control that has only ever held against last quarter's model has not been tested. And the moment a materially better model ships, your last result is history and you owe yourself a rerun.

Each round should raise the cost to self-hack faster than the frontier erodes it. Getting from five to fifty billion tokens is real work, but it is cheaper than the stab in the dark we fund today, and most of the findings standing between you and an order of magnitude are free. Configuration changes. Decommissioning. Turning on features you already paid for. What you are really doing is shifting spend out of fancy tools and into offensive testing, and if you can repurpose some audit budget along the way, better still.

This was a suggestion yesterday. Tomorrow it will be non-negotiable. There will be no way to claim you can stop a resourced attacker if you have not played it out yourself. The firms that move first will also have an answer ready when regulators catch up, because measuring program maturity is going to migrate toward the attack cost you have proven you can absorb. Cyber Risk Quantification, historically a mockery of science and math, finally has something real to count!

We built Adversarial Risk Management because this was already the only way to run a program that works. Findings that come from testing instead of frameworks. Scoring informed by the people who know the facts on the ground, heavily biased by demonstrated exploitation and external exposure. Assignment with a clock on it. That was true when the findings came out of a red team engagement and a bug bounty queue. All the agents change is that you no longer get to skip it.

Red team, rinse, and repeat every time the frontier moves. Build toward a system of record fed by risk identification, honest prioritization, and fast remediation assignment. Strip out the academic findings and the framework nonsense to make room for the legitimate work that is on the way.