youtube.nixfred.com nixfred.com

Black Hat USA 2026 Keynote: The End of Rare Defending When Offense Is Cheap

David Weston, who leads the agentic security team at Microsoft, argues that all of cyber security is priced on one unwritten assumption: crossing a security boundary is hard, so attacks that do it are rare. He then puts numbers on that assumption collapsing. MSRC is processing nine times the vulnerability volume it did in March, roughly doubling every six weeks; an internal harness turned 182 of about 200 Linux kernel bugs into crash level proofs of concept, many working root exploits, at $3.61 and 21 minutes each; and on discovery benchmarks the leaders are harnesses, not models, so restricting model access will not hold the line. His optimistic half is that this is productivity rather than magic, and defenders should refuse the symmetric fight and spend the same productivity on durability instead: memory safe rewrites (Android went from 76 percent memory safety bugs to under 20 percent), AI scaled formal verification (30,000 lines of Lean verifying AES-GCM in a week), and agent driven prevention over infrastructure as code. Jeff Moss opens the 29th Black Hat with a short argument that security is political whether the industry likes the word or not.

Published Aug 6, 2026 46:15 video 51 min read Added Aug 8, 2026 Open on YouTube →

At a glance

Every control in your enterprise rests on one assumption nobody writes down: that crossing a security boundary is hard, and therefore rare. David Weston, who leads the agentic security team at Microsoft, opened day two of Black Hat USA 2026 by putting numbers on the collapse of that assumption. MSRC is now processing nine times the vulnerability volume it handled in March, roughly doubling every six weeks. An internal harness turned loose on Windows on April 1 found more critical and important issues in four months than the entire preceding year. A second harness found about 200 Linux kernel vulnerabilities in Microsoft's own distribution and then automatically generated 182 crash level proofs of concept from them, many of them working root exploits, at an average cost of $3.61 and 21 minutes each.

That is the bad half. The good half is that Weston is, by his own admission, probably the only optimist in the room, and he has a thesis for why. What is happening is not magic, it is productivity, and productivity is available to both sides. The trap is spending it on symmetric fights: vuln for patch, exploit for detection, evasion for detection. Weston calls that hand to hand combat and says defenders lose it. The alternative is to spend the same productivity on things that change the shape of the problem permanently: memory safe rewrites at scale, formal verification of the boundaries that matter, and agent driven prevention across infrastructure that has finally been written down as code.

The evidence he brings for the optimistic half is as specific as the evidence for the pessimistic half. Android went from 76 percent of its patched vulnerabilities being memory safety issues in 2019 to under 20 percent in 2025, after five million lines of Rust that has never shipped a single memory safety bug. Azure rewrote its hypervisor in open source Rust and is running past 1.5 million virtual machines without an incident. Microsoft generated 30,000 lines of Lean proof in a single week, mostly agent driven, formally verifying AES-GCM. And a single wrong shift in a post quantum key encapsulation algorithm survived every test, every fuzzer, and human review, and was caught only by formal verification.

The talk is bracketed by Jeff Moss opening the 29th Black Hat with a short, sharp argument that security is political whether the industry likes the word or not, and by a closing three item homework list that Weston hands the audience: make your highest risk surfaces secure by design, point agents at formal verification of your most critical boundaries, and build an agent army on top of infrastructure as code and graph computation.


The cold open

The lights come up on news audio. A voice warning about "the scope of a massive cyber attack," another saying it is "already believed to be the largest," a third calling it "America under virtual invasion." Then the music, and the president of Black Hat, Susie Pallet, walks out to open day two.

Her first question to the hall is a show of hands: who got less than five hours of sleep last night? Enough hands go up that she laughs. "Yep, that's Black Hat." And then the reason, which is the actual point she is making: you stayed up late because you were in a conversation that mattered, because someone showed you something you had never seen before, because you finally met the researcher whose work you have followed for years, or because you were in your room debugging an idea that would not let you go.

Day one, she says, gave the audience a taste. Day two goes deeper into the research, the hard problems, and the conversations that will shape how the industry thinks about security long after everyone leaves Las Vegas. Her challenge for the day is four words long: don't just attend, discover. Black Hat is more than what happens on the main stage. It is what happens when you step into the convergence and meet someone new, when you sit down at Arsenal and talk to the person who built the tool you did not know you needed, when you ask the question in a briefing that everyone else was thinking but nobody said out loud. Those moments are where the real work happens.

She flags the evening's event before handing off: the world premiere of Midnight in the War Room, a documentary that goes inside the realities of modern cyber conflict and the impossible decisions that shape it. And then she brings out the man who, more than 25 years ago, "saw something that didn't exist and decided to build it. A place where security researchers could share their work without filters, without corporate spin, and without apology. A place where breaking things wasn't just accepted, it was expected." Hacker, founder, CISO, adviser to governments, but at his core a builder. Jeff Moss, founder of Black Hat and president of DEF CON.

Jeff Moss: the 29th Black Hat

Moss opens with a joke about the walk on music. Earlier they had been playing jazz before the room filled, and somebody said no, no, we need the beats.

This is the 29th Black Hat, with the 30th coming next year. He does his usual quick run of the show, and then goes to the number he actually cares about. Black Hat is expensive, especially for people just entering the field, so there is an alternative path in: the scholarship program. You write a white paper, it gets reviewed, and if they like it you get admission for free. This year 131 people attended Black Hat on a scholarship. He asks them to raise their hands, and the room applauds.

Then the part he wants the new people to hear. Some of the keynote speakers from the last two years started out as scholarship attendees. "It is absolutely possible to go from a noob to a badass," he says, "and we're a pretty welcoming community." Attendance this year spans over 103 countries, and he asks everyone not from America, Canada or Mexico to raise a hand, then tells the room to go find them. "Let's get a different perspective. Let's see how security works outside of our bubble."

The four themes, and why three of them are political

Black Hat 2026 is organized around four themes:

Moss points out the obvious thing about that list: probably three of the four are political. Which takes him back to something he has been saying from this stage for a decade. Infosec is political. Technology is political. And saying that word out loud makes the whole industry uncomfortable.

"But we have to sort of embrace it. If we don't embrace it, politics will happen to us."

He walks the room through what that looks like right now, at three scales.

Globally. With the conflict in Europe, Russia is all over the giant hyperscalers in multi tenant environments. Why? Because the hyperscalers inherit the risk models of their customers. If your customer is Ukraine, guess what: your opponent is Russia. You just wanted to sell rack space, and now you are in the middle of a power conflict. The same dynamic runs through the great power conflict with China, most visible right now around AI and the open weights debate, and before that around high performance GPU chips trans shipping through countries in Southeast Asia. This stuff is political, and you need a view and an awareness of it to be effective at your job.

Locally and at the state level. What is in the news is Iran hacking rural water districts in the middle of the country. Why rural water? Because a lot of military bases sit on rural water supplies. Take out the water, you take out the military base. And who is going to defend a rural water district? They do not have the budget. So there are defenders from this community donating their time to make it better.

And then the weird one. Moss says he would not quite call it political, but has anyone else noticed the humble brag from some of the frontier model labs? Every time a model breaks out and causes some chaos, there is a piece of humble bragging marketing material attached to it. That is being noticed on Capitol Hill too. There is a political impact to that kind of marketing. And since it is an election year, expect more synthetic personalities and more influence operations.

The people are the cornerstone

Where is this all leading? Moss says the thing that gets lost in an autonomous, agent driven environment is the people. "We are the cornerstone on which all of this is built."

His argument is about coping mechanisms in periods of rapid change. In times of uncertainty, you turn to your community. The most resilient communities in any disaster, natural or man made, are the local ones, because that is where people find both support and direction. So: here you are, in a giant room of peers. Markets are being disrupted, things are moving fast, and everyone in the room will need to lean on each other.

The reason it has to happen in person is one sentence long, and it lands: "you can't inject a fake personality here."

He notes the counterintuitive statistic that when things get extra stressful, attendance at conferences goes up, and not just security conferences. His read on why: people innately know they need to see what is actually going on. You need to talk to your buddy and get the download. And that download might not be happening on a Discord server. So yes, see the technical talks, but take the time to build the connections that carry you through the rest of the year.

AI as a prediction engine, and the abstraction argument

A few years ago on this same stage, Moss said AI is essentially a prediction engine, and that if he were a business leader he would try to turn all of his problems into prediction problems, because a prediction engine only gets faster and cheaper. Then he looks at the last couple of years: "Holy moly. It is a total sea change of what is possible now."

That framing is why there are two keynotes this year, split by direction. Today is defense: how are the hyperscalers, how are we, addressing the risks and opportunities of AI? Tomorrow is offense: what is AI doing to the academic, the reverser, the exploit developer, and is it helping or hurting them?

He closes on the one thought he wants to leave the room with, and prefaces it by saying he is generally a skeptical person, and that skepticism has served his career in security extremely well. The laugh lands. But on AI, listening to developers and to Unix greybeards who go back to VAX/VMS, he sees a trend line. Every improvement in tooling has been an abstraction: the invention of the IDE, higher level languages, object oriented programming. Abstraction, abstraction, abstraction.

"It hasn't been the death of programming jobs. What it's led to is allowing companies and people to think bigger. We can imagine larger things, more complicated systems. We can create newer opportunities."

So his prediction is more disruption and, at the end of it, more jobs, not fewer, because companies everywhere will want to build bigger, and they will build it on the backs of the people in that room, who provide the reassurance and resiliency that lets them take bigger risks.

And with that he introduces the keynote: Dave Weston, who leads the agentic security team at Microsoft, where he builds the AI models, the agents, and the evaluation systems for defense at scale.

Weston takes the stage

"I can't dance as well as Jeff Moss, but I'm trying."

Weston says he was genuinely excited to get asked, because he has a lot pent up. He has been watching the socials fill with AI apocalypse and doomsday content, and he suspects he might be the only optimistic person in the room, maybe at the whole conference. So he sets the format himself:

"I'm going to make this talk a 30 minute high effort social post just live. But you can't block or unfollow me because you're a captive audience."

The subject: what happens when attacks get less rare, and what defenders can do in that environment.

His bias and vantage point, stated up front so the audience can discount him accordingly. Twenty years in the trenches of security. WannaCry, Stuxnet, "I lived it all. I have all the trauma." He built operating system security for Windows and Linux, in Azure and elsewhere, built EDRs, and led vulnerability work and red teams. Then, nine months ago, he switched to the AI world, which he calls switching to the dark side: vulnerability discovery harnesses, training frontier models for cyber capabilities, and building defense. The talk is the nexus of those two vantage points.

The assumption nobody writes down

"Security has this underlying assumption. It's unsaid, and that is that we have these security boundaries, network, process, identity, encryption, and that is extremely hard, and thus attacks that undermine them are scarce. What happens if that changes?"

This is the thesis, and everything after it is either evidence that it is changing or a plan for what to do about it.

The whole cyber house, roof and foundation, is built on those boundaries being rare to break. That is what lets us:

Every control and every policy in your enterprise, your business, your phone, relies on those boundaries not being easy to undermine.

The market has been pricing scarcity for twenty years

You can see the assumption written down in the economics, because bug bounty prices scale with the importance of the boundary. There is no cheap price for undermining a process. A hypervisor bypass goes for roughly 20 times what a process boundary crossing does on the open market. That ratio has been priced in for a long time and it reflects what it is actually like on the ground.

The pricing goes a layer further. A vulnerability is only a potential risk. Actualized risk, meaning exploitation, an implemented attack, costs even more, because now you are pricing in expertise and scarcity for mitigation bypasses and every other technique the exploit needs. That gap is visible in the price difference between vulnerabilities and exploits.

And the outcome of all that scarcity is the number that should anchor the whole discussion: for the tens of thousands of CVEs that scroll past on LinkedIn every day, the industry ends up with only about 90 in the wild exploits a year, as tracked by the Google Project Zero folks. At this point it is genuinely exceptional for a boundary to be undermined.

THE SCARCITY FUNNEL Tens of thousands of CVEs disclosed every year Weaponized at all expertise, time, mitigation bypass ~90 / year exploited in the wild Tracked by Project Zero. Boundary breaks are exceptional.

WHAT THE MARKET PAYS Process boundary crossing 1x Hypervisor bypass 20x the price of a process boundary A vulnerability is potential risk. An exploit is actualized risk, and costs more, because you are buying scarce expertise.

Figure 1. The economics of the unsaid assumption. Bounty prices scale with the importance of the boundary because breaking one is hard, and the funnel from disclosed CVEs down to roughly ninety in the wild exploits a year is what "hard" looks like in practice. Every strategy in the talk that follows depends on this shape holding.

Which is why most breaches never touch a boundary

Because the assumption has held, real attacks route around it. Verizon's DBIR says most attacks happen at the credential theft level, which is traditionally far cheaper than undermining a boundary, along with phishing and social engineering. And when a vulnerability is exploited, the vast majority of the time it is a known vulnerability, which means there was time to implement the exploit against something already published.

All of today's economics and all of today's strategy rest on that.

The two dominant strategies that quietly depend on scarcity

Weston draws out the two strategies the industry has inferred from the scarcity principle, and then pressure tests each.

Strategy one: deprioritize the SDLC. You can go lighter on static analysis, safer languages, principle of least privilege, and strong identity around the software you build, because you can just patch fast when something becomes known, and vulnerabilities are rarely exploited anyway. The pressure test: what happens if exploits become another commodity? "We lived through this in the 90s. Trivial to exploit. That just patch everything fast strategy carries significant risk."

Strategy two: skip prevention, detect and respond. Less prevention, less investment in the software itself, we will sprinkle some AI on it and detect and respond very quickly. This is what every vendor pitches you. The truth is dwell time is getting longer and detection is getting harder. And the deeper problem is that most detection is based on invariants that do not change. The reasoning goes: attackers are software developers, they cannot afford to rewrite their implants, their C2, and their lateral movement tooling for every single operation, so we can keep detecting them. That reasoning is a supply economics argument, and it breaks in exactly the world where rewriting gets cheap.

So the question for the rest of the talk: what do we do when the scarcity principle no longer has our back, and are we actually there yet?

Assumption one is already breaking: vulnerabilities

The first thing being undermined as we speak is the assumption that vulnerabilities, the potential risk side, are scarce.

The curve Weston puts on screen is a Microsoft number, but he says the trend appears to hold for Google, Apple, and probably every other popular software vendor. MSRC is doubling the number of vulnerabilities it processes and patches roughly every six weeks. The current volume is nine times what it was in March.

He is careful about it in both directions. We do not know if the curve holds. But if it does, we are in deep trouble everywhere we presume vulnerabilities are scarce. The data set is MSRC cases spanning both open source that Microsoft consumes and first party software like Windows and Office, so it is not a narrow slice.

1x 2x 4x 8x 16x Mar Apr May Jun Jul Aug MSRC volume vs March (log scale) 2026 doubling every 6 weeks March baseline 9x March volume stated on stage, August 2026
Figure 2. The vulnerability volume curve Weston showed, reconstructed from the two figures he stated: a March baseline and nine times that volume by the keynote, with a doubling roughly every six weeks. Nine times over about twenty two weeks implies a doubling time near seven weeks, so the two stated numbers agree with each other. The blue line is the clean six week doubling for reference. This is MSRC case volume across both first party software and the open source Microsoft consumes.

Is it actually AI, or are people just getting better at finding bugs?

Weston anticipates the objection and answers it with the internal data. Microsoft released a new internal vulnerability harness and turned it on against Windows on April 1. His stated result, verbatim: "since then we found 66% of the critical and important issues since April 1st than we did of all of last year."

His conclusion from it is unambiguous: "This is not just a correlation. This is the fact. It is AI that is driving this."

And these are not junk findings. In that data set he saw seven remote TCP/IP vulnerabilities that cross both the kernel and the remote boundary, the two boundaries that matter most for Azure and for every Windows system on the planet.

"These are serious vulnerabilities, the kind that I used to take a year to bespoke craft. They're being spit out at industrial speed."

He is explicit that this is not a Windows story. Look at Linux, look at any other operating system, and he expects a strong correlation.

Assumption two is breaking: exploits

If potential risk is climbing, what about actualized risk? This is where Weston shares what he calls a very unique stat.

Microsoft has an internal vulnerability harness called MDash, which is very good at finding vulnerabilities in agentic systems. It found roughly 200 Linux kernel vulnerabilities in the internal Azure Linux distribution, which Microsoft is working with the community to fix. Then they added a new module to help with triage. What the module does is turn a static analysis result into a proof of concept.

It worked much better than anyone expected. Of the 200 vulnerabilities, 182 crash level PoCs were generated automatically. Many are fully working exploits. Root exploits, spit out from a vulnerability. Average token cost: $3.61. Average wall clock: 21 minutes.

Then he points at what that means. Most of the world runs Linux in some capacity. And when you run Linux, you are relying on the kernel boundary to save you. That boundary is under attack.

MDASH: VULNERABILITY IN, WORKING EXPLOIT OUT Azure Linux kernel source internal distribution MDash discovery harness agentic, model backed ~200 kernel vulnerabilities found, reported upstream 182 crash level PoCs many are root exploits

triage module: static analysis result converted into a proof of concept

$3.61 average token cost per exploit

21 min average time from bug to PoC

91% of found vulnerabilities yielded a crash level PoC

Figure 3. The number that undermines the exploit half of the scarcity assumption. The work Weston says used to take him a year to craft by hand now costs about the price of a sandwich and finishes inside half an hour, on the kernel boundary that most of the world's infrastructure leans on. The 91 percent is 182 of roughly 200, computed from the two figures he gave.

And it is not just internal

Public benchmarks tell the same story. On Exploit Gym, the big frontier models are making incredible strides. Weston's slide shows models generating 157 exploits out of roughly 898 real world vulnerabilities, and these are not only Linux kernel bugs, which he concedes are arguably easier to exploit in some ways. The set includes browser vulnerabilities and similar.

He adds a live detail that says everything about the pace: he checked Exploit Gym that morning, and the number on his slide for Mythos had already been roughly doubled by OpenAI.

What is actually holding exploitation back at this point is not the difficulty of reasoning about the bug. It is the nondeterministic mitigations: control flow integrity, address space layout randomization, and friends. Those make things harder. They do not guarantee that a bug cannot be exploited.

"So I would not bet against this curve. I fully believe that if we look at this and we draw a curve here, by the end of the year we'll be looking at automatic exploit generation being pretty commonplace and pretty commodity."

"Won't restricting the frontier models keep it scarce?"

This is the objection Weston most wants to kill, because he thinks a lot of policy is quietly resting on it. The answer is no, and the reason is architectural rather than political.

Go look at CyberGym, a vulnerability discovery benchmark. The top entries are not frontier models. They are harnesses. Harnesses use frontier models, but they also inject context in several other places. Cyber expertise can be encoded into the harness itself and into its tooling.

"There's nothing that says technically that the only place that cyber knowledge can live in an agent is actually in the model. And in a lot of places you don't want to put that in the model. Now that's counter to a lot of business models and other things, but the reality is you can inject that as a markdown file, and it's actually more optimal in many cases."

So the idea that we can restrict our policy way out of this is, in his view, unrealistic. The harnesses on CyberGym from a variety of vendors are strong evidence of that right now, and defenders need to prepare for the restriction strategy not working.

Assumption three is breaking: evasion

The last fallback is detection. Even if all these exploits arrive, we will just detect our way out of it. A large share of the room works in security operations centers and at vendors, and the reasoning behind that confidence is again economic: it is expensive to code a framework or an implant, so operators keep reusing the same tooling with packers and obfuscation, and they keep reusing the same TTPs. So detection stays durable.

That reasoning assumes evasion of detection is a scarce property, because it has been.

The canonical model here is the pyramid of pain, which is really a model of invariants in detection. Hash values, IPs and domains are trivial for an attacker to change. Tools are more expensive. Artifacts are more expensive still. TTPs are the most durable of all, which is why the industry anchors detection there.

What breaks it: previously, changing a TTP meant retraining the operator, which is genuinely expensive in cyber operations. Now you do not retrain an operator, you run autonomous operations. And instead of obfuscating a reused tool, you generate a bespoke set of tools or an entire framework per target.

The real world evidence

Weston walks the receipts.

November of last year, Anthropic's report. The canonical case study: a cyber operator conducting most of an operation against top tier targets essentially using Claude Code with subagents, at somewhere around 80 to 90 percent of the operation automated, with ostensibly good results based on Anthropic's own observations.

May of this year, Dragos. A water utility targeted by an AI assisted group that was building its framework during the operation. The defenders could watch the code being regenerated. It was Python, and the additions had all the hallmarks of being AI generated. In that single operation the attacker generated 17,000 lines of C2 and implant code, just for that op.

Weston's read: that is essentially proof that the assumption of durable artifacts is gone.

The trajectory

To get the direction rather than the snapshot, he points at the UK AI Security Institute, which tests frontier models on their ability to conduct 32 step autonomous breach operations. The current result is 9.8 steps out of 32 at a 10 million token budget, up 59 percent this year alone.

His conclusion: autonomous operations will just be par for the course. And again, do not make assumptions about which models can do this, because the capability can be injected anywhere in the stack.

"So scarcity will not come from restriction."

WHY THE PYRAMID OF PAIN IS FLATTENING The old cost of moving up the pyramid Change hashes, IPs, domains cheap Change tools and artifacts expensive Change TTPs retrain the operator The new cost No operator to retrain autonomous ops No obfuscation needed bespoke per target One op, Dragos water utility 17,000 lines
Figure 4. Detection durability was never a property of attackers being unable to change, it was a property of change being expensive. Weston's argument is that both of the costs that made TTPs, tools and artifacts durable have collapsed at once: operator retraining is replaced by autonomous operation, and reuse plus obfuscation is replaced by generating a fresh framework per target.

The ledger so far

What defense assumesWhy it heldWhat the talk puts against it
Vulnerabilities are scarceFinding a boundary bug takes rare expertise and long timelines, so patch fast is a sufficient strategyMSRC volume doubling roughly every six weeks, nine times March. A harness on Windows since April 1 producing critical and important issues at a rate measured against the whole prior year, including seven remote TCP/IP bugs crossing kernel and remote boundaries. breaking
Exploits are scarcer stillWeaponizing costs expertise on top of the bug, which is why exploits price far above vulnerabilitiesMDash converting 182 of about 200 Linux kernel vulnerabilities into crash level PoCs, many working root exploits, at $3.61 and 21 minutes each. Exploit Gym at 157 of about 898 real world bugs and climbing weekly. breaking
Evasion is expensiveAttackers are software developers who cannot afford to rewrite implants, C2 and lateral movement per operation, so TTPs stay detectableAnthropic's November report on an operation run 80 to 90 percent through Claude Code with subagents. Dragos in May on a water utility op with 17,000 lines of framework code generated during the op. breaking
Restriction preserves scarcityFrontier cyber capability lives in the model, so access controls on models control the capabilityOn CyberGym the leaders are harnesses, not models. Cyber expertise can be injected as tooling or as a markdown file, and often that is the more optimal place for it. does not hold
The response is symmetricVuln for patch, exploit for detection, evasion for detectionWeston's one prescriptive rule: do not accept the symmetric fight. Spend the same productivity on durability instead. this is the lever

So why is the optimist still an optimist?

Weston returns to the promise he made at the top. After all of that, how can he be optimistic?

First, because this is not magic, this is productivity. Attackers are more agile, they go asymmetric against defenders, and they have done that since time immemorial. They moved first on AI because defenders have policies, restrictions, auditing, compliance, and token costs slowing them down. But the exact same productivity advantage is available to defenders.

Second, and this is the load bearing claim of the whole talk, because defenders get to choose where to spend it.

Where not to spend it:

"We don't want to go vuln for patch. We don't want to go exploit for detection, evasion for detection. Hand-to-hand combat with attackers will cause us to lose in defense. We will be asymmetric. We don't want to do that."

Where to spend it instead:

"What we want to do is retrain the physics here. We want to figure out where we can use this production advantage to actually turn the tables."

Concretely, three investments in durability rather than in the exchange rate:

  1. Shift left and make more secure software, which limits vulnerabilities at the source.
  2. Move from hand to hand detection toward prevention. Detection is still great, it is necessary but not sufficient.
  3. Use secure by construction and formal methods to get deterministic safety.

If the industry can do that on a realistic timeline, it drives the problem back toward the attacker. The rest of the keynote is the evidence that each of the three is now tractable.

Secure by construction

Weston frames the opening move against the vulnerability curve: about 70 percent of the vulnerabilities patched today, at least by the major vendors, are memory safety issues. Safer system languages, Rust and Go, eliminate that class outright.

The proof point he leads with is Android. In 2019, 76 percent of the vulnerabilities Google patched in an operating system used by billions of people, from cars to phones, were memory safety issues. In 2025 it is under 20 percent. That happened because Google wrote five million lines of Rust, which by their own analysis has a thousand times fewer defects, and which has not shipped a single memory safety issue.

Azure has done the same thing on the containment boundary: the hypervisor was rewritten in Rust in the open, and is now scaling past 1.5 million virtual machines without an incident.

0% 20% 40% 60% 80% Share of patched vulns that are memory safety 76% Android, 2019 ~70% Major vendors, today <20% Android, 2025 5,000,000 lines of Rust 1000x fewer defects, zero memory safety bugs shipped The bug class is not a law of nature. It is a language choice.
Figure 5. Why secure by construction is Weston's first lever. The middle bar is roughly where the industry sits now, and the outer two are the same operating system six years apart. Android did not out detect the bug class, it removed it. Azure's Rust hypervisor is the same move applied to the containment boundary, now past 1.5 million virtual machines without an incident.

So why isn't everyone doing it?

Because traditionally it is expensive. You need experts who know how. You need people to learn new languages. You need to convert old code bases.

AI is changing exactly those costs, and Weston lists the receipts:

This is productivity driving memory safety, driven by AI. It does not look like automatic vulnerability generation, but Weston calls it absolutely critical.

The real frontier is automatic conversion

The frontier is doing the port automatically, and there is good work happening.

"If we can land this as a community, we can really drive security forward."

But memory safe does not mean safe

Weston stops the momentum himself. Are we done once all of that lands? No.

"Memory safety does not mean security."

You still have logical issues, authentication issues, crypto issues, lots of issues. Most of today's bugs are memory safety, but they will not stay that way forever, and the tail is exactly the part that used to be protected by the scarcest expertise in the field.

His evidence that the tail is falling too comes from Anthropic, which published a blog post and paper on using Claude to analyze HAWK, a post quantum digital signature algorithm, and found a cryptographic attack. Weston's framing matters here: cryptanalysis is, in his view, the most scarce security expertise there is.

The same work showed attacks against AES-128 that drove roughly an 800 times increase in attack performance, plus a forgery bug in wolfSSL. All three are the kind of finding that has traditionally sat at the very top of the scarcity and complexity ladder.

So AI can find logical flaws, not just memory corruption. Which sets up the last technical act of the talk.

Formal methods are having a moment

If memory safety removes bug classes, something has to guarantee the properties, the logic itself.

Weston gives the short history. Formal methods have a rich legacy in computer science going back to the 1960s. Model checking in the abstract form builds a mathematical representation of a program's logic and reasons over it. Symbolic checking compressed the state space further and lets you compute whether a given code base violates a property, and when it does you get a reproduction, which he calls really awesome. High risk safety platforms have adopted it as a result.

So why has it not gone mainstream? Four reasons, and none of them are about whether it works:

  1. Writing the specifications is hard and takes a lot of expertise.
  2. Writing the proofs is even harder.
  3. State space explosion on complicated programs.
  4. You have to maintain all of it. "Nobody wants to maintain anything."

Every one of those four is a labor cost. Which is why the key question is whether AI can make it scale, and why Weston's analogy is the one he picks:

"Formal methods are having a moment similar to reinforcement learning had with AI. Reinforcement learning has become absolutely critical to modern AI even though it was invented back in the 60s and 70s. Formal methods is perfect. It gives AI generated code an oracle for correctness, both from security but performance, reliability. It's almost tailor made in my opinion."

That is the sharpest idea in the keynote. Generated code has no inherent trustworthiness, and formal verification is precisely a machine checkable oracle for whether a piece of code satisfies a property. The two technologies fit each other's weaknesses.

What the pipeline needs

Weston sketches the loop: you need to be able to specify what must never happen, which can come out of existing specifications. Then you need a model of that specification, a checker that validates the models and proofs, and a way to supply the verdict back.

The tooling exists. CBMC is what Amazon and AWS have used for checking libc and their crypto libraries. But getting good results out of it still takes a lot of maintenance and effort, which is the same labor wall as before.

The bugs only formal verification found

Both Apple and Microsoft have put serious work into this, and both found real bugs in their most scrutinized code.

And this is not confined to crypto: AWS is doing it at scale with Cedar for their role based access control policies, which Weston calls amazing.

Scaling the proofs with agents

The remaining question is how to scale formal methods, and here is where the numbers get striking again.

Aeneas can take code bases like Rust, or even specifications, and convert them into Lean proofs. Lean is a functional language for describing the proofs used in formal verification.

Microsoft demonstrated this on real code in SymCrypt: in a single week, mostly driven by agents, they generated 30,000 lines of Lean that formally verified AES-GCM.

What that verification buys is the whole point:

Out the other side you have mathematical grounding in the soundness and reliability of that crypto. "If we could have that for all of the boundaries I talked about previously, we'd be in a different game with respect to AI."

The case in point: the OpenAI sandbox escape

Weston connects it to the incident everyone in the hall had been talking about, extrapolating from public information around the OpenAI escape and the Hugging Face incident.

His assessment is that OpenAI did all of the right things. They had a properly isolated, strong boundary around their model evaluation. They gave exactly one proxy out, which was necessary to reach packages.

And that was enough. A capable model found what looks like nine different vulnerabilities on the fly, all of them logical vulnerabilities.

"What that tells us is you can follow the best practice out there, the best boundaries, but if you can't guarantee your code is free from logical issues, which is a tall order today, you're simply not going to be able to guarantee safety."

Then the call to action, which is the reason the incident is in the talk at all: if that package system had been formally verified with the right proofs, we might be talking about something different.

"But what about the meantime?"

Weston voices the objection to himself: Dave, you have talked a big game about prevention, but what happens when we cannot find all the bugs, and we cannot find them all. Everything described so far means taking 20, 30, 40 years of software debt and converting it to memory safe and formally verified software. What do we do while that happens?

The answer starts with an uncomfortable fact about how systems actually fall over today: most infrastructure, cloud and on premises, is owned without exploits at all. It is configuration.

And unfortunately, very little of that infrastructure can be reasoned about by agents, because very little of it has been converted into infrastructure as code or into other persisted, checkable policy. That is the gap. Prevention is the goal, because "we cannot get into a hand to hand combat detection, we will lose." So the job is to reduce attack surface, improve posture and configuration, and shift left on the infrastructure side. The less that reaches production, the less that is reachable in production, and the less there is for attackers to go after.

Give the agents a graph

Agents can reason about infrastructure holistically when it is in the right format. That means building graphs or ontologies of the assets and network flows in your organization, which unlocks a set of things humans cannot do at scale:

The number he offers as evidence that this is not hypothetical: an analysis of using agents with Checkov, which checks infrastructure as code against various mechanisms, showed that 78 percent of the findings can be resolved by agents.

Figure 6. The chronology of everything the keynote cites, offense and defense on the same rail. The two halves are running at the same speed, which is Weston's entire argument: the productivity is symmetric, only the strategy is a choice.

Three things you can do today

Weston closes with homework, and it is deliberately concrete.

"Attackers have changed the economics of what we're doing today, but as defenders we can choose to change the physics."

  1. Focus on your highest risk surfaces and make those secure by design. It does not have to be only memory safety. The same move applies to web.
  2. Figure out what your most critical boundaries are and start using agents to work on formal verification for those. There are many open source tools capable of it now, and you can start preparing yourself.
  3. Start using infrastructure as code and graph computation to build an agent army that can help with prevention.

"If we invest in this, we change the physics, we change the economics, and we lead the pack. Thank you."

The close

Susie Pallet comes back out. "Great energy." Housekeeping: the next session is in the same room at 10:30, and all of the main briefings start upstairs at 10:15.

Tomorrow's keynote is in the same room, with Yan Shoshitaishvili, team captain of Shellphish, speaking on vulnerability research in the agentic era. That is the offense side of the pairing Jeff Moss described at the top: today the hyperscaler view of defense, tomorrow what all of this does to the reverser and the exploit developer.

And at 6:30 that evening, the premiere of Midnight in the War Room in Oceanside A on level two.

Where it stands

A few honest notes on the talk, kept here at the end rather than threaded through the reconstruction.

The strongest part of the argument is the part that is hardest to argue with. The pessimistic half rests on internal Microsoft telemetry, which the audience cannot audit, but the individual numbers are checkable in kind: public benchmarks like CyberGym and Exploit Gym move in the same direction, the Anthropic and Dragos reports are public, and the UK AISI evaluations are published. The claim that harnesses beat raw frontier models on discovery benchmarks is the most consequential and the most verifiable, because it is the one that kills the "restrict the models" policy answer on technical grounds rather than political ones.

The vulnerability curve is presented with its own caveat, which is to Weston's credit. He says twice that we do not know whether it holds. A doubling every six weeks is not a sustainable rate for anything, and part of the current spike is certainly a backlog being drained by a new capability rather than a permanent new rate. What matters for his thesis is not the exponent but the floor it settles at, and nothing in the talk claims to know that.

The $3.61 number is doing a lot of work and deserves a footnote. It is the token cost of generating a crash level proof of concept from a vulnerability already found, in a kernel Microsoft controls, with a harness built for it. That is not the same as the full cost of a weaponized exploit against a hardened, mitigated target in the wild, which is exactly what Weston says next when he names nondeterministic mitigations as the remaining obstacle. The direction of travel is the point, not the decimal.

The optimistic half is a research agenda, not a product roadmap. Sila, Rustler, Aeneas, TRACTOR and the Lean work are real, and the Android and Azure hypervisor results are shipped and load bearing. But the gap between "30,000 lines of Lean verified AES-GCM in a week" and "your organization's boundaries are formally verified" is the same gap that has kept formal methods out of the mainstream for sixty years. Weston's bet is that the labor cost was the only thing in the way. That is a genuinely plausible bet and it is not yet a proven one.

And the closing homework is the honest test of the talk. Two of the three items, secure by design on your highest risk surfaces and infrastructure as code with graph computation, are things a competent security organization could start this quarter. The middle one, agents doing formal verification of your critical boundaries, is the one that will separate the labs from everyone else for a while.

Key takeaways

Chapters

Notable quotes

"Technology is political and that makes us all uncomfortable to say that word out loud. But we have to sort of embrace it. If we don't embrace it, politics will happen to us." Jeff Moss, 6:40

"The hyperscalers inherit the risk models of their customers. And if your customer is Ukraine, guess what? Your opponent is Russia. Like, you just want to sell rack space, but now you're in the middle of power conflict." Jeff Moss, 6:40

"Take out the water, you take out the military base." Jeff Moss on Iran targeting rural water districts, 7:38

"You can't inject a fake personality here." Jeff Moss on why the community still gathers in person, 8:41

"It hasn't been the death of programming jobs. What it's led to is allowing companies and people to think bigger. We can imagine larger things, more complicated systems. We can create newer opportunities." Jeff Moss, 11:31

"I might be the only optimistic person in this entire room right now, maybe at this whole conference." David Weston, 13:17

"I'm going to make this talk a 30 minute high effort social post just live. But you can't block or unfollow me because you're a captive audience." David Weston, 13:17

"Security has this underlying assumption. It's unsaid, and that is that we have these security boundaries, network, process, identity, encryption, and that is extremely hard and thus attacks that undermine them are scarce. What happens if that changes?" David Weston, 14:53

"The entire premise of cyber is based on this scarcity and supply economics around this." David Weston, 16:44

"These are serious vulnerabilities, the kind that I used to take a year to bespoke craft. They're being spit out at industrial speed." David Weston on the Windows harness findings, 20:55

"The average cost from a token perspective, $3.61. 21 minutes on average." David Weston on automatically generated kernel exploits, 22:13

"I would not bet against this curve. I fully believe that if we look at this and we draw a curve here, by the end of the year we'll be looking at automatic exploit generation being pretty commonplace and pretty commodity." David Weston, 23:27

"There's nothing that says technically that the only place that cyber knowledge can live in an agent is actually in the model. And a lot of places you don't want to put that in the model. Now that's counter to a lot of business models and other things, but the reality is you can inject that as a markdown file and it's actually more optimal in many cases." David Weston, 24:27

"So scarcity will not come from restriction." David Weston, 28:25

"This is not magic. This is productivity." David Weston, 29:03

"We don't want to go vuln for patch. We don't want to go exploit for detection, evasion for detection. Hand-to-hand combat with attackers will cause us to lose in defense." David Weston, 29:34

"What we want to do is retrain the physics here." David Weston, 29:34

"Memory safety does not mean security." David Weston, 35:00

"Formal methods are having a moment similar to reinforcement learning had with AI." David Weston, 37:34

"Nobody wants to maintain anything." David Weston on the fourth barrier to formal verification, 37:34

"A single shift that was wrong. Passed all the tests, passed fuzzers, passed human review. Only formal verification found it." David Weston on a post quantum key encapsulation bug, 38:43

"You can follow the best practice out there, the best boundaries, but if you can't guarantee your code is free from logical issues, which is a tall order today, you're simply not going to be able to guarantee safety." David Weston on the OpenAI evaluation sandbox escape, 40:34

"Attackers have changed the economics of what we're doing today, but as defenders we can choose to change the physics." David Weston, 44:24

"If we invest in this, we change the physics, we change the economics, and we lead the pack." David Weston, closing line, 44:24

Resources mentioned

The event and the people

Data sources and reports cited

Benchmarks and harnesses

Memory safety and automatic conversion

Formal methods and verification

Prevention and infrastructure

Full transcript
======================================== [music] A new warning. tonight about the scope of a massive cyber attack >> already believed to be the largest >> America under virtual invasion. Please welcome to the Black Hat main stage the president of Blackhat, Susie Pallet. >> [music] >> Good morning. By show of hands, who got less than 5 hours sleep last night? [snorts] Yep, that's Black Cat. You stayed up late because you were in a conversation that mattered because someone showed you something that you'd never seen before. because you finally met the researcher whose work you've been following for years or because you were in your room debugging that idea that wouldn't let you go. That's what makes this community different. Welcome to day two of Black Hat USA. If yesterday gave you a taste of what this week is about, today we go deeper into the research, the hard problems, and the conversations that will shape how we think about security long after we leave Las Vegas. Here's my challenge to you today. Don't just attend, discover. Black Hat is about more than what happens on stage. It's about what happens when you step into the convergence and meet someone new. When you sit down at Arsenal and talk to the person who built the tool you didn't know you needed. When you ask questions in a briefing that everyone else was thinking but didn't say out loud. Those moments, that's where the real work happens. Today we have a full slate of programming on the main stage, keynote presentations and highlevel conversations exploring some of the most pressing challenges facing our industry. Take advantage of it. Push back. Ask hard questions. This community gets stronger when we're willing to challenge each other. And tonight, don't miss the world premiere of Midnight in the War Room, a ground groundbreaking documentary that takes you inside the realities of modern cyber conflict and the impossible decisions that shape it. It's going to be one of the most talked about moments and tells your story. I hope you'll join us. But with before we dive into this morning's keynote, I want to bring someone onto the stage who needs no introduction in this room, but deserves one anyway. Over 25 years ago, this man saw something that didn't exist and decided to build it. A place where security researchers could share their work without filters, without corporate spin, and without apology. A place where breaking things wasn't just accepted, it was expected. where the focus was always on the research, the truth, and pushing the boundaries of what we understood about security. That vision became Black Hat. And that spirit, that relentless commitment to real research and real impact is why we're all here today. He has worn many hats over the years, hacker, founder, CISO, adviser to governments and organizations worldwide. But at his core, he's always been a builder. someone who believes that security gets better when we're willing to ask the hard and tough questions, challenge the status quo, and never settle for good enough. Please join me in welcoming the founder of Black Hat and president of Defcon, Jeff Moss. Thank you. Yeah, you didn't hear it, but earlier they were playing jazz um before you came into the room and you're like, "No, no, we need we need the beats." Okay, so um welcome. This is the 29th Black Hat. We're on the buildup to next year, which will be the big 30th celebration. And uh for those of you who haven't been to a black hat before, let's see a show of hands. Who's this? who's new here for keynote. Okay, fantastic. So, let me just give you a quick run of the show. I'm going to give you some uh some opening remarks, then we're going to kick it off to the keynote. Um, normally I talk about how big uh the community is. And every year we do something where if you don't have the financial capability to attend Black Hat because it can be expensive for people just entering the field, we have an alternative path to try to bring in uh new talent and that's our scholarship program. You write a a white paper, we review it. If we like it, you get admission for free. And this year 131 people uh attend Black Hat through scholarship. So, if you're here in the room and you got here by a scholarship, please raise your hand. Let's give them a round of applause. [applause] Right here. Right on. Um, [applause] it's uh some of our past keynotes in the last two years actually started as scholarship uh attendees. So, it is absolutely possible to go from a noob to a badass. Um, and we're a pretty pretty welcoming community. Uh we also have people attending from over 103 countries. So if you're here not from America, Canada or Mexico, raise your hand. Right on. Okay. So if you see them, right? Let's get a different perspective. Let's see how security works outside of our bubble. Now Black Hat this year, you might notice a couple of themes. There's four themes. And as you would expect, AI and autonomous threats, cyber conflict and live operations, sort of information operations, how do you build engineering resilience in the face of this in identity, trust, and control? Sort of the four organizing principles around a lot of the talks uh this year. And if you notice, probably three of those are pretty political. And this goes back to some of the things we've been saying over the past decade is that infosc and security is political. Technology is political and that makes us all uncomfortable to say that word out loud. But we have to sort of embrace it. If we don't embrace it, politics will happen to us. Just think about it right now globally with a conflict in Europe. Russia's all over giant hyperscalers in multi-tenant environments. Why? Because the hyperscalers inherit the risk models of their customers. And if your customer is Ukraine, guess what? Your opponent is Russia. Like you just want to sell Rackspace, but now you're in the middle of power conflict. Same thing with China. Great power conflict going on right now. that's most visible around AI and open weights debates, but remember when it was like high performance GPU chips trans shipping through countries in in Southeast Asia like this stuff is political and we need to have a view and an awareness around it if we want to be effective in our jobs locally or state right now. What's in the news? Iran. Iran hacking rural water districts in the middle of the country. Why? A lot of military bases are on rural water supplies. Take out the water, you take out the military base, right? Who's going to help defend? They don't have the budget. So, there's a lot of defenders donating their time from our communities to try to make things better. So, um, and then there's another kind of a weird thing. I don't say it's political, but like who here has noticed this weird humble brag from some of the Frontier models? Every time a model breaks out and causes some chaos, it's like this he humble bragging marketing material. Um, that's being noticed on Capitol Hill, too. There's a political impact to that kind of of marketing. And finally, we're in an election year, so expect more synthetic personalities, more influence operations. And where is this all leading? Um, I think one of the things that we can lose sight of in this whole autonomous agent-driven environment is the people, us, I mean, we are the cornerstone on which all of this is built. And there's rapid change right now. And in times of rapid change, in times of uncertainty, some of the coping mechanisms are you turn to your community. The most resilient communities in any kind of a natural disaster or man-made disaster are your local communities. That's where people find support and they find direction. So here you are. You're in a giant room of peers. You're in a giant community. Yes, things are moving very quickly. Things are changing. Markets are being disrupted. But we're all in this together and we will need to lean on each other in our community. And we do this by gathering in person because you can't inject a fake personality here. And the weird things are when things are extra stressful, attendance at conferences, not just security conferences, attendance at conferences go up. And I think that's because people, we innately know we need to see like what the going on. Like we need to talk to people. I need to talk to my buddy and get the download on what's happening. And that might not be happening on some Discord server, right? And so I want you to make sure like saying what uh Susie brought up. Yes, see the technical talks, but take time to build those connections that are going to support you throughout the year. A few years ago, I mentioned on stage that AI is essentially a prediction engine. And if I was a business leader, I try to make all of my problems prediction problems because if my problems can be predicted, AI is perfect for me. And the prediction engine will get faster and cheaper. My problems can be solved faster, more efficiently. Great. Holy moly. in the last couple of years. It is a total sea change of what is possible now. And um and so thinking about our keynotes, we have two keynotes today and tomorrow. Today I wanted to have a keynote that focused more around defense, how are hyperscalers, how are we addressing the risks and opportunities um of AI? And tomorrow we're looking more at attacks. what is it doing to the academic the reversing the exploit developer? How is it helping or hurting them? And finally, one thought I want to leave you with is I'm I'm generally a pretty skeptical person. It's turned out really well for my career in security. [laughter] But for AI, I think um in listening to developers and programmers talk, I'm I'm talking Unix greybeards talking about you know vax VMS to today the trend line is every time there's an improvement in say IDE or the invention of IDE and development languages, object-oriented programming, so on and so forth. Abstraction, abstraction, abstraction. It hasn't been the death of programming jobs. It hasn't been the death of AI, I mean of uh it. What it's led to is allowing companies and people to think bigger. We can imagine larger things, more complicated systems. We can create newer opportunities. And so I think what's going to happen is yes, there'll be a lot of disruption, but at the end of the day, there'll be more jobs, not less. Because all these companies around the world are going to be wanting to build bigger and it's going to be built on the backs of what we do, providing the reassurance and the resiliency for them to take bigger risks. So with that, I'm really interested and excited to see what David Wesson has to say. Dave now leads the Aenic security team at Microsoft where he builds the AI models. He builds the agents and the evaluation systems for defense at scale. And uh so please put your hands together and let's welcome Dave Wesson. Thanks man. I can't dance as well as Jeff Moss, but I'm trying. So really hyped to be here. When I got asked to do a keynote, I was really excited. I have a lot of pent up. I saw on all the socials a lot of AI apocalypse and doomsday stuff. And I might be the only optimistic person in this entire room right now, maybe at this whole conference. And so I said, I'm going to make this talk a 30 minute higheffort social post just live. But you can't block or unfollow me because you're a captive audience. So, I'm going to talk to you about what I think happens when attacks get less rare, less uh scarce, and what we can do to defend in that environment. So, I want to give you kind of my bias and my vantage point. I spent the last 20 years in the trenches of security. I'm talking Wukry, Stuckset, I lived it all. I have all the trauma. So I built operating system security for Windows, Linux, you name it in Azure, all those other places, EDRs, and I've led vulnerabilities in red teams. And the last nine months, I've switched over to the AI world. I've switched to the dark side. I've been working on vulnerability discovery harnesses, training frontier models for cyber capabilities, and also building defense. And so I thought I could offer some of the things I've learned from these two vantage points, kind of combining the nexus of where I think the future is going. So security has this underlying assumption. It's unsaid and that is that we have these security boundaries network process identity encryption and that is extremely hard and thus attacks that undermine them are scarce. What happens if that changes? Cyber security the entire house that we live in the cyber house the roof the the the foundational floor is based on these being rare. So we use these boundaries to isolate our networks to separate trust domains in places like the cloud to create authentication encryption data protection policies and ultimately to contain failures with things like hypervisors or process sandboxes. So every control or policy in your enterprise in your company in your business on your phone relies on these not being easy to undermine. And you could see this in the economic models around things like bug bounties. So for example, we have no easy way or no cheap prices for undermining things like processes etc. Our prices scale with the importance of the boundary. So if you want to have a hypervisor bypass, that's going to be 20 times as much on the open market as a boundary crossing for something like a process. And so this economic model has been priced in for a long time and is reflective of what it's like on the ground. And even further, we've made the assumption that if a vulnerability is a potential risk, actualized risk, which is exploitation and implemented attack, is even more costly because then you price in expertise and scarcity for things like mitigation bypasses or all the other techniques. And again, that's reflected in the cost pricing of exploits versus vulnerabilities. And as a result, for the tens of thousands of CVEes that you hear about every day on LinkedIn or your news of choice undermining existence, we actually usually only end up with 90 or so in the wild exploits every year as tracked by the Google folks. So that's it's really exceptional at this point for a boundary to be undermined. And so the entire premise of cyber is based on this scarcity and supply economics around this. Now as a result this has really held true for a long time. Most the breaches today occur far above the boundaries. Right? Verizon's breach report tells us most of the attacks actually occur at the credential theft level which is significantly cheaper traditionally than undermining a boundary. It's also happening through things like fishing and social engineering. And when a vulnerability is exploited, the vast majority of the time that's a vulnerability that's exploiting or being a vulnerability being exploited that's known, which means there was time and uh to to to implement that. So all of our economics, all of our strategy is based on this today. Now we can infer a couple dominant security strategies that are wholly reliant on this principle of scarcity. The first one is you could ignore the SDLC or at least have less priority on it. And what I mean is static analysis, safer languages, principle of lease privilege, strong identity around our the software we build because we can just patch fast when something's known because vulnerabilities are not often exploited. Now the pressure test against that is what happens if exploits just become another commodity. We lived through this in the 90s. trivial to exploit that just patch everything fast strategy has carries significant risk. The second inferred dominant strategy is, hey, we can do less around prevention, less around software. We'll just detect and respond fast, right? This is what all the vendors pitch you. We're going to sprinkle some AI here and very quickly we'll just detect and respond. The truth is the dwell time is getting longer and it's getting harder to detect. And the pressure here is most detection is based on invariants that don't change. The idea is, hey, attackers are software developers. They can't afford to change their software, their implants, their C2, their lateral movement management tools every OP so we can continue to detect them. But in a world where that becomes cheaper, supply economic supply. So what do we do when this scarcity principle no longer has our back? What does it look like to defend in that environment? And the bigger question is, are we actually there yet? So the assumption that vulnerabilities are scarce, right? This is a potential risk is being undermined as we speak. What you're looking at here is a curve from this year around the pressure that AI is putting on vendors. This is a Microsoft number, but as you see, this trend seems to be holding true for Google, Apple, and I bet just about every other popular software vendor. MSRC is doubling the number of vulnerabilities that they are processing and patching every less every six weeks. That is an incredible number. We're nine times the vulnerability volume that we were in March. And again, we don't know if this curve is going to hold true, but boy, if it changes, we're in deep trouble in the places where we presume vulnerabilities are scarce. And again, I would urge you to think about all the different software vendors out there. This is representative of the MSRC cases that combine open- source that Microsoft uses and our first party software like Windows and Office. But this is a significant jump and it's something we'll have to continue to look at. So the next thing is, is this actually correlated to AI? You might ask yourself, or is this just people are getting better at finding bugs? Our internal data says yes, it's heavily correlated. We released a new internal vulnerability harness and we've turned it on to Windows on April 1st. And since then we found 66% of the critical and important issues since April 1st than we did of all of last year. So this is not just a correlation. This is the fact. It is AI that is driving this. And these are real vulnerabilities. a good example when it comes to boundaries. In this data set, I saw seven remote TCP IP. Vulnerabilities that cross both the kernel and the remote boundary, which are both absolutely critical for Azure, you name it, and every other Windows system on the planet. So, these are serious vulnerabilities, the kind that I used to take a year to bespoke craft, they're being spit out at industrial speed. And it's again, it's not Windows. you look at Linux, you look at any other operating system out there, I think you'll see a pretty strong correlation. So the next question is if potential risk is is is driving up dramatically is actualized risk are exploits. And here's a very unique stat I'm sharing. We again we have an internal uh vulnerability harness called Mdash and it's very good at finding vulnerabilities in this gentic system. It's found roughly 200 Linux kernel vulnerabilities in our internal a Azure Linux distribution that we're working with the community to fix. We added a new module to this to help us triage. Guess what it does? It turns a static analysis result into a P. That's worked much better than we ever thought. Of the 200 vols, we can automatically generate 182 crash level PC's. Many of them are fully working exploits. I'm talking root exploits automatically spit out from a vulnerability. And the average cost from a token perspective, $3.61, 21 minutes on average. And again, most of the world runs Linux in some capacity or another. And when you run Linux, you're relying on the kernel boundary to save you. This is under attack. And it's not just internal. If you go look at Exploit Gym, what you'll see is the big frontier models making incredible strides. In fact, I looked at Exploit Gym this morning, and this number for Mythos had already been doubled by OpenAI roughly. So, of 898 real world vulnerabilities, and these aren't just Linux kernel, which are arguably easier to exploit in some ways. These include things like browser vulnerabilities, etc. They can generate 157 exploits out of roughly 900. What's really holding back at this point are non-determin nondeterministic mitigations, control flow, as randomization, ASLR, etc. Those things are non-deterministic. They make things harder. They will not guarantee these can't be exploited. So, I would not bet against this curve. I I fully believe that if we look at this and we draw a curve here, by the end of the year, we'll be looking at automatic exploit generation being pretty commonplace and pretty commodity. But these are advanced frontier, pick your uh term, dour cyber models. We're restricting them, right? Isn't restrictive access going to maintain scarcity? Isn't that going to save us? Nope, it's not. And I'll tell you why. If you go and look at CyberJ, which is at least a vulnerability discovery benchmark, the top entrance are not actually frontier models. They're harnesses. Harnesses make use of frontier models, but they also inject context in a few other places. They can inject cyber expertise and through tooling. It can be encoded into the harness. There's nothing that says technically that the only place that cyber knowledge can live in an agent is actually in the model. And a lot of places you don't want to put that in the model. Now that's counter to a lot of business models and other things, but the reality is you can inject that as a markdown file and it's actually more optimal in many cases. So the idea that we're going to sort of restrict our policy our way out of this I think is unrealistic and in many ways we need to prepare for that not being the case and I think these harnesses on Cyber Gym from a variety of vendors are really strong evidence of that for now. There's also another assumption that I see played out all the time which is hey we can just detect right even if all these exploits start coming we'll just detect our way out of this many you're employed in security operations centers at vendors etc and the idea here is like it's super expensive to code a framework or an implant so people just keep using packers and obuscation tools on the same stuff and they keep using the same TTP so we'll work against that and that's going to give us durability and detection Again, we're making the assumption here that evasion of detection is somehow a scarce property because it has been. If we look at this model called the pyramid of pain, which is an interesting model around essentially the invariance in detection, the idea is that TTPs, tools, and artifacts are the most expensive to change. You can change hash values, IPs, domains, no problem. But TTPs stand durable. Your tools are more expensive to change and certainly your artifacts. But now instead of having to retrain the operator, which would have been expensive for cyber operations, we can just use autonomous operations. Instead of obfiscating, we can create a bespoke set of tools or frameworks per target. So we are challenged just like exploits and vulnerabilities on this assumption that somehow this is going to be hard for attackers. It's not going to be scarce. And again, we have real world evidence of this. If you look at the canonical sort of case study around this was anthropic reporting back in November of last year, a cyber operator conducting most of these operations against, you know, top tier targets 80 to 90%, essentially using cloud code with sub aents and getting ostensibly good results based on anthropics observations at 80 or 90%. Then we saw in May of this year, Drago reported a water utility being targeted by an AI assisted group who is building their framework during the op. So they're able to see the code keep getting regenerated ostensibly. This is Python. So you can see the additions there were all the hallmarks of this being generated in AI. And this single op the attacker generated 17,000 lines of C2 implant etc code just for this operation. That's essentially proof that evasion and the assumption that artifacts are going to stay the same just isn't there. And if we extrapolate this a little bit more, if we calculate the trajectory, we can see from groups like the UK is SI, which does a sort of testing of frontier models against their ability to conduct 32step autonomous breach operations. We're getting to 9.8 steps out of that 32 at 10 million tokens. And that's up 59% just this year alone. So our trajectory is autonomous operations will just be part of the course. And again I would not make any assumptions about any models doing this. We can inject this at any place. So scarcity will not come from restriction. So I told you at the beginning I'm probably the only optimist in the room. How can I be optimistic after all that? Well, a couple reasons. First is this is not magic. This is productivity. The attackers are much more agile. They go asymmetric to defenders. They've done that since time immemoriam. And so they move first on AI because you have policies, restrictions, auditing, compliance, and token costs that slow you down. But defenders have the exact same advantage. But let me tell you where we don't want to go. We don't want to go van for patch. We don't want to go exploit for detection, evasion for detection. Hand-to-hand combat with attackers will cause us to lose in defense. We will be asymmetric. We don't want to do that. What we want to do is retrain the physics here. We want to figure out where we can use this production advantage to actually turn the tables. And I think we can do this. We can use the same productivity advantage but invest in durability. We can shift left and make more secure software that'll limit vulnerabilities. We can move away from handtohand detection. Detection is still great. It's just it's necessary but not sufficient. We can move to more prevention mechanisms and we can use secure by construction and even formal methods to get deterministic safety. If we could do that along a realistic timeline, then we can turn the tables and drive this problem towards attackers. And I really believe that. And let me show you why. First is secure by design and construction are really coming into their own. There are more tools than ever to use safe system level languages. Formal verification is almost tailorade for AI. It gives AI and Oracle to write software that is provably safe against the set of properties that you need. Perfect for boundaries. And prevention is getting much simpler because using tool sets like infrastructure as code, you can actually reason about the safety of your infrastructure and do things to fix it at less cost with less people. All of these things would drive based on the data durable change in attacker economics and turn the tables. So let's talk about secure by construction. You've all heard this, but I want to put this in front of your face as you see that vulnerability curve. About 70% of vulnerabilities that are patched today, at least by the major vendors, are memory safety issues. Safer languages like Rust and Golane eliminate those. And we have strong evidence of this. Google has done an amazing job of proving this in Android. In 2019, 76% of the vulnerabilities they patch in this operating system used by billions of people from cars to phones and everything in between, 76% were me safety issues. In 2025, it's less than 20%. And that's because they wrote five million lines of Rust code, which based on their own analysis has a thousand times less defects. And they haven't shipped a single memory safety issue in that code. That's amazing. And Azure has done something similar in the containment boundary. rewrote the hypervisor in open source in Rust and it's now scaling past 1.5 million virtual machines without an incident. So if the world's most popular mobile operating system and a very popular hypervisor can do this, why isn't everyone doing it? Well, traditionally it's expensive. You need experts who know how to do this. You need to learn new languages. All sorts of reasons. You need to convert old code bases. But AI is changing this, right? There's a project from Microsoft research called Rust Assistant. It showed that it could take 74% [snorts] of the failures, compilation failures in Rust code and fix them automatically without user intervention, passing tests improving. Similarly, for existing C code bases in a variant of C called safe C or check C rather, that's pretty similar to some of the things that you'll see from clang and Apple. They were able to infer memory safety contracts in this codebase and they generated 86% of those code contracts that'll check for what are called spatial safety vulnerabilities or buffer overflows. This is productivity driving memory safety driven from AI. It doesn't look the same as generating automatic vulnerabilities, but it is absolutely critical. The real frontier though is doing this automatically and there's great work happening here. Sila is a really good example, great paper from sort of Google, Microsoft on taking SIMP, which is a library in Windows that does most of the core encryption, TLS, you name it, and converting it automatically to safe Rust. They were able to take a Shaw 3 method, which is a few thousand lines of code, convert it automatically to Rust, and then run tests. All the tests pass and the performance was within 1%. Rustler is a more sophisticated project. I can do this with dependency graphs context and what it does is generate Rust and then work through those error checking until you have valid code that passed tests. And it's not just Microsoft or Google doing this. DARPA, who you could argue really showed us first where the agentic vulnerability discovery was happening, actually has a project called tractor where they're sponsoring and giving out data sets for automatic conversion. If we can land this as a community, we can really drive uh security forward. But with memory safe, are are we safe? Are we done when we do all this? Unfortunately, no, we're not safe. Memory safety does not mean security. You still have logical issues. You have authentication issues. You have crypto issues. You have lots of issues. Most of the issues today are memory safety, but they won't stay that way forever. And Anthropic has done some really interesting work here. They released a blog post and a paper recently about using Claude to do an analysis of a postquantum encryption algorithm uh for digital signatures called Hawk where they were able to to find a cryptographic attack which is traditionally the most scarce I would say of security expertise is people who can do cryp analysis and they were able to do that with Claude. They also were able to show that they could do some attacks against AES128 that drove around an 800 times increase in attacks against that in terms of performance. And then they also had a great bug in wolf SSL around forgery. These are traditionally I would say top tier in terms of scarcity or complexity to reason about. And what we're showing is that AI can find not just memory safety issues but logical flaws. So if memory safety removes bug classes, we need something that's going to guarantee safety prophecy properties, guarantee the logic. Can AI help us with that? Formal methods have a really rich legacy in computer science. You know, started in the 60s, we had model checking in the abstract form which essentially creates a you know mathematical representation of the logic in your program and then does abstract reasoning on that. Symbolic checking took that further further compressed state space and actually allows you to compute whether or not a given code base violates a property. If that property is violated in this mathematical proof, you get a reproduction which is really awesome. And as a result, we've seen high-risk safety platforms start to adopt this. The reason it hasn't gone mainstream is writing the specifications is hard, takes a lot of expertise. Writing the proofs is even harder. And then you have state space explosion for complicated programs. And then finally, you have to maintain this. Nobody wants to maintain anything. So the key question is, can AI make it scale? And from my standpoint, formal methods are having a moment similar to reinforcement learning had with AI. Reinforcement learning has become absolutely critical to modern AI. Even though it was invented back in the 60s and 70s, formal methods is perfect. It gives AI generated code in Oracle for correctness both from security but performance reliability. It's almost tailorade in my opinion. So what do we need to do to get there? Well, we need to be able to specify what must never happen. That can come from things like specifications e etc. You then need to create a model of that specification, a checker that validates those models and proofs and you need to be able to supply the verdict. There are awesome tools out there like CBMC forc that Amazon and AWS have used for checking of lib C their crypto libraries etc. But you still requires a lot of maintenance and effort to get the good results from this. But we're seeing both Apple and Microsoft doing a ton of work around this. Both Apple and Microsoft have taken their core crypto libraries and done formal verification and both of them through formal verification have found really interesting crypto issues. Uh in Apple they found uh an error in MLDDSA with Microsoft. We recently found in protochem which is a key encapsulation algorithm for PQC a single ship that was wrong. passed all the tests, passed fuzzers, passed human review, only formful verification found it. And this is not just for crypto. AWS is doing this at scale with Cedar for their arbback policies, which is amazing. So the key question, at least in my mind, is how can we if we think formal methods is going to be really helpful, how do we scale it? And again, we're seeing really interesting research here. Anias [snorts] is able to take code bases like rust or even specifications and convert them into what are called lean proofs. So, lean is a functional language for uh describing proofs around uh code for formal verification and Microsoft demonstrated on real code and simcript that in a single week mostly driven by agents we could generate 30,000 lines of lean code that formally verified AES GCM and that means it's free of logical issues. It's free of cryptographic issues at least as the proofs stand and it's free of memory safety. So we come out of this having mathematical grounding in the soundness and reliability of that crypto. If we could have that for all of the boundaries I talked about previously, we'd be in a different game with respect to AI. Now how can this apply to sort of the the questions duour? So I extrapolated a little bit from the public information around this uh open AI escape, right? and the hugging face incident. And what OpenAI did, at least from my assessment, is they did all of the right things. They had essentially a isolated um strong boundary around their model evaluation. They gave one proxy out which was necessary to get to packages. And unfortunately, that was enough for a capable model to find what looks like nine different vulnerabilities, all forms of logical vulnerabilities on the fly. What that tells us is you can follow the best practice out there, the best boundaries, but if you can't guarantee your code is free from logical issues, which is a tall order today, you're simply not going to be able to guarantee safety. This is a call to action. If that package system was formally verified and with the right proofs, we might be talking about something different. So Dave, you've talked a big game about how we can change and prevent things, but what happens when we can't find all the bugs and we can't find them all, right? Everything I'm talking about here is taking 20, 30, 40 years of software debt and converting it to memory safe and formalized software. What do we do in the meantime? Well, today most infrastructure cloud onrem is owned without exploits. most of its configuration issues and unfortunately very little of that infrastructure can be reasoned about by agents because very little of it has been converted into infrastructure or code or other persisted policies that can be checked through a number of scalable mechanisms with AI but we know that prevention is key we cannot get into a hand-to-hand combat detection we will lose so we need to be able to reduce attack surface improve the security posture and configuration and ultimately shift less in the infrastructure side. The less that reaches production, the less that's reachable in production and that means less for the attackers to go after. And we know that agents can reason about infrastructure holistically when they're in the right format. So creating graphs or ontologies of all the assets and network flows in your organization can allow them to do things like create uh locality or understand centrality of the biggest uh risks out there and focus remediation on that. So for example, you could compute a identity style attack graph and then use that to figure out which uh accounts are are overprivileged. You can do the same thing for posture, right? Which which devices or uh uh uh infrastructure pieces in your organization have the biggest attack surface? And you can allow agents to validate that. On top of that, once you've computed a graph, edges become more obvious. So if you've computed a graph of normalized authentication across your organization and all of a sudden you have a connection from accounting to the DC, one that creates a point of analysis that agents can operate on that. Two, agents can scale out and do more proactive hunting than you ever could do with humans. So in the end, we believe agents can drive meaningful productivity here. And we have good examples of that. a check off an analysis of using agents for check off which checks against um uh various IA mechanisms have showed us that 78% of the findings in that check off can actually be resolved by agents. So attackers have changed the economics of what we're doing today but as defenders we can choose to change the physics and I want to leave you with three things that you can do today. One is you can focus on your highest risk surfaces and make those secure by design. It doesn't have to just be memory safety. It can also apply to web. Number two, you can figure out what your most critical boundaries are and start to use agents to work on formal verification for those. There are many open source tools that are capable of doing that now and you can start to prepare yourself. And then finally, you can start to use infrastructure as code and graph computation to build an agent army that can help with prevention. If we invest in this, we change the physics, we change the economics, and we lead the pack. Thank you. Right on. Great energy. Hey. Okay, everyone. That was fantastic. It was everything I'd hoped it to be. I um I want to give you a couple of housekeeping notes, then we move on to the rest of Black Hat. Um in this room at 10:30 will be the next session and all of the main briefings talks will start upstairs at 10:15. Um tomorrow's keynote will be here again uh where Yan Shisho Shiovali who was also the team captain for Shellfish um is going to talk on vulnerability research in the Agentic era. Then tonight 6:30 there's a special premiere of Midnight in the War Room. It's happening. It's a movie screening. It's happening in Oceanides A which is on level two. All right. Well, thank you for coming to the keynote. It's been fantastic welcoming in everyone for day two and I look forward to seeing you tomorrow. Thank you.