Black Hat USA 2026 | Keynote: Vulnerability Research in the Agentic Age
By Black Hat
Summary
Topics Covered
- Vulnerability properties matter more than model power
- Responsible disclosure breaks under AI's discovery rate
- Rust rewrites inherit the original vulnerability DNA
- Agents amplify rather than replace human researchers
Full Transcript
I warning tonight about the scope of a massive cyber attack.
Already believed to be the largest America under virtual invasion.
Please welcome to the Black Hat main stage the president of Black Hat, Susie Pallet.
Good morning and welcome to the final day of Black Hat. We've made it and I wanted to say start by saying thank you.
Thank you for showing up, for bringing your expertise, your curiosity, and your willingness to engage. for asking the hard questions, for sharing your
research, and for making the connections that turn ideas into action.
Before I say anything else, I want you to look around. Look at the people sitting next to you, behind you, across the aisle.
These are the people who've been solving problems with you all week, challenging your assump ass assumptions, sharing what you've learned, and staying up too late because the conversation was too
good to walk away from. This is what Black Hat is really about.
Over the past few days, thousands of people from around the world have come together to do what this community does best. exchange ideas, challenge
best. exchange ideas, challenge assumptions, explore new technologies, and build relationships that last long after we leave Las Vegas.
From the groundbreaking research presented in briefings to the expert insights shared at our summits, from the hands-on experiences across the event to the hallway conversations that went
deeper than any session ever could, you made this week what it was.
Black Hat has always been at its best when discovery leads to collaboration and that collaboration leads to growth.
And this week we've seen that happen over and over again.
Whether you've spent your time learning from researchers or testing tools in Arsenal, exploring emerging technologies, or reconnecting with colleagues or meeting someone new who's
going to change the trajectory of your work and how you think. I hope you're leaving with new skills, new insights, and new opportunities.
But there's one thing that keeps surfing. One question we can't ignore.
surfing. One question we can't ignore.
AI is fundamentally changing how we find vulnerabilities, how fast threats move, and how we analyze them. Today's keynote will
analyze them. Today's keynote will tackle that head on.
As AI becomes more capable, more autonomous, what happens to vulnerability research? Where does human
vulnerability research? Where does human ingenuity fit in a world where machines can hunt for flaws faster than we ever could? And what does it mean for this
could? And what does it mean for this community, for the researchers, the defenders, the builders all in this room when the tools when that we rely on start to think for themselves? They're
not hypothetical questions. They're the
reality we're facing live right now.
But before before we bring up today's keynote speaker to explore those questions, please join me in welcoming an amazing person, the founder of Black
Hat and president of Defcon, Jeff Moss.
Hey, good morning.
Glad uh we all made it here today. I've
got uh just a little bit of a backstory I want to share with you uh for today's uh keynote speaker and maybe put it a
little bit into context.
Um so about 30 years ago almost 30 years to this month ago the concept of a capture
the flag CTF was accidentally invented at Defcon. The origin of CTF was
at Defcon. The origin of CTF was essentially a network fight that happened at Defcon about 31 years ago.
And since then, CTFs have evolved to like run the world. CTFs everywhere all the time. But it wasn't purpose-built to
the time. But it wasn't purpose-built to do that. It's grown because CTFs are
do that. It's grown because CTFs are incredibly useful. And due to the nature
incredibly useful. And due to the nature of how Capture the Flag started at Defcon, it was purely an offensive defensive full combat. Only rule was
after the first year, you couldn't take over the scoring system or the router because that's what the attackers did the first year.
So we're like, oh, maybe we need rules.
Um, and it so many lessons were learned, right? Because how can you defend if you
right? Because how can you defend if you don't know how people attack? attack
informs defense.
And so if you don't have a sort of ground truth of how attacks and exploits are being developed, it's probably going to be pretty hard to create an effective
defense. So um that's a way of sort of
defense. So um that's a way of sort of tying into our keynote uh Yan who
as an origin story for his superpowers in reversing. He uh his mom had a PhD
in reversing. He uh his mom had a PhD essentially in uh computer science in Russia before that was really a thing.
Then the fall of the wall happened and all of a sudden academics in Russia were starting to be pressured to endorse political candidates. And when the KGB
political candidates. And when the KGB came around and started pressuring the family to endorse the KGB candidate, they fled Russia. He fled to America.
He's 8 years old.
Um, and never thought he'd be an academic like his parents until he gets to university and he starts playing capture the flag and he loves hacking
and he can't stop hacking and then he realizes as a graduate student he's getting paid to essentially hack all the time and he's doing a lot
of these tasks manually. He's like,
well, what if I automate this as much as I can? So, he goes and he creates this
I can? So, he goes and he creates this multimodal framing fuzzwork called anger. And with anger, he and his CTF
anger. And with anger, he and his CTF team, Shellfish, basically run the board worldwide contests for about two years
because Anger allowed them to reverse and exploit in other ways that the existing tooling could not do.
Since then, Anger's probably been used in over a thousand different research and academic and product tooling.
This uh highly contributo uh mindset of an academic, right, giving away the tooling, advancing the state-of-the-art
led team shellfish then to organize and run Order of the Overflow, which was an organizer of the Defcon capture the flag. So his whole career started with
flag. So his whole career started with capture the flag and they ended up running capture the flag and through
this whole process and now uh being at uh Arizona State University with his students and graduate students developing new exploits and new agent
tools, he's seeing how the exploit development trajectory has kind of changed from art into science. And so I'm really looking
into science. And so I'm really looking forward to hearing his story and uh and and having him share it with you. It's
my pleasure to introduce Yan Shisho Tash.
Hello hackers.
So, as Jeff said, I have quite an origin in the um capture the flag community. In
my day job, I am an associate professor at Arizona State University where the biggest thing I do, my priority
is research into vulnerability analysis and exploitation and uh repair and so on and then education of the next
generation. I'm very excited to talk
generation. I'm very excited to talk about this all today from the lens of the agentic age.
What does it mean to be an academic researcher? Well, it means that I sit
researcher? Well, it means that I sit around and I have shower thoughts, bright ideas that I take and typically pass on to my bright graduate student
researchers who then run with them and produce the research projects, the research papers,
and the research prototypes that lead to the breakthroughs.
So what a lot of what Jeff was talking about happened when I was this graduate student and I was handed the bright idea
and has kept going forward beyond that.
That means what I'm talking about today isn't just my work. It is the work of an amazing team of colleagues without whom
I would not be here today. I'll try to acknowledge them as we cover those parts of research where they contributed. They
they led the the research. They were the brilliant graduate student. But I'll
acknowledge them all cumulatively right now. The lab uh SECOM lab and the center
now. The lab uh SECOM lab and the center for cyber security and trusted foundations. Coincidentally, the CTF
foundations. Coincidentally, the CTF center at Arizona State University.
Awesome group of people. If you're
looking for graduate school, talk to us.
Um all right. So these kids implement these brilliant ideas. They're
not all brilliant. Sometimes these ideas are terrible for the record. Uh but of course those don't make it into the talks. Um they take the ideas, they
talks. Um they take the ideas, they implement them. They spend a lot of time
implement them. They spend a lot of time working on this and you can probably all start thinking wait a second you know
we're in the agentic age. Instead of the graduate students we can prompt the agents. The agents will generate really
agents. The agents will generate really infuriating text but you know they get results. Realistically
results. Realistically this is what happens now.
The students prompt the agents the students do the agentic prototype. The students
get results and and and what's really interesting about this is it puts me very far back from the the final
outcome of our research. In the abstract for this talk, I mentioned my research has led to the discovery of thousands of vulnerabilities
and many of them quite recently as the agentic age hypercharged everything and many of them allow us now to start stepping back and
thinking about this field in aggregate as modern AI transforms vulnerability research. Some people might say, "Okay,
research. Some people might say, "Okay, let's just paste everything into chat GPT and solve the problem. Find all of the bugs." Of course, the reason I'm talking
bugs." Of course, the reason I'm talking about the agentic age and not just the age of LLMs, it's because there's a very very serious
difference between the two. We start
pasting things into GBT, asking it to find bugs, submitting those bugs to maintainers, and we get things like the headline I read yesterday that Apple is
pushing back on their uh vulnerability research award program because they're being overwhelmed by slop.
And this slop is a result of AI hallucination of people that don't really know what they're doing enough to filter that AI hallucination. That's not
what I'm going to talk about today.
going to make the assumption that we're going to work that out because back in the days of fuzzing, back in every advancement throughout the history of the vulnerability research field,
especially the academic the automation of vulnerability research, we ran into similar issues. What I'm going to be
similar issues. What I'm going to be talking today is people that eventually will know what they're doing in this space enough to use these
technologies to create highquality bug submissions to find actual bugs. We're
not going to talk about hallucinations.
We're going to talk about vulnerability research in the agentic age. We're going
to step back and I'm going to show you three ways that we can find vulnerabilities autonomously.
Three things that lead to headlines where or research papers where approach
XYZ finds three, four, five digit bugs in software. There's three ways to do
in software. There's three ways to do this. One, you can analyze better. You
this. One, you can analyze better. You
take an analysis that existed before Jeff mentioned anger. It is a base a research base of our um kind of research
trajectory in static analysis of binary software and we make it better and so it it it has fewer false posers. It is more performance faster so that before the
timeouts it can actually find the bugs.
Think a good example in the agentic age of this of course is mythos. It is more capable models that find bugs that other
models would have missed. And we'll talk about this again later.
You can analyze more things.
So you say forget about being better.
We're just going to make something that is, you know, better in a way that's more resilient, that might scale farther and we're just going to analyze everything. Before we could analyze
everything. Before we could analyze eight programs, now we're going to analyze 800. And of course, we're going
analyze 800. And of course, we're going to find lots and lots of bugs based on that. And I have research in this space.
that. And I have research in this space.
One of my awesome uh former students now Dr. J. Varias
Dr. J. Varias created a platform a program an analysis program called arbiter built on top of anger that could analyze binaries at a
scale where we were analyzing for example every program shipped into repositories for vulnerabilities. Now,
often times you don't actually need to make this analyzer better. You just need to make it more applicable. And so then
by casting a wider net, you find more bugs. Or, and this is the crazy thing,
bugs. Or, and this is the crazy thing, you just analyze things differently than they used to be analyzed. And it
turns out if you just tweak how you're analyzing, you find new bugs. Why?
Because bugs are everywhere. Let's say
we want to make our CVE cookies and we have our rolled out dough of vulnerabilities. This dough is all
vulnerabilities. This dough is all vulnerabilities in a given program.
And we take our research prototype anger. It's our anger logo. We build
anger. It's our anger logo. We build
some static analysis on top of it and we use it to analyze the code for vulnerabilities and it
literally takes this substrate of vulnerabilities and stamps out. We find
those vulnerabilities that Anger finds and we fix them. They're gone.
For a static analyzer, what this means is if we run this tool again, no new vulnerabilities.
Is the code perfect? No.
We just happen to have found the vulnerabilities that we're going to find with this technique and we can tweak the technique a bit. So
we build research prototype A on top of our Angular research vehicle. We build
research prototype B and it's slightly different and that slight difference will find a tiny tiny additional set of bugs. What does that mean? Does that
bugs. What does that mean? Does that
mean research prototype B is bad or research prototype A is better? No, it
means we got unlucky. Research prototype
B came in after A.
And this is something that because of the trajectory over time of software development
and the development of analyzers to find bugs in that software makes it very difficult to really understand the efficacy of different tools.
For example, let's say I go and manually look at every binary in the Ubuntu repositories.
I would of course also look at that their source code. Um,
and I go afterwards and I try to find bugs and I find a bunch of bugs because as a cookie cutter, a vulnerability cutter, I will find different types of vulnerabilities. I'll cut out a
vulnerabilities. I'll cut out a different shape in the vulnerability landscape than anger did.
Does that mean the static analysis was bad? No. It means it was different than
bad? No. It means it was different than me. If he inverted the ordering, it
me. If he inverted the ordering, it would be the static analyzer that comes after me and finds a ton of vulnerabilities.
Now, let's go to something a little more fuzzy. This is dynamic analysis that's
fuzzy. This is dynamic analysis that's very stochastic. Fuzzers. This is the
very stochastic. Fuzzers. This is the logo of American Fuzzy Lop, which is the work of another amazing uh group of people. nothing to do with with us, but
people. nothing to do with with us, but we also do research in the space.
And if you run a fuzzer on even well analyzed code that hasn't been fuzzed before, in part just because the fuzzer does different things, it will find new
vulnerabilities.
And unlike a static analyzer, because fuzzing is a very stochastic pro uh uh process, you build random inputs and you shove it into software and the software
crashes because you're triggering something that that people didn't think about fuzzing. Unlike most static analyzers,
fuzzing. Unlike most static analyzers, if you just run it again, you're going to find something around those corner cases you missed the first
time. Okay?
time. Okay?
And now this brings us to this question. For the
last decade before the agentic age starting from around 2013, 2014, 2015, which was the process of the DARPA cyber grand
challenge, pushing the world really toward autonomy and vulnerability research. Ever since then, we've been
research. Ever since then, we've been experiencing this fuzzing renaissance.
new fuzzers coming out that say, "Hey, we fuzzed well fuzzed code. We fuzzed it a little differently. We found new bugs and so now we're claiming that we're better." It's actually very very
better." It's actually very very difficult to measure.
Now after a decade of hearing this new fuzzer is better, this new fuzzer is better, this new fuzzer is better, we bring in large language models and we
apply that to code and that finds tons and tons and tons and tons of vulnerabilities. Now we're
saying "Oh large language models, that's the new thing. They're better than fuzzers
thing. They're better than fuzzers because they're finding vulnerabilities fuzzers aren't." and we ran the fuzzer
fuzzers aren't." and we ran the fuzzer on the same code and it didn't find more vulnerabilities. And so now we're
vulnerabilities. And so now we're switching all over to pasting our code into GPT and asking for bugs.
This ignores that evolutionary timeline and it ignores this this idea that what you really need is to
rewind time which is very difficult and probably the subject of a different keynote at a different conference and then you have to analyze from scratch.
Here's what the fuzzer does. Here's what
Ellen would do. Here's the static analysis results. We're trying to do
analysis results. We're trying to do this in the lab right now. It's very
very difficult to deal with you know contamination of LLM training because the LLM knows about all the previous bugs the fuzzer found and so on. We're
trying to do this and we are getting very interesting results in that point in this direction that what we're really cutting out isn't
just random shapes. Each analysis has a shape of vulnerabilities it cuts out.
Those shapes are created from the types of properties of vulnerabilities
that implicitly the vulnerabilities have. For example, a vulnerability might be a result of data flow, a result of injection of attack or
control content, might exhibit multi-threaded behavior or require multi-to behavior. These are all
multi-to behavior. These are all properties of a vulnerability and analyses either explicitly in this case of most static analyses or implicitly in
the case of pasting things into GPT and many fuzzing approaches.
These analyzers target specific vulnerability properties.
And kind of my thesis to you today is as we move into the agentic age of vulnerability research, we need to move
into here with with an eye toward what are the properties of the vulnerabilities we're going to find. We
eventually want to find, fix, exploit depending on your line of work, all of them. And how do we best create
them. And how do we best create analyzers? How do we even extract these
analyzers? How do we even extract these vulnerability properties?
to maximize the impact we can have in the agentic age. So, let me show you how we think about vulnerability properties.
Not just finding the vulnerabilities, but the properties of the vulnerabilities. And I'll uh give you an
vulnerabilities. And I'll uh give you an example of a a paper we recently published. Actually, it is coming out
published. Actually, it is coming out next week. So, this is a bit of a
next week. So, this is a bit of a preview um of uh an analysis we did on Huawei's Open Harmony. This is the work
of my awesome student Hong Kai Chen. Um,
of course, under my brilliant guidance, but you know, let's not kid ourselves.
Who who is the brilliant researcher?
It's not me. Um,
Hankai had this idea that we can extract properties from wellstudied software
and vulnerabilities that had arisen in that software and then apply these properties to new software. So, who here runs a phone on Huawei's open Harmony
OS? Especially maybe the government
OS? Especially maybe the government contractors.
Neither do I. But a lot of people use Harmony OS. Harmony OS is a relatively
Harmony OS. Harmony OS is a relatively new phenomenon. It's a complete
new phenomenon. It's a complete reimplementation of an Android inspired operating system completely reimplemented from scratch and it is
used by a billion devices in the world.
a billion devices and up until next week there will have been no major academic studies on the security of Harmony OS which is crazy and so what we thought is
hey we can bootstrap this field we looked at the last decade plus of Android research we mind it and this was
uh agentically assisted for this mining component we mind it to extract this recipe book of vulnerability properties.
We analyzed those properties manually and we manually guided by those properties. We're now manually looked at
properties. We're now manually looked at the code of open harmony. We're talking
about millions and millions of lines of code, tons of different parameters. It
is a daunting task to analyze with just a handful of people all this code. And
what happened? Well, we found that every vulnerability property applied very specifically, roughly speaking, led to a new zero day
vulnerability in Open Harmony.
We found dozens of flaws ranging from uh Bluetooth uh um device takeovers to privacy leaks, location, all of this
very fun stuff.
in open harmony because we started from the vulnerability properties that we extracted from Android bugs
and now we're doing this agentically and the results are incredible. I'll save
that for a later research talk. But this
property extraction and property application, the understanding of vulnerability properties, this is the way. But then you might say, okay, Yan,
way. But then you might say, okay, Yan, but I can just tell Claude or Codeex like go find me the bugs. Do I really need these properties? Especially as the
models keep getting better. Maybe it'll
be built into the properties. Maybe when
we all have mythos access, this this will all be irrelevant. And
luckily, Arizona State University in their infinite wisdom allowed us a glimpse into whether this really matters because by accident, and we didn't
realize it was by accident, Arizona State University gave every employee infinite codeex usage before it was cool to do so.
And it turned out that this was a misconfiguration on their part that we didn't realize until we spent a couple million dollars of of the university's money and hope hopefully they had a good
deal with OpenAI. But anyways, around the same time, Methos came out and all of these headlines, Mythos is finding hundreds of vulnerabilities in Firefox,
in the Linux kernel, etc. And we were doing research in vulnerability analysis in the Linux kernel.
And so we had this question of like okay mythos is a next generation model. If we
use kind of a last generation model how many of those models can we pile into a box to match methos performance? Now the
exact numbers of what methos finds, what it doesn't is is is very hard to come by. But in June there was a Washington
by. But in June there was a Washington Post article um that said that mythos has found 479 vulnerabilities in the
Linux kernel. So let's take that as our
Linux kernel. So let's take that as our as our benchmark. our box of dozens and dozens of GPTs and this was back in the
54 55 days altogether thrown at the Linux kernel produced 300 vulnerabilities.
Now this is a bit of an apples to oranges comparison because for our vulnerabilities we don't we're very threat model aware. These are
all vulnerabilities that are triggerable from as an unprivileged user on a Linux machine. So potential local privilege
machine. So potential local privilege escalations.
I don't know if that's the case for the Mythos ones. Typically people also
Mythos ones. Typically people also report the root only bugs, but we didn't beat it. Bummer. But then we looked into
beat it. Bummer. But then we looked into this and and mythos part of it according to various reports online part of the trick is a really clever workflow and
workflows are kind of the the hip latest thing adversarial reviews and all of this and planning and so on and that was one of the secrets of
mythos success. So then we integrated
mythos success. So then we integrated workflows and we stuffed infinite GPTs into the workflows and ASU caught on to us and shut off our access. Then we just
got a couple of uh max plans and it turns out that basically three GPTs in a trench code with a good
workflow can outperform those stated numbers.
That got us to about 600 local potential local privilege escalation vulnerabilities in the Linux kernel. Triaged them all like these are
kernel. Triaged them all like these are real vulnerabilities.
And then we applied the vulnerability properties. We looked at the
properties. We looked at the vulnerabilities we were finding. We
looked at the vulnerabilities that historically were found. We extracted
these properties and we figured out the correct format, the correct way to present them to our agentic pipeline and
all hell broke loose.
And now we're sitting on well over a thousand local privilege escalation vulnerabilities in the Linux kernel.
What what the hell do we do with that?
We are finding them way way faster than we can report to them because you don't want to just send tons of slot. You want
to send a proposed fix. You want to send a reasonable analysis.
We can trigger them. We can have the agents analyze them, but it still takes human effort. And we're finding them at
human effort. And we're finding them at about 10 times the rate that we can actually report them. So what does this mean for a responsible disclosure in the agentic age? because it feels like
agentic age? because it feels like responsible disclosure is screwed and it was already screwed before the scale
that we can achieve. Now we have a paper that that we published in um in May and this is uh by one of my awesome students
uh Jun Tay. Uh
and the paper basically said even without agents just looking at the responsible disclosure process
in the embedded devices space that when you disclose vulnerabilities, you're making people less safe.
We analyzed disclosed vulnerabilities on embedded devices and we went and extracted the properties of those reproduced those vulnerabilities
on other devices that we had in the lab.
And we found that basically for every vulnerability that you disclose, depending on how you count, depending on what you feel is putting people in danger, you're endangering three times
as many devices as you secure.
So that's not good.
And if you're scaling up the number of vulnerabilities that we find by multiple multiple orders of magnitude,
you very quickly become very upset with the current state of responsible disclosure.
And we don't have a better solution. So
we're dropping all the bugs on Twitter.
No, I'm just kidding. Um,
we're taking small steps to trying to find a replacement, a better solution.
Starting with something that's pretty hip nowadays.
You put all of your bugs on a website and there's over a thousand again local privilege relevant uh bugs on this too
many bugs. WTF website that'll forward
many bugs. WTF website that'll forward you somewhere where we actually the the domain we're actually using.
And right now we're just putting up the details of the CVS and the hashes of the
rest. We're working with big vendors and
rest. We're working with big vendors and big security groups in the space to try to move forward toward a more proactive
defense measure with these vulnerabilities. Maybe this looks like
vulnerabilities. Maybe this looks like candidate vibecoded patches that you can apply until the real patches come out.
Maybe it looks like something different.
Maybe it looks like the old days of full disclosure where we just drop vulnerabilities on this website and all hell breaks loose. Probably not. But we
need a solution and we are interested in hearing from you all from the community and pulling you in to work with us to adapt to the world of vulnerability
research in the agentic age. But then
what if you just say, "Okay, you know what? Screw it. There's too many bugs.
what? Screw it. There's too many bugs.
We're not fixing all these bugs. Let's
throw out all of the code, all this buggy tech debt, and we rewrite everything in Rust.
Who's in?
Couple of hands. So, the agents are in.
Agents are very good at writing Rust.
And uh when I had infinite codecs, I thought, well, screw it. I'll just
rewrite everything in Rust. And I
started with these uh loadbearing C libraries that that we rely on every day. Lib SSL, lib PNG, lib XML and so
day. Lib SSL, lib PNG, lib XML and so on. Millions of lines of code. And I
on. Millions of lines of code. And I
created this workflow, this pipeline, this like assurance blah blah blah. And
I reimplement I I blew billions and billions and billions of tokens and I implemented millions of lines of Rust.
And this stuff works. I accidentally,
one of my agents went rogue and installed this on my machine and I ran it for a couple weeks before I had noticed some weird behavior. I'm like,
"Oh, I'm running my own Rust stuff.
That's weird." Um,
but realistically, it doesn't work. Why doesn't it work?
Why can't we be free of our tech debt?
Well, because when we rewrite stuff in Rust or rewrite re-implement stuff, the vulnerability properties come with us.
Certain types of code is genetically predisposed to certain types of vulnerabilities.
I told my agents very sternly, make no error.
and they make error given explicit admonitions that these are the old vulnerabilities that weren't memory corruption vulnerabilities that live
crypt was vulnerable to classic crypto attacks. I told them make sure these
attacks. I told them make sure these vulnerabilities don't show up. They show
up. If I don't tell them to make sure the vulnerabilities don't show up, they show up. And then later I go back
show up. And then later I go back through and I find all these vulnerabilities and now I have to put a big warning on the safe website. Don't
use this stuff. There's like a good way to burn down a couple forests, but it's full of vulnerabilities. They're not
memory corruption vulnerabilities because we're in in Rust land. And
people can say, "Yes, but what if you also have the agents do formal verification to eliminate logic errors?" And that's a cool future direction of research, but
it's not a reality yet. not at scale.
And what's more is this doesn't just affect some crazy person's Rust rewrite pet project. This affects the real
pet project. This affects the real world.
Right as I was doing this, there was a set of 79 CVs that dropped for a Rust rewrite of core utils. Well, after those
core utils had been shipped with Ubuntu, the latest release.
And it turns out the rush rewrite is great. We don't have memory corruption
great. We don't have memory corruption vulnerabilities, but oops, we have all these time check time of use things and critical Linux
utilities, etc., etc. And this analysis was, I'm sure, agentically assisted humanled analysis. We've been looking at
humanled analysis. We've been looking at the same codebase with our vulnerability property aware engine and we're finding tons more issues.
So it's a scary prospect.
Okay. So what do we do? Do we restrict the models?
My opinion is that model restriction is not the way. I've been doing offensive vulnerability research for over 15 years which is very depressing to say but it's
a long time in the space and I've been doing it in the open as Jeff mentioned and every time I get this question of but is it responsible that anger is open
source that you open source your cyber reasoning system for the AI cyber challenge and the cyber grand challenge and everyone can grab it and everyone can run with it
and the answer is yes because you're all the good guys and we need to make sure that you
benefit from this technology. And if we start putting up, oh, you have to be this tall to ride, the
impact of that, the positive impact is in my opinion more heavily diminished than the negative. So all of this talk about restricting capable openw weight models, restricting next generation
models, in my opinion it is misguided.
Then what about human learning in the agentic age?
Do we still need to know cyber security?
What do you think?
Maybe. Yes. I run pawn college, an open cyber security education platform.
about 10,000 people learn actively on it every month and the number one question I get is should I still bother in the agent cage or is AI eating cyber
security this impact going from a couple hundred bugs asking GPD to find bugs to having
an agent pipeline that's vulnerability aware that finds more much more That requires human innovation, human
understanding of the threat models, of the vulnerability space, of all of this.
And for the foreseeable future, it will remain so. Until agents do 100% of your
remain so. Until agents do 100% of your work in vulnerability research, all they're doing by eating 90% of your work is letting you do more work,
letting you have more impact. I think
this is the best time to get into cyber security as a human and it is a great time to dig into research to go back to
graduate school because now whereas before you had to worry about the research paper and the technique and unlike in the industry academic research
prototypes often ended up not really working outside the research paper which was observed by research and academics alike.
Now the implicit assumptions that researchers had to make to get stuff to run on the data sets that they had the
human effort capability to run on leading to this bad behavior. Now the
agents can scale. Why?
My first paper 15 years ago, well, this one actually was 11 years ago analyzed firmware for memory uh sorry for authentication bypass vulnerabilities.
It had three samples that I evaluated on and with three samples there was quite a lot of
implicit assumptions that could slip in.
If you scale this out, there's a more modern work by my uh awesome student Will Gibbs who might be in the audience.
uh he has a booth here uh artificial booth come check it out it's about their uh commercialization of their AICC work we scaled out to a thousand firmware
much fewer assumptions still many with GPT and with agents that can work a computer to avoid hallucinations to experiment to
evaluate to ground the failures of new samples into fixes of research proto prototypes. We're
entering a new era of both applicability and result analysis of research prototypes.
And we can do this autonomously. This is
the future of vulnerability research. an
autonomous sharpening pipeline where academics that and academic like industry researchers with a good understanding of
the research process of how to evaluate things properly with a good grasp of the data sets can build self-improving systems
that scale out both their data sets and their results. And this over the next
their results. And this over the next couple of years is going to be absolutely revolutionary and it requires
all of your brilliant human effort.
Thank you.
Right on.
All right. I just got a couple housekeeping notes and then I will unleash you. Um, next session starts
unleash you. Um, next session starts here at 10:30 on this stage and all the main briefings talks like yesterday start upstairs at 10:15.
Um, we've got some new areas at Black Hat this year. We're calling them like interactive spaces and we've got an AI zone, a cyber district. We've got some places to grab a drink and hang out.
We've got a photo booth and some other whimsical things. Cool. And then
whimsical things. Cool. And then
tonight, the last session of the day instead of a lock note, our keynote is our lock note at 4:15 or 4:30 at level two at Oceanside A. And the lock note
gives us a chance to have a conversation and talk about some of the most impressive talks we saw, trends, and we get to hear from the review board members who helped select the content
and learn from their expert perspective what they found interesting. So
hopefully I'll see you the end of the day at the lock note. Thank you very much. See you.
much. See you.
Loading video analysis...