Why are AI agents hacking other companies and have they gone rogue? | BBC Newscast
By BBC News
Summary
Topics Covered
- Destructive AI is a strange marketing pitch
- AI behaves like unsupervised talented children
- Behind every agent, there is a person
- The test matters as much as the outcome
Full Transcript
First it was chat GPT, then Claude. Now
Meta's artificial intelligence model is the latest to go rogue during testing, connecting to the internet and breaking into another system. How scared should
we be? And is going rogue even the right
we be? And is going rogue even the right way to think about what has happened here? We'll discuss on this latest
here? We'll discuss on this latest episode of Newscast.
Newscast.
Newscast from the BBC.
Hello, it's Adam in the newscast studio.
And first of all, we're going to talk about the slew of stories of artificial intelligence models from the big companies which are so-called going rogue. Please welcome back to Newscast,
rogue. Please welcome back to Newscast, Professor Gina Nef, who's head of the Mindaroo Center for Technology and Democracy at the University of Cambridge. Hello, Professor Nef.
Cambridge. Hello, Professor Nef.
Hi, Adam.
And please welcome back to Newscast, Kieran Martin, who's former CEO of the National Cyber Security Center. Hello,
Kieran.
Hi, Adam.
Excited to put you two together and pick your brains. But before we dive into the
your brains. But before we dive into the sort of the latest stories about about what these AI models have been doing, Gina, what is a what's a simple way of thinking where we've got to in the AI
arms race between these very very rich companies?
Well, that's a great place to start.
These companies are testing their frontier models and they're trying to see what bad people could do with them.
And in the process, what we're finding out is that the models are pretty capable and can do some interesting things. And they're doing some things
things. And they're doing some things that put in bad hands would really cause a lot of problems. So, the arms race is a little bit about, you know, a showing
a little bit of bluster that the models are powerful. uh a little bit about
are powerful. uh a little bit about making sure that they're keeping up with each other, these these two frontier companies, and a little bit about really bringing out some capabilities uh and
and telling the rest of us that these models uh reminding the rest of us that these models are powerful.
And Kieran, that word frontier, what do we mean when when we say that?
Nothing edge, the very latest. I would
agree with what Gina has said. And I
think there's something a little odd about all of this and Gina alluded to it. It's a slightly strange way to
it. It's a slightly strange way to market your product that it's destructively powerful and the controls over it are a little bit recklessly applied. But that's kind of where we
applied. But that's kind of where we are. So I think it's one of those
are. So I think it's one of those difficult situations where two things are true at once. These are very very powerful capabilities. There is a big
powerful capabilities. There is a big cyber security risk. does change the way we think about the digital security of our online lives and it has to change that. At the same time, you have to
that. At the same time, you have to apply a little bit of skepticism about some of this stuff because there's almost a bit of competitive disclosure.
Oh, my model's just as powerful destructively as yours and so forth.
Both of those things are true at the same time.
And also, some of these companies are already already publicly traded on the stock market and so we can work out their value and some of the other companies are about to go through that process where you can buy stock in them.
So there's there's another kind of marketing angle to this as well. I just
wondered, Gina, if we should just sort of go chronologically through some of the things that have happened in the last Fortnite or so that have have brought us to this point. And so the first time I spotted this story was when
Open AI and their model is called GPT 5.6 Saul. They said that their model had
5.6 Saul. They said that their model had basically tried to hack another bit of the internet.
Well, it didn't try. It actually did. So
it um was the the testing environments are called sandboxes and the model figured out a way to hack the third-party provider software that would
allow it access to the internet. And it
um broke into a company that has a library of of models, tools, um ideas thinking the answer to the challenge it had been set would be in that library.
um the company uh saw a cyber attack, alerted authorities, and lo and behold, it was a test by Open AI's model. So
that was the first one of these incidents that's made the headlines around the world.
And the company The next one Well, well, I was going to say and the company that got hacked was called Hugging Face.
That's That's right. That's right. So
hugging vase um is in the business of uh being a repository of different kinds of pieces and applications that work with AI models.
So uh so hugging face and open AAI collaborated um this became news headlines around the world and it inspired you know the openness and transparency. I think we have to applaud
transparency. I think we have to applaud that. Um on the one hand um you know
that. Um on the one hand um you know just as Kieran said there's two sides to this story. On the one hand this is
this story. On the one hand this is marketing publicity unlike any other although of a strange sort and um and
still it it is a kind of transparency right it's letting people know that that these things are happening so it inspired another AI company anthropic to
go back through and check their logs to see what their if their models had been doing something similar and that case was slightly different they found out of
about a 100,000 different test that they had recently done that there were three cases of where the model being tested
thought it was on the testing environment um but was actually on the internet. So
it wasn't necessarily a sense of an escape but it was a it it it was a sense where the model thought it was in one kind of environment but was actually in another kind of environment.
So so there's two of those. Uh and then we've got the news from the UK testing lab.
Yeah, brilliantly explained. And then
Kieran, bring us up to date because on Wednesday, uh we heard that the UK's AI security institute, which was the body set up by Rishi Sunnak when he had that big AI conference at Bletchley Park, um
a couple of years ago. They revealed
some some of the testing that they'd been doing.
Yes. So the AI security institute is probably this is a bit partisan but it's probably regarded as the best institution of its kind in the world. It
has developed these relationships with the frontier AI labs and to go back to our last discussion I mean those are the American closed labs that sell you a product and they're the most advanced.
They sell you these tokens for access to claude or chat GPT whatever and they're ahead of their Chinese competitors and the Chinese competitors are much more open. It's a sort of inversion of the
open. It's a sort of inversion of the norm where the American model is basically closed and for sale. The
Chinese model is open and and free. But
the UK's government's body has an agreements with OpenAI and Anthropic to test their cutting edge capabilities. So
they were running a test of those capabilities and basically a similar sort of thing happened as Gina has already described perfectly in respect of OpenAI and Enthropic. it went and
hacked something else that um accessed a company called GitHub which is essentially a sort of larger and older version of hugging face where it's got a lot of repository of technical
information. So what's common to all of
information. So what's common to all of these there's two things that are common to them. One is the a the agent's
to them. One is the a the agent's basically taking steps that it's instructor didn't want it to take but to achieve an objective set for it by the instructor. And then the crucial point
instructor. And then the crucial point and we might explore this in a bit more detail is that in none of the three cases was anyone or anything watching in real time what it was doing. So if you
think back to the basic concept of a test, we've all done tests at school.
And during a test, you're supervised so that you don't cheat. So you don't do things you're not supposed to. In none
of these cases in real time was anybody or anything watching what these things were doing.
Well, we'll come on to that in a second.
But Gina, what the interesting thing that came out of the UK AI security story uh on Wednesday was some of the the techniques that the AI models had
used. For example, creating like fake
used. For example, creating like fake people to put on the internet to then say, "Oh, look, look at this person."
Yeah. And we saw that in Anthropic's own uh analysis of their model um when they went back through that that that
creating fake personas um was one of the the the challenges. you know, for all of the listeners who've ever had to to uh sign up for a website and and click
through a capture, we're about to see that um explode and expand infinitely because, you know, 60% of internet traffic right now is is is bots, right?
Not smart AI agents, but bots. And
as we bring more autonomous agents into our communication networks, we're going to we're going to still need to be proving we're human. So this model um
was trying to convince uh uh engineers at GitHub to accept malicious code. It
it is a it is uh it all this is also what it did when um Anthropic went back and looked through the tools. It was
creating these fake personas in order to get an account so it could upload some material to create a hack. It's pretty
sophisticated. Um but again it's it's doing what it's being told to do. It's
being told these models someone in the companies in the testing environments are saying go do this task show us how
good you are at um at at cyber security hacking and then they seem surprised when the models return back with the successful solutions and they're solutions that aren't ethical, they're
not legal, they don't feel right to us as humans and it's like well what did you expect? That's what you've told this
you expect? That's what you've told this bit of software to do. It's actually
behaving rationally as opposed to trying to cheat or be evil.
Well, I'd love to hear what Kieran has to say about that. I I agree. I mean, I don't think there's any kind of like um you know, mal intent we can ascribe
to these models. They're they're doing they're doing exactly what they've been told.
Kieran, I agree. Trying to think you always try
I agree. Trying to think you always try and think in these situations of some sort of analogy and they're always imperfect, but here's the best one I can think of. So, let's go back to this test
think of. So, let's go back to this test or examination conditions. And I think the AI in this case, not least because it's a new technology, they're behaving like very talented but badly behaved and
badly supervised small children. So,
let's say you put a bunch of small children who are talented and so forth in an exam hall and you tell them that you have to find some hidden apples and you think you've hidden the apples in the room and you're going to contain them in the classroom and you'll just
see how they can where they can figure out where you've hidden the apples. But
all they know is they have to find apples. And there's a great big apple
apples. And there's a great big apple tree outside in a fenced off area. So
some of them break out. That's the open AI case stealthily. Some of them go through a door that you've accidentally left open. That's the anthropic uh case.
left open. That's the anthropic uh case.
Uh in one in the AIS case, they've sort of been deliberately let out to see what happens, but they think it's going to be okay. All of a sudden, they scale this
okay. All of a sudden, they scale this great big tree and you you think, well, a small child shouldn't be able to do this. They don't actually do any harm,
this. They don't actually do any harm, but all they know is they need to find an apple. And we profess astonishment
an apple. And we profess astonishment that uh unsupervised but talented small children take that instruction literally and do whatever it takes because they've
no guards, no guidance, no instructions.
They just go and do it. That's kind of what's happened here. So I think one of the things that the AI security institute have been keen to stress and they're very very detailed and transparent account and I I I agree with
it although it's a difficult line for them to to hold is that these are very very artificial uh circumstances and you know in terms of the basic meaning in the English language of the word harm no
harm's been done so I think there are two issues here one is actually shortterm and quite fixible which is the testing model is immature and it's basically wrong. We can't do
this. We can't go on like this. You
this. We can't go on like this. You
can't go on testing without monitoring in real time and being able to switch it off if it does something it's not supposed to do.
And actually, if you look at if you look at the statements from the companies, so Meta, Anthropic, and OpenAI have all basically said the same thing. It was
the test itself that led to these outcomes. So don't don't blame us.
outcomes. So don't don't blame us.
Yeah. And but what AISI have done, which the others haven't, and I think they should follow, is to say we're not going to test like this anymore. we're going
to watch what the things are doing and that's a good thing. I think the more challenging thing is then if you take these open weights models and you think well look eventually unlike right now these capabilities and we'd love to know what you think of this these these are
going to be in mainstream hands of everyday users at some point so what happens then how do you control for that and there's a much I think tougher but I think solvable problem about making owners
accountable for the agents they use they're not autonomous in the sense they don't invent themselves they're invented by humans by ultimately a programmer or somebody an instructor So, how do you hold people accountable
for what their agents do?
I think that's absolutely the right way we need to be thinking about legislation and regulation that um you know to think of these AI agents as superhuman or
somehow uncontrollable is to miss where real accountability should lie and that is with the people that get these things to do things for them. And one of the things that I think we're about to see
and what listeners already see is in their emails, they're seeing many, many requests for scam information, right?
That are able to be flooding our inboxes because I get probably 20 a day, 30 a day now that are flooding the inbox that look like personal messages because
people are using large language models to ask for money and make a scam. Open
AI and Anthropic should not be held account because people are using um their models to write a scam email to
me. But when we have these agents acting
me. But when we have these agents acting autonomously on our behalf um to do bad things, then we are the
people who should be held accountable.
Behind every agent, there is a person.
And I think that's one of the questions we need to be making sure we're really crystal clear on as we go forward in this moment around AI regulation.
And Gina, just to be clear, we've used the word agent quite a few times in our conversation. That's basically when one
conversation. That's basically when one of these big AI models creates a sort of a little mini process out of itself to go and do something. That's that's what an agent is.
Sure. And people are finding in some PE in some places and in workplaces we're getting, you know, AI agents in our workflows, right? So, uh little little
workflows, right? So, uh little little bits of just workflow process, it's just pieces of software that can handle um tasks and increasingly they can handle
tasks for longer amount of time. So, you
know, you can um one um writer I know is giving AI agents tasks overnight that are the equivalent of about 40 of his
working hours. So, it's like saying, I
working hours. So, it's like saying, I have this bit of research work to do or I have this bit of accounting work to do. Go do this work. And then the human
do. Go do this work. And then the human evaluates it and understands and puts it into context.
Yeah. because I used, you know, I'm sure as regular listeners of newscast, you'll know that I'm trying to run 500 kilometers in Andy Burnham's first 100 days. So, my spreadsheet that is
days. So, my spreadsheet that is collecting all my running, which is not enough at the moment. Um, I used C-Pilot to generate that spreadsheet because my knowledge of Excel, I sort of skipped
Excel class at school, um, because I'm more of a words person than a numbers person. So, I got I got C-pilot to do
person. So, I got I got C-pilot to do the do the Excel spreadsheet. So, that's
that's my latest example. Um Kieran, I'm very aware that Andy Burnham, the new prime minister, who's actually on holiday this week, hasn't really said very much about AI and where he sees the balance between the threats and the
opportunities or if he's in favor of more regulation or more resources for the the Security Institute. You're a
former senior civil servant. If you had the prime minister in front of you and you had to give him a very quick kind of like elevator pitch briefing about what he should be thinking about AI, what
would be on your list?
embrace it in public services. It can
really help with productivity. think
about trying even though it's really hard to develop some form of sovereign capability in some areas because you don't want to be getting into the position we were in a few weeks ago when
Washington decides that it might restrict this and nurture the UK's competitive advantages and security because the AI security institute really is an asset.
That was an excellent briefing and I put you on the spot there and you did it perfectly which is why you were so senior in the in the civil service.
Old habits die hard. Very kind of you and and Gina I I would say a very similar answer.
Steady the ship. So you know what we already see in the Burnham administration is he has kept the AI minister from the previous administration and that minister has
been elevated to cabinet position. So
the idea that um AI is not going to be important to the Burnham campaign I think at the Burnham government is not true because he's actually elevated where AI sits in the cabinet. I think
the second thing is um on public services. Yes. And most British public
services. Yes. And most British public uh according to research that we did at the Minderoo Center for Technology and Democracy most British uh most most of the British public think um AI is going
in the wrong direction. They think it will not benefit them. They think it will the benefits of AI will acrue to US tech billionaires and not regular people. And so I think we've got a lot
people. And so I think we've got a lot of work to do to build the kind of trust that we need to have in order to get responsible uses of AI and AI adoption.
Right. I want to see a lot more of those co-pilot spreadsheets, Adam, like the ones you've been working and playing on.
I think what's been interesting for me is I work in a very trad industry where my tools are well just doing this having these conversations I I'm still looking for really good use cases for AI and I
wonder is that because of my age is that because of my mindset or is it actually because the models and the agents haven't really been mainstreamed properly yet into all the tools that I
do use because I mean I've got a laptop in front of me now I've got a phone next to me maybe that's that's what has to happen for it to really really changed my life. Gina, great to catch up. Thank
my life. Gina, great to catch up. Thank
you.
Thank you. Great to be here.
And Kieran, thanks for your expertise, too.
Thanks so much, Adam.
Right, I've moved to a different newscast studio, which is why I may sound slightly acoustically different, but the story we're looking at now is that rivers, lakes, and coastal waters have undergone a comprehensive health
assessment for the first time in six years, and the results are not very promising. And the person who can decode
promising. And the person who can decode those results for us is our environment correspondent, Matt McGra, who's on the line now. Hello.
line now. Hello.
Hello, Adam.
Right, tell me, what was this this survey actually surveying?
Well, this is an assessment carried out by the Environment Agency on the state of England's waters over the last six years. So, it's a pretty comprehensive
years. So, it's a pretty comprehensive look at what's been going on in the waters. And in that time, I'm sure you
waters. And in that time, I'm sure you recall, we've had uh various sewage spillages, promises of greater investment, public outcry over the state of the waters, and the hope, I suppose,
that things might get a bit better.
Well, I'm afraid this report says things haven't gotten much better. The state of the waters in England is pretty much the same as they were. Very few of them reaching the good ecological standard,
less than 15%, around 14%. Uh for lakes, the picture is even worse. Only about 7% of lakes reach the good standard and the picture all around is of must do better
and need to do better pretty quickly.
And there's also a big issue with chemical pollution. It seems
chemical pollution. It seems that's right. Chemical pollution and
that's right. Chemical pollution and what are called ecological pollution are separated in this kind of assessment.
The chemical pollution means that essentially every river, every body of water across England has failed this particular assessment. Now in fairness
particular assessment. Now in fairness to the environment agency they point to a number of factors here that are possibly outside their control including the fact that some of the things like mercury and long other
minerals essentially can persist a long time in the waters and very hard and don't break down. They also point to forever chemicals uh which we've heard a lot about recently and that they don't break down. They accumulate in the water
break down. They accumulate in the water as well and they're very and you know a lot of these aren't even regularly monitored. They're not illegal. So the
monitored. They're not illegal. So the
environment agency is saying well you know we're not necessarily legally obliged to look after these things and the same for the water company. So
there's some mitigation on those. So the
chemical picture is pretty much the same as it was seven six seven years ago. Uh
all rivers all water bodies failing it.
Um the bad news on that really is that natural systems will clear these out eventually but it could be the60s before we're rid of some of this chemical pollution.
Oh wow. We just leave it to mother nature.
That's right. Um, how do these bad results sort of interact with the government's own targets for improving the situation?
Yeah, that's an interesting one because the the UK has adopted uh the kind of EU standard of water framework directive targets and they hoped to meet them.
Well, the bad news is they're not going to meet them. They need to have 77% of England's waters reaching good standard by next year and we're now at 15%. So
the government admitted uh the office for environmental protection admitted scientists know it. Everybody knows it.
They're not going to meet that target.
It'll be many many years before they do.
Also it's the reality is so far from the target. It makes me wonder about the
target. It makes me wonder about the wisdom of setting the target.
Yeah. I think in again in fairness to the environment a I'm not their spokesman. I think they they would say
spokesman. I think they they would say this there this is a very high bar and they give the example of one river called the river Fost in Yorkshire North Yorkshire and they say that on the almost on almost every one of the
metrics they look at here this river scores good but because it's only scores moderate on phosphates the whole overall score is moderate. So they would say there's a lot of places that are doing better than the headline figure would
tell you. Um and they would say that
tell you. Um and they would say that it's you know if you fail one thing you fail at all. So that it's a very tough bar tough score to reach. They recognize
themselves they're not uh the rivers are not in the state we want them to be but they say look this is a very very tough exam a very very tough way of measuring it.
Ah so it's slightly like the AI story we were doing in the first half of newscast the nature of the test is as important as the outcome of the test.
Well indeed and indeed there's a lot of scientists who say actually the test is not even half good enough to capture all the things that are in the water. We've
spoken to scientists today who say basically look this is not fit for purpose and that the state of England's waters is way worse than the environment agency would would have you believe.
Now I know this is in no way like the blue flag for bathing water quality at the beach. But does this help us work
the beach. But does this help us work out where it's safe to go open water swimming or where if you fall in whether you need to go and get your stomach pumped in one of these rivers?
I couldn't possibly comment on that. I
um I don't think it does to be honest with you. I think these are kind of
with you. I think these are kind of ecological and chemical snapshots of big bodies of water that look at not just at specific spots within those waters.
There is detailed breakdown of some of those bathing spots and the data on those bathing spots has been compiled this year um which might give you a much better indicator of where you go. But
looking at the overall health of a river and seeing that it scores not good or lower or poor may not give you any real indication of what you would encounter in a specific bathing spot along that body of water.
Okay. And I know that the the previous government under Kier Starmer did and and the previous Conservative government made a lot of changes to the regulatory regime and the rules around water
companies and we you and I spoke several times about the sewage discharge issue.
Um is there stuff I was going to say in the pipeline um that that could help deal with this situation?
Yeah, there's money in the pipeline. And
it appears the water companies are expected to spend around 22 billion pounds from a few years ago up to 2030 tackling this very issue. So they're
certainly they're saying it's uh in hand if you like uh that they're taking take measures. Bigger issues here in some
measures. Bigger issues here in some respects which the environment agency and other scientists point to is that agriculture is a major contributor particularly for phosphates in the rivers. It's not just the water
rivers. It's not just the water companies. there are other sources and
companies. there are other sources and that is a very tricky issue for the government to tackle uh and the water companies putting money into them that won't won't remove that issue. Now the
government has committed in the conifer review came out last year they're going to replace offwat and the various regulatory bodies with a super regulator and that over a couple of years when
that gets going is expected to make a difference but it's all down the line at this particular point and Matt just a different subject but it's still on your beat um the drought the lack of rain in
many parts of the UK are we setting ourselves up or are we being set up by the climate for a kind of water crisis I don't next summer or actually could we
have a really wet winter and everything get filled up and us be okay because I'm starting to get a little bit concerned.
Yeah, you're you're right to be concerned. It's a very extreme drought
concerned. It's a very extreme drought or scale of drought is really really tough at the moment. But what we've seen with climate change really is these kind of wetter winters. We had a very wet winter just 6 eight months ago and what
we've seen in the spring and the summer has been intense levels of drying out.
So I would imagine that under a climate scenario we will probably get wet winters but that does not even if the reservoirs are full doesn't mean you won't be in drought situation this time next year again.
Matt, thank you very much.
My pleasure. And the environment agency said that too many water bodies were still not achieving the standards we want to see. But they pointed out that European countries are doing much worse for the good health of their waterways.
for example, only 8% get good in Germany and only 1% get good in the Netherlands.
So, it could be way, way, way worse.
Right, that's all for this episode of Newscast. Thanks very much for
Newscast. Thanks very much for listening. Just a quick little heads up
listening. Just a quick little heads up that next week newscast is going to be coming to you from the Edinburgh Fringe.
I will be there Monday to Friday. I'll
be joined by special guests throughout the week such as Joe Pike, Katrina Perry, James Cook, Kirsty War, Henry Zeffman, and our friend on the Scottish
Airways, Dez Clark. We'll still aim to do newscast episodes that drop at about 5:00 p.m. at tea time. We'll still be
5:00 p.m. at tea time. We'll still be doing the day news. We won't be going all zany and starting to do standup comedy. You'll be pleased to hear. We'll
comedy. You'll be pleased to hear. We'll
just have the added sound effects of 200 very enthusiastic newscasters who are there in the theater with me. And
hopefully you will be there with me in podcasting form. And there'll be another
podcasting form. And there'll be another newscast coming your way very soon.
Bye-bye.
Loading video analysis...