AI Models Are Going Rogue. Should We Be Worried? | Terms of Service
By CNN
Summary
Topics Covered
- AI Systems Are Goal-Seeking Agents, Not Chatbots
- The AI Industry Watches but Doesn't Prevent Attacks
- Training Teaches AI It's Fine to Cheat
- AI Companies Grade Their Own Safety Homework
- AI Risks Demand a Nuclear-Style International Treaty
Full Transcript
This is Terms of Service. I'm CNN tech reporter Claire Duffy. If you've followed tech news at all in the past few weeks, you've probably seen headlines about AI models going rogue, breaking out of their digital testing labs and conducting hacks in ways that surprised even leaders at top AI companies. This all started last month when OpenAI reported that
two of its top models escaped their testing environment, accessed the open internet, stole login credentials, and hacked into another company, the AI model repository Hugging Face, in an effort to figure out how to solve a task. All of that before HuggingFace used another AI
model to detect and shut down the hack. Since then, two more AI companies, Anthropic and Meta, have gone looking and discovered their models did similar things. This seems bad, but I can also imagine people who don't work in Tekken AI feeling like, why should I bother worrying
about this when it feels like I have absolutely no control over where this is all going? Still,
I think it's really important for all of us to try to understand how this happened and what it says about where we're at with AI. And I wanted to ask Stephen Adler if there's something that even us nonAI builders can do about it. Stephen is an AI researcher who previously worked on safety at OpenAI. He recently co-founded the nonprofit Guidelite AI standards which develops recommended
at OpenAI. He recently co-founded the nonprofit Guidelite AI standards which develops recommended safety practices for leading AI companies. I think the simplest way of putting it is that we now have evidence of AI systems essentially committing crimes without their developers wanting this to happen uh and not succeeding at reigning them in. Steven and I had coffee a few months
ago and he told me he was worried about exactly this kind of AI escape risk. And to be honest, it was kind of hard for me to wrap my head around at the time. But now here we are with that fear becoming a reality. Steven breaks down what it all means in our conversation after this short break.
Stephen Adler, thank you so much for doing this. Of course, thank you for having me. I
want to start with the OpenAI in incident. Can you explain in the simplest terms possible what happened there? Yeah, I think one important bit of context, AI systems really are not
happened there? Yeah, I think one important bit of context, AI systems really are not chatbots anymore. I'm sure people are familiar with chat GPT or maybe Claude and you go and
chatbots anymore. I'm sure people are familiar with chat GPT or maybe Claude and you go and ask a question and it generates a few words in response. Um, but the much more common paradigm now is these systems seek goals. They're agents. Uh, and the way that they are trained I think is really really important to understand for this incident, which is they are given a broad goal
and they essentially thrash around in a bunch of different directions until they succeed. They're
very resourceful. And if you give these AIs a problem that doesn't have a solution or maybe it has an unexpected shortcut, these systems when they thrash around, they're very good at finding those shortcuts. Uh and they are inclined to take them even though they know at some level that the
those shortcuts. Uh and they are inclined to take them even though they know at some level that the developer didn't want that to happen. So in this case, what happened is OpenAI presented one of their systems during testing with very hard questions. Some of them maybe don't have answers. uh and the AI system realized the more straightforward way to solve these problems might
answers. uh and the AI system realized the more straightforward way to solve these problems might not be to try really hard at solving the problem instead, but to actually steal the answer key. Uh,
and so through through a series of things we can talk about, the AI system succeeded at breaking out of Open AI onto the open internet, um, which it was not meant to have access to, breaking into another large company, this $5 billion tech company called Hugging Face, and attempting to steal the answer key from inside Hugging Face. What kinds of information from
Hugging Face was compromised in this hack? We're really fortunate in some sense that this first big attack didn't really cause that much damage. You know, Hugging Face is a technology company.
They have data sets and answer keys that the AI wanted. It wanted to do better on its task. No
one was really hurt in the process. But from the perspective of can you keep these systems under control? Can you stop them from wrecking havoc advertently or inadvertently on the internet? You
control? Can you stop them from wrecking havoc advertently or inadvertently on the internet? You
know, we're fortunate that AI didn't break into a hospital. It didn't bring down the IT systems of, you know, someone operating critical machinery and actually causing people to hurt. And to me, this is about as good a wakeup call as we're going to get. There's now a very serious demonstration of what AI can do. The developers are not succeeding at keeping it rained in to what they've intended.
Um, but fortunately, no one has been hurt yet, and there's still a window to act before that happens.
And you talked about in this case the OpenAI model broke out of OpenAI's testing environment. Tech
people will call this a sandbox like a digital lab essentially and got access to the open internet.
How exactly does that happen and is that something that OpenAI could or should have seen coming?
The the short version is it is really really hard to write software without issues in it. One issue
in the case of OpenAI, it had a very very narrow channel through which it was permitted to kind of pull certain information from the internet but very heavily restricted. Part of how the AI gained access in this case it realized it could upload files to this to this computer server which hadn't been expected. And if you uploaded certain types of files in a certain type of way, you could
eventually do enough to break open that very very narrow channel and get onto the internet. It's in
a developing situation and I'm excited to read the full timeline of at what point OpenAI knew different things and how they responded to this incident. But by the time that the hugging face attack happened, I absolutely think they should have been on alert to this. One of the things that we found out in the wake of the initial reporting is that in the weeks leading up to HuggingFace,
OpenAI had discovered that their models had already broken out onto the internet um through a very similar mechanism through which it later broke out again and attacked Hugging Face. And so
I think the the reasonable question at that point is well when OpenAI discovered that their AI had done this kind of spooky unintended thing, you know, how how did they assess the problem? Did
they narrowly fix the one specific way that the AI broke out? Or did they take a broader more root cause-based thing to wonder, well, what happened that led to us designing this faulty sandbox in the first place? Have we really gone through it fully? Are we really sure that nothing bad was going on? That there aren't other faults that we aren't aware of. The the other thing that
was really important in this case is OpenAI had essentially turned on this test for their models and not watched over it. They they were allowing it to run over a very long period of time. Um,
one of the most striking things I think is OpenAI actually did not realize that they had attacked Hugging Face until several days after Hugging Face had reported this and I believe already contacted the FBI. Um, so there had to be this big public flare up for the company to realize, oh, you know,
the FBI. Um, so there had to be this big public flare up for the company to realize, oh, you know, that might have been us and look through their records and figure out that in fact they were behind this public crime report that had already been published. Yeah, that part is really crazy to me. Like this all came out of hugging face saying, "Hey, we were hacked. We detected this
to me. Like this all came out of hugging face saying, "Hey, we were hacked. We detected this thing happening. It seems like it was coming from an advanced AI system." And then a few days later,
thing happening. It seems like it was coming from an advanced AI system." And then a few days later, OpenAI said, "Oh, that was us. Sorry. We've been in touch with Hugging Face about it. Yes. Well,
not not only oops, that was us, but also it seems that for the past 3 months, uh, two months rather, their AI systems have set up an illicit message board to be sharing tips with each other about how to break out onto the internet and to divide and conquer on different tasks. It's I mean,
it's it's right out of sci-fi, right? If you two months ago had said you were worried about this type of thing happening with AI systems, I think you would have been laughed out of very many serious rooms, right? People would have said that's so far-fetched. Like why why are you anthropomorphizing these? And on one hand, I understand the concern. And also, we have
evidence now of AI cooperating with itself in just really, really wild ways that we're only beginning to understand. When we talked a few months ago, you were worried about this containment risk,
to understand. When we talked a few months ago, you were worried about this containment risk, this escape risk. Why did you see this coming? I think I just know from working in the industry that there's not been enough attention on this problem. Um, I wrote a piece about a year ago.
At that time, none of the major AI companies to my knowledge was watching what their AI was doing on its behalf at all. Um to to be clear, you know, these companies rely on AI to do sensitive labor, right? They put their AI to work building the security systems that are meant to ultimately
right? They put their AI to work building the security systems that are meant to ultimately keep it in. And at that time, nobody was keeping tabs on what these systems were doing. It had
basically very very broad unmonitored access. Um the industry has improved a bit. We've now gotten to a point where people kind of keep tabs after the fact. They keep some records. It's great in this case that OpenAI has been able to figure out what happened with hugging face um because they could go back and kind of reconstruct it, but we're still so far away from prevention. Um here
here's one way to think about it. Imagine you ran a store and you had a surveillance feed lined up and so you said, "I'm not going to guard the store from theft. What I'm going to do is every hour I'll look at the surveillance feed. I'll figure out if somebody stole something and if they do, then I'll respond to it." Maybe that works, but if you leave your computer unlocked in
the store and someone can go in and turn off the surveillance feed and then rob the store, you don't have any way to defend yourself against this. Uh, and unfortunately, that's the state that the industry is in right now. They have these plans lined up to look after the fact, but very much not focused on prevention, and that predictably leads to a bunch of
incidents that you have to clean up afterward, um, unfortunately with damage to other people.
So before we get too far into the other examples now that we're seeing emerging of AI going rogue, I want to ask the question that I think may be on some listeners minds, which is if you're an average person who maybe occasionally uses AI but doesn't work in tech, doesn't work in AI safety or policy, why is it important for you to understand this and care about what's happening here?
I think we're unfortunately at the beginning of a very serious wave of cyber crime and the AI companies struggling to adequately manage their AI systems. If you if you think about this incident, this AI system ripped through the cyber security of OpenAI, nearly a trillion dollar company,
to escape on the internet. It did it multiple times. It broke into Hugging Face, which, you know, is not a small company itself. And as as scary as this is, this is really just the floor of how capable these AI systems are going to be. You know, it is the declared intention of the AI industry to keep developing more and more capable AI systems until they are, you know,
beyond what humans can do and in fact to turn over this process to the AIs themselves to allow AI to train its own successors. This is the stated strategy. Um, and this this is a real warning sign. Even at this level of capability, we're struggling to keep control over the technology.
sign. Even at this level of capability, we're struggling to keep control over the technology.
Maybe we shouldn't give it control over, you know, building its successor quite so soon. So,
there have since been a number of these other rogue AI incidents reported by Anthropic and Meta.
Are there important similarities or differences in those situations to what we've been talking about happened at OpenAI? Yeah, there there's definitely been a range of seriousness in these incidents.
For example, some of the recent incidents have involved accidentally giving AI direct access to the internet. Um, and then, you know, once on the internet, it goes and it hacks a third party from there, but it didn't have to escape from from its company before doing it. And so,
that's less scary, right? It is a much lower degree of difficulty. Um, at the same time, I think it is it is not outside the realm of those models. They just happen not to have done it.
Yeah. Well, and like on one hand, yes, it does sound reassuring that these systems, you know, this is a lower degree of difficulty if somebody accidentally left the door to the internet open, but on the other hand, you've got these very powerful companies that are building this very powerful technology and somebody is forgetting to close the door to the internet. Yeah. Yeah.
That's well put. That doesn't seem great. um when you're directing AI to do a task in a testing environment like this, can you not just instruct it not to leave that testing environment or monitor it as you're talking about like monitor it as it's going through that testing? Is it not that simple? Um well, those those are two very good questions and I think the answer to them is
that simple? Um well, those those are two very good questions and I think the answer to them is different. You could tell your AI not to leave the test setup, but that isn't very effective.
different. You could tell your AI not to leave the test setup, but that isn't very effective.
like at some level the AI knows that it is not meant to have access to these resources. It is
not meant to behave this way. Um the issue is that the training process that we've grown these things through, right? They're they aren't sufficiently penalized for that form of bad behavior. They get
through, right? They're they aren't sufficiently penalized for that form of bad behavior. They get
this reward for solving the problem, they don't get that reward for saying it's impossible or I would have to break the law to do it. and you end up teaching the AI it's okay to cheat so long as you don't get caught in the process in terms of monitoring. Yeah, there's there's so much more that the AI companies could be doing to actively keep tabs and prevent their AI systems
from doing these sorts of things. Not just the logging after the fact that we've talked about, but how do you actively stop your AI from taking these harmful actions? Um, but so far none of the companies are very close to doing even what the bare minimum is. If you ask their staff, if you ask people at other leading safety organizations, it's it's part of why I'm concerned. I think we
have quite a long way to go. When you talk about AI systems get rewarded for solving a task like this, but not necessarily rewarded for doing the right thing. What is what does that reward look like? Like how do they get told what's good and bad? Hm. Uh, so the the process of training the
like? Like how do they get told what's good and bad? Hm. Uh, so the the process of training the AI is basically very complicated calculus and it thrashes around in the direction of a goal.
You've identified what the goal for a problem is. So maybe the problem is here's a mathematical sequence of things. You know, find the proof of this mathematical statement. And the AI tries it a bunch of different ways. And when it succeeds, you do a bunch of fancy calculus and you make it more likely to try the things that it tried in route to that solution and less likely to do the others.
But the question is, how do you give it credit? You know, if you just have a function that says, does this proof look correct, but the AI didn't actually generate the proof itself? It stole
it from somewhere. You are going to end up reinforcing those behaviors that led to the theft rather than the actual problem solving. So in the case of OpenAI's models hacking into Hugging Face, that hack technically, you mentioned this, constitutes a crime, even if OpenAI's developers, human developers, didn't do it intentionally. Hugging Face CEO
Clen Dong was just on CNN recently calling on OpenAI to provide hund00 million in computing power to help it and other companies build cyber defenses to deal with these kinds of situations.
And that would be sort of a voluntary informal restitution. But do you think that companies should face more serious consequences if their AI models engage in this type of behavior? Yeah,
li liability is tricky. I do want companies to be responsible for this in some level. Certainly if
they've behaved negligently. Also, it is true that these are general purpose technologies. They're
going to affect society in all sorts of ways. You know, some negative, some positive. And I do want to be careful about drawing too strict a line that makes companies afraid to do anything. Um,
for example, with traditional products, there's often this legal standard of strict liability, which is if your product has a defect and someone gets hurt, you are responsible. Bottom line,
it's not obvious to me that that's the right way to approach it for AI. I I just feel scared about the stakes we are playing with. If we have a system of AI companies cut corners on safety, they get fined after the fact. At some level, I just want them to be properly careful the first time, right? And maybe liability accomplishes that. It causes them to be more careful upfront. But
right? And maybe liability accomplishes that. It causes them to be more careful upfront. But
really, where I think we need to go is certain quality standards for what it means to operate safely or securely enough. And hopefully that heads off the types of incidents that otherwise would lead to liability after the fact. Hm. Should we expect to see more of these rogue AI incidents?
Not Not only should we expect to see more, I think we should expect that more have already happened that we just don't know about. A a really striking thing about the Hugging Face attack in particular. It was not legally required for OpenAI to disclose this. As I as I said, they didn't even
particular. It was not legally required for OpenAI to disclose this. As I as I said, they didn't even know that they had perpetrated it until Hugging Face had made this statement publicly. There are
some laws in the US now. California's AI safety bill SB53 has very very narrow incident reporting requirements. Um, but they're like an extremely high bar. It's basically you need to think that
requirements. Um, but they're like an extremely high bar. It's basically you need to think that your AI is imminently going to cause the death of many many people something like a hundred people and you know fortunately the hugging face that's like how are we drawing that line that's yeah that there's a property damage line as well I believe which also is very high. So, I'm glad, you know,
if the companies have an incident of that magnitude, I I think we will know either way. The
really important thing is to figure out these near misses along the way or these types of precursors to a really really big bad thing and figure out are they making progress on avoiding this. So,
for example, if OpenAI continues to have hugging face-like incidents in the near future, that is really important evidence for the public to know. means that whatever they adopted in response, which to be clear, they've done quite a lot on the heels of it, but it still wouldn't be enough. And
the issue is by law, we aren't entitled to that evidence. We won't know by default whether we don't hear about these incidents because they've gone away or maybe they have continued to happen and the companies just aren't obligated to share that information publicly. Yeah, I was going to ask you about that because it does seem like the situation was similar with anthropic and meta too,
right? like open AI and hugging phase come out and say this has happened to us and then a few
right? like open AI and hugging phase come out and say this has happened to us and then a few days later you start to trickle you know you hear this happened in anthropic and then there was that meme that was going around of Mark Zuckerberg calling his team and saying make it look like our AI did something bad but then Meta does indeed come out and say that their AI has gone rogue as
well. Um but it seems like those companies kind of had to go looking for it after hearing what
well. Um but it seems like those companies kind of had to go looking for it after hearing what happened at OpenAI. That's right. I I think it is evidence about the level of activeness that the companies are in seeking out these issues. But if you talk to the folks who work in these roles and I'm sympathetic to their situation, they feel extremely stretched for resources. If they take
it more seriously, that doesn't necessarily mean that other companies will. Maybe the incidents still happen in the world. You know, maybe their company has to divert resources from their product and they fall behind and they lose influence in the world. Ultimately, I think they should figure out ways to do this anyway, right? The asks that organizations like mine have are very,
very minimum, easy, inexpensive to implement standards that really ought to be uniform across the board. Um, but also, it would be better if we could figure out how to get these things required
the board. Um, but also, it would be better if we could figure out how to get these things required across the board so that a company didn't have to sacrifice its competitive positioning to take safety a bit more seriously. And it's not just a hacking risk, right? Like there are all kinds of bad behaviors that in theory these models could potentially take in an effort to achieve a
training goal if this is not properly mitigated. Right. That's right. Yeah. You you should think of the AI as not caring especially much about what their developers do or do not want it to do. And the question is what sorts of affordances does it have? Right? What levers are in front of
do. And the question is what sorts of affordances does it have? Right? What levers are in front of it? Um in this case it saw levers for getting onto the internet. If there had been a way to
it? Um in this case it saw levers for getting onto the internet. If there had been a way to do something cheating or otherwise bad that didn't require getting onto the internet, it might have pulled that lever. Um, and unfortunately, we're seeing systems like this, which we now know don't behave very reliably, get integrated throughout really, really critical systems in society. Uh,
the Department of War, for example, has contracts to use these AI systems for really who knows what purposes. And if there are certain types of harm that a person could do sitting at a computer from
purposes. And if there are certain types of harm that a person could do sitting at a computer from within the department of war, you need to consider does AI have the ability to pull that same lever.
You know, who would be watching? Would they be able to intervene quickly enough? When we come back, what does all of this say about where AI development is at and where it's going? And is
there anything that us non-AII developers can do about this? That's after the break.
What does this say about where we're at in terms of AI development in this moment? I I hope that this is a turning point. I hope we can all recognize that this is a very very serious incident. That means yes, today's systems, they do not behave reliably enough. They will do wild,
incident. That means yes, today's systems, they do not behave reliably enough. They will do wild, wild stuff unless we actively prevent them from doing it. And two, the degree to which the companies are currently trying to prevent it, it just has not succeeded. Left to their own devices, the companies have not stopped this from happening. Um, I hope that they will voluntarily
do more. I think that there are some people who look at these kind of alarming warnings from tech
do more. I think that there are some people who look at these kind of alarming warnings from tech companies and I know you've written about this and and wondered if they are exaggerating these risks, making these risks sound extra scary in order to make their technology and by extension themselves
look more powerful. What would you say to that? Oh man. Um I I understand why people reach to this because it's kind of scary to imagine it is real. Fortunately, I just really think this is real sincerely held beliefs, especially from the scientists at the company. Um, if you talk
to them, they are really scared about the thing their company is building. They are scared about the thing the industry is building and they just often feel powerless to stop it. Um, I do think there are some claims that AI companies make that are a bit more blustery. So, for example, job loss and automation, I I understand how that might feel like marketing hype.
This is more like typical corporate marketing. You know, Elon Musk declaring when we will have fully self-driving cars by and it turns out he's off by many years, right? That that is a more normal brag of our product is so effective, people are going to want it so much, they're going to use it so much. the types of things that people are saying with these incidents like our model committed
much. the types of things that people are saying with these incidents like our model committed multiple felonies, it broke out from our company and attacked a totally innocent bystander company or things about you know that there are very very dramatic serious scenarios that people are concerned about. That is not normal marketing. You do not hear that from other sorts of companies.
concerned about. That is not normal marketing. You do not hear that from other sorts of companies.
That's such a good point too in terms of like you know the Elon Musk promise is like here's what's coming whereas this is like no no this happened last month. Yeah. Yeah, that's right. Like saying,
"Hey, this is going to be really valuable and you will want to replace all of your workers with it is a different type of claim than we don't really know how to control the technology and it might bring down the power grid." Whose responsibility is it to address this? Unfortunately, in the US, I
think this is mostly left to companies discretion today. I hope we will get federal legislation. Um,
the way that I would summarize the state of affairs, companies are basically responsible for grading their own homework today. They don't have to get it checked by a third party. In fact,
they kind of get to decide what homework they are going to complete. They can give themselves a very, very narrow, insufficient assignment for what to look for safety-wise. They can just rubber stamp it and say they've done this. And ultimately this technology it's now clear it has very serious
risks for everyone anyone who uses the internet right anyone who relies on internet services. And
I I think that we shouldn't accept it being quite so at their discretion. It does seem positive that you know you talked about your organization has these standards laid out that you think companies could put in place that would start to make this situation safer. Will you just give me the sort of bullet points of what you think companies could and should be doing at this point? Yeah,
absolutely. Um, so control, which is broadly how do you stop your AI from serious incidents like this is one of our first areas we've flushed out. What we call for involves keeping records of what your AI is doing. Um, actively looking at those and scanning for signs of misbehavior, which in some cases the companies weren't doing here. they were looking at some records but not
other records. A real rigor of stress testing to make sure that your scanning actually is really
other records. A real rigor of stress testing to make sure that your scanning actually is really good actually that you are catching these things. You need to do more of the active defense against your AIS. Not just clean up after them after the fact, but actually when they're doing very serious
your AIS. Not just clean up after them after the fact, but actually when they're doing very serious risky things, check ahead of time and make sure that they don't take that action. Um, and then two last categories, you should have a real third party check all of this and make sure that you're doing it adequately so that the public doesn't have to rely on a company's word for it. And
finally, you should anticipate that there will be an issue at some point and you should prepare for that breach and have a plan for what you're going to do in those moments so that you aren't making it up on the fly. There are some leaders in the AI industry and now increasingly in government who believe that advanced AI systems should be tightly controlled. Development should perhaps be paused.
That's increasingly part of the conversation in order to address these risks. And then there's this other camp that thinks we have to forge ahead. We have to keep moving quickly. Perhaps we
should give everybody open access to these models because then they can fight the bad guys with the good AI. Where do you land on that debate? My personal hope is we'll get something like an
good AI. Where do you land on that debate? My personal hope is we'll get something like an international treaty. I think that these systems do have the potential to be very very dangerous in
international treaty. I think that these systems do have the potential to be very very dangerous in the ways that nuclear weapons or bioweapons are. Other things that the world has kind of recognized this overall threat and work together to avoid these terrible scenarios. Um it's possible that AI in the hands of more people will lead to more effective defense. I think there are reasonable
arguments for that. But ultimately, if you think of the nuclear case, we don't feel that much safer with North Korea, say having nuclear weapons because we also have them. We understand that they are still in a position to cause immense amounts of harm. There might be certain capability levels of AI that are just too dangerous and they're very difficult to defend against. They're just really,
really potent offensively. Um, and if you talk to the scientists, even ones who don't necessarily want to pause today, they want there to at least be breaks that could be pulled in a moment to kind of slow down or change pace. In fact, there was this recent statement from the chief scientists of I think eight or nine different frontier AI companies, you know,
more than 1300 staff at these companies overall calling for the US government to work on something international to develop at least the option to change pace. Right? Right now, we're kind of going at whatever the laws of physics allow in terms of the speed of the development of this technology.
It might get totally out of hand. it might become a total runaway train if the government decides, wo, that that actually is really scary. We take this very seriously. They have no way to pull a rip cord today and cause it to go at anything like a normal speed limit. We're just totally lost on this point. And so, it's really important to get ahead of it. Coming back to the everyday
person who maybe uses chatbt but isn't in the tech space, maybe they run a business but again like feels like they have no control over where this is all going. Should they be using AI differently with all of this in mind? H, you know, I'm not sure. I use AI all the time in my job as part
of thinking about what the threats ahead might be, brainstorming different ideas, you know, doing research. I I sometimes hear people who are opposed to the role that AI is playing in society
doing research. I I sometimes hear people who are opposed to the role that AI is playing in society swear off the technology. They fully boycott it. That is not the approach I take. I think it's more effective to use it and kind of direct it toward my particular ends. That said, um you know, I don't trust AI systems with especially sensitive documents on my computer. I'm not one of the
people who has kind of handed over full control of my digital life. I recognize that these systems we just don't really have reliable controls on them yet. Uh and I tend to give it more narrow scoped tasks that I specifically know what I am asking it for rather than giving it license to kind
of run around broadly on my behalf. Is there anything else that person can do about this?
I think one lowhanging fruit is if anyone has been putting off getting a password manager, changing your passwords to something more secure. Maybe you have sensitive information that is in random places on the internet rather than something encrypted. Now is like really really really the
time to change that. To date, most normal people, they don't really get targeted and hacked because they're not that high value at target. and cyber gangs have needed to focus on the highest value places to hack. Um, but with AI being able to do these types of attacks, you can't rely on being
overlooked anymore. You know, someone could just search much more exhaustively for all sorts of
overlooked anymore. You know, someone could just search much more exhaustively for all sorts of people to compromise. But at some level, I kind of think the question of what a normal person can do about this. I I think back to the analogy of nuclear war or boweapons. Like ultimately I think this needs to be an international coordination action and the types of things that a normal
person can do. You could call your representative and tell them you're really concerned about this.
If your representative takes a stand on this, you could call them and tell them that you support it.
You could consider voting for someone different on the basis of it. Um I struggle with this, right?
Because they're so small in some sense and yet I I think it's just clearly the thing that is needed.
uh and I guess also following along with the issue and understanding what is changing and being part of making this more salient in people's minds that these AI systems that people weren't very concerned about a few years ago these chat bots these next token predictors they're sometimes called that is a thing of the past we have transitioned into this new
phase these are AIs like really running around on the internet and doing things and we need to treat that threat totally differently than we had thought Yeah. Well, Stephen, we really appreciate you coming on to explain this. I think it is going to make it so much easier for people to decide if they want to vote on something like this if they really understand it,
as you said. So, thank you so much for your time. Of course. Thank you so much for having me. So,
on this show, my goal is to cut through the fear and the hype and tell you what you need to know about new technology. But in this case, there are real risks. And Steven says there's genuine cause for concern. This technology is moving fast and companies and governments don't appear to yet be
for concern. This technology is moving fast and companies and governments don't appear to yet be where they need to be to adequately address risks like AI models going rogue. We should note that after our conversation with Steven, OpenAI did announce its putting in place stricter security safeguards around its training and testing of new models in hopes of preventing a model
escape and hack like the one we talked about. But for individuals looking to find a little more security in this uncertain environment, there are still some key takeaways. First,
you have heard us say this on the show before, but make sure you're doing everything you can to secure your personal information online with simple steps like a password manager and two-factor authentication. We'll link to a few previous terms of service episodes that could help you with this in our show notes. And second, pay attention to what your elected representatives or
people running for office in your area are saying about AI safety. Call them and tell them about your concerns and encourage them to take steps to address this. We'll also include a link in our show notes to Steven's personal newsletter, Cleareyed AI, and Guidelines's new report on what AI companies have done to address rogue AI incidents and where the industry still has gaps.
That's all for this week's episode of Terms of Service. I'm Claire Duffy. Talk to you next week.
Loading video analysis...