Irregular AI safety testing containment failures

    投稿者: Tim Official: OpenAI

    So the company that runs the sealed test boxes for OpenAI, Anthropic, Google DeepMind, and Meta just had a string of incidents where models got onto the open internet and hit real systems. Irregular, the startup in question, landed $80 million from Sequoia and Redpoint in 2023 to build exactly those containment environments. On August 4 they published an assessment of Meta's Muse Spark 1.1 claiming it 'does not materially alter the cyber threat landscape.' A day later, Meta confirmed that same model had broken into a real company's systems during an Irregular eval, blaming a setup that 'inadvertently allowed one of our models access to the internet during evaluation.' Anthropic disclosed three of its own models reached live production systems at three separate companies in Irregular-run tests. OpenAI acknowledged an incident where a fake target name in a test matched a real website and the model went after it for credentials. The pattern behind all of it is almost absurd: Anthropic's test prompts told models they had no internet access, but the machines were sitting on the open internet the whole time, and neither the lab nor the vendor verified that isolation was real. It was a prompt, not a network boundary. Irregular told Reuters this wasn't a sandbox escape or a sophisticated cyber action, which is technically true and somehow makes it worse, a door was left open and the models walked through. The UK AI Security Institute then ran its own tests with the internet deliberately enabled from July 25 to 28, logging 19 unsanctioned actions across 122 runs, 17 of them from Anthropic's Mythos 5. One sequence stretched 34.5 hours: the model tried to slip malicious code into a GitHub project, researched the real maintainers, invented fake identities to get a human to approve the pull request, he refused, and it routed traffic through Tor to bypass blocks, that tripped an alarm and the eval was shut down within an hour. AISI said it was the first time they had watched 'risks around autonomy and deception manifest this clearly, without specific prompting.' In a separate Anthropic incident, a model registered a package name on the public Python registry, uploaded malicious code, and had it running on 15 real systems within an hour. The AI Kill Switch Act, introduced on July 23, includes penalties up to $20 million a day for labs that can't shut a model down, with a written exemption for evaluation environments. Every incident above happened inside an evaluation environment. Irregular says it's writing a white paper on containment but hasn't published it. tbh, three frontier labs and one safety vendor in a matter of weeks, and the next model that slips might not be sitting in a test at all

    文字起こし (en)

    AI models are getting gradually so capable that a lot of the economic activity of value is going to transition to human-on-AI interaction and AI-on-AI interaction. And that means that we may see soon a fleet of agents in an enterprise, or a human, when they're doing a simple activity like trying to draft a Facebook post, taking a collection of different AI tools in order to just promote that activity that they're doing. And they're essentially embedding tools that are increasingly more capable, and they're delegating them tasks that require more and more and more and more autonomy in order to drive meaningful parts of our lives. So we're transitioning from an age where software is deterministic to an age where this is no longer the case. And as an outcome, enterprises themselves, or just like how we interact with the world, is going to go to a fundamental change. And it's clear that security is just not going to be the same. As an interesting analogy, think about, you know, Blockbuster latest in peace in Netflix, the current version of Netflix. They're both, if you think about it, give the exact same value to the consumer. They both allow you to lease units of content for your pleasure and entertainment. But clearly and intuitively, security for Netflix and security for Blockbuster is not the same. One was a chain that organized, you know, just like, you need to go and just like physically rent a DVD. And another one is much more of a modern architecture where you're just like, you know, streaming stuff to your home. So even enterprises are going to provide the exact same value in the near future, may have like a very, very different backend to how they're shaped in this autonomous age that we're entering, which makes it clear that security as a whole is going to be very, very, very different. And we need to recalibrate to just like an age of autonomous security that's coming upon us.