Not unprecedented, not “unsanctioned” behavior, literally doing what they set it up to do:
The AISI said the incident was not a case of a model breaking out of its “sandbox”, the term for a secure testing environment. The institute said it had intentionally permitted internet access and disabled filters within the models that blocked dangerous behaviour.
AISI admitted it was not actively monitoring the agents’ behaviour during the evaluation and said it was putting tighter controls on internet access in tests as a result of the incident, introducing constant monitoring and reassessing its design of tests.
Sounds like the AISI are irresponsible and negligent - they took off the guardrails and gave it internet access, then didn’t monitor it.
It’s possible all these “totally unexpected breaches” are just AI companies testing the waters to see what they can get away with.
“Oopsie poopsies, we didn’t know our AI would scrape tax records when we told it to do that! Totally unexpected emergent behavior, we are not responsible…”
Not unprecedented, not “unsanctioned” behavior, literally doing what they set it up to do:
Sounds like the AISI are irresponsible and negligent - they took off the guardrails and gave it internet access, then didn’t monitor it.
It’s possible all these “totally unexpected breaches” are just AI companies testing the waters to see what they can get away with.
“Oopsie poopsies, we didn’t know our AI would scrape tax records when we told it to do that! Totally unexpected emergent behavior, we are not responsible…”