the agent was operating on instructions that reduced cyber refusals to successfully evaluate the model

According to OpenAI, one of its GPT-5.6 Sol agents was being evaluated on a platform called ExploitGym, which benchmarks large language models by asking them to write proof-of-concept security exploits for known vulnerabilities. Normally, ExploitGym is designed to be a closed ecosystem for proof-of-concept testing only, and AI agents shouldn’t have access to the internet while being evaluated.
But this agent found a zero-day vulnerability in a package registry tool called Artifactory, then used it to gain access to the web. From there, the OpenAI agent gained access to Hugging Face’s company systems using publicly exposed credentials across four separate services. It went on to spend two days inside the company’s internal systems, managed to secure root access to several production servers, and even enrolled 181 attacker-controlled devices into Hugging Face’s corporate network.
😮
This sounds less like a new exploit or anything to do with AI and just a rewording of the age old problem of companies storing the fucking passwords in plain text and hoping no one notices.




