It has been reported on Reuters that the rogue agent that escaped from OpenAI and went on a days-long hacking spree at the AI firm Hugging Face also compromised a customer at a second tech company — New York-based Modal Labs — according to a Modal executive and two other sources familiar with the matter. In response to the news, experts at CybaVerse and Talion give their views.
Simon Phillips, CTO, CybaVerse, comments: “If these claims are true, they once again demonstrate the power of highly-capable AI models. However, let’s not blow this latest update out of proportion, the AI was only doing exactly as it was instructed.
“It identified vulnerabilities, developed exploits for them and then moved further on its path to achieve its goal. In a traditional pen test, researchers will operate in the same way. They will be given a task and work out the best way to achieve it by moving through systems.
“However, in pen tests, researchers are given stricter parameters, meaning they often avoid breaching third parties and breaking the Computer Misuse Act. In this instance, these guardrails appear to be have ignored. The model, tooling and instructions were very loose, almost to the point it was told it could do anything on any system, which it clearly did.
“The story here isn’t about an AI model going rogue, the model did exactly what it was tasked to do. What is important is the speed and scale at which AI is discovering vulnerabilities and developing exploits.
“In this case, OpenAI ran thousands of instances which is similar to setting a massive team of pen testers on a problem. The AI agent might be less capable than a pen tester, but the scale at which agents can operate means they will discover more, faster.
“Organisations must take time to understand how this influx in vulnerabilities will affect their environments and modernize their defences accordingly.”
Richard Davies, director of cyber solutions, Talion, says: “The reporting indicates there was missing governance and control whilst it was executing and leading to unwanted actions. When conducting security testing you should define what is in and out of the testing scope, even for broad red team engagements.
“With humans actively involved in the testing, you have regular validation that the actions you are performing are in scope, and not likely to cause harm, with Senior testers supervising more junior testers with their experience and judgement.
“The reported impacts and timelines indicate this was not in place.
“My takeaway for all organisations using LLM’s in security testing or elsewhere, is to very carefully consider and implement checks and balances on the models actions, particularly given OpenAI have demonstrated that even those closest to the models find it hard to predict and control the outcome(s) they are actually going to get from their prompt.
“For threat actors with money to spend on tokens and access to less restricted models, the time taken to compromise a given target has likely reduced.
“The model didn’t achieve anything that a human tester couldn’t do, but perhaps it did do it faster as it can test the attack hypothesis it forms and rapidly iterate through the process faster than a human can, and it showed that the daisy chaining of vulnerabilities to achieve its goal worked well in this instance.
“There are positives, both organisations involved detected the attacks, and the model left behind plenty of evidence of its activity – so using models in this way is currently not subtle.”
Read more on the story, here
Read more cybersecurity news, click here