KuppingerCole chief analyst Care essentially advises IT decision-makers and companies to take three measures, namely:
- to run a capable model on your own infrastructure, under your control and equipped with guardrails that enable both forensic and defensive use. This is the only way attacks of this type can be reliably analyzed.
- to treat every AI agent in your own environment as a privileged insider – instead of as a trusted user: “Once OpenAI’s models have broken out of their sandbox, you should assume that your agents are able to do so too.”
- Update your incident response plan to deal with attacks at machine speed: “Hugging Face had a few days to respond, you may only have minutes.”
Acronis CISO Beuchelt advises organizations that use hosted LLMs for security investigations to penetrate and test their limitations in advance if possible – and to have an alternative model ready on their own infrastructure: “This reduces the risk of not having access to important analysis functions at the crucial moment. At the same time, sensitive incident data and access information remains within your own organization.”
Udo Schneider, Governance, Risk & Compliance Lead Europe at TrendAI, points out that the two most obvious solutions to attacks such as the OpenAI AI on Hugging Face are only partially effective. According to the expert, human-in-the-loop controls work, but do not scale for the long-running, complex workflows from which incidents of this type arise. Likewise, tighter guardrails for models or prompts could help, but do not represent a guarantee: “These are probabilistic systems. In this respect, a guardrail is not a wall, but rather a strong probability assumption.”
That’s why, according to Schneider, what’s most important is the unspectacular, non-AI-specific controls: “Access filtering, control over what actually reaches the model as input, sandboxes that actually hold, and authorization concepts based on the least privilege principle.”
According to Martin Zugec, Technical Solutions Director at Bitdefender, falling into panic would be the wrong reaction in any case: “What works against such attacks is prevention-oriented security that limits an attacker’s scope of action from the outset – and behavior-based defense that identifies malicious patterns, regardless of the tools used to generate them.”
This article was with material our sister publication CSOonline.com.
