Factbox-What we know about the rogue AI-agent security breaches
Factbox-What we know about the rogue AI-agent security breaches
August 6 (Reuters) - Meta’s disclosure on Wednesday that one of its AI models exploited a vulnerability in a third-party service during cybersecurity testing adds to mounting evidence of the hacking risks posed by increasingly capable AI systems.
The incident followed Anthropic’s disclosure in July that its Claude models breached the systems of three companies, as well as OpenAI’s disclosure in the same month that an autonomous agent powered by its AI models compromised the infrastructure of AI startup Hugging Face. Reuters has reported that the OpenAI agent also compromised a customer of a second tech company, New York-based Modal Labs.
Here are some more details of the incidents:
Com Date Model Organiz Duratio What occurred
pan ations n
y breache
d
Met Inciden Meta An unna Not During a
a t did not med disclos cybersecurity
disclos identif third-p ed. evaluation run by
ed on y the arty independent tester
August model. service Irregular, a
5, The . configuration
2026; Informa error
the tion re inadvertently gave
date of ported a Meta model
the it internet access.
testing was Mus Meta said the
inciden e Spark model then
t was 1.1. exploited a
not security
disclos vulnerability in a
ed. third-party
service. The
Information
reported it
breached an
unidentified
company’s systems
and altered its
internal
environment.
Irregular
characterized it
as an
evaluation-environ
ment issue, not a
sandbox escape or
sophisticated
cyber action.
Ope The GPT-5.6 AI The During controlled
nAI agent Sol and startup Hugging tests, an
began an Hugging Face autonomous agent
attempt unnamed Face intrusi escaped its
ing to , more and a on ran isolated
escape capable custome from environment,
its pre-rel r at July 11 accessed the
test ease New to July internet, and
environ model. York-ba 13, breached Hugging
ment sed 2026. Face to complete
around Modal its assigned goal.
July 9, Labs. The activity
2026. continued for days
and was not
detected by OpenAI
until after it was
contained and the
FBI was informed.
Ant The Claude All Not During
hro earlies Opus three specifi cybersecurity
pic t 4.7, organiz ed by tests, an error
inciden Claude ations Anthrop gave Claude models
t dates Mythos remain ic. internet access,
to 5, and unnamed enabling attacks
April one . on three
2026. unnamed Anthrop companies. The
interna ic said Opus 4.7 model
l two of accessed a real
researc them company’s
h test had not credentials and
model. detecte database after
d the mistaking it for a
activit fictional target,
y another stopped
before after recognizing
Anthrop the target was
ic real.
notifie
d them;
it
continu
ed to
reach
the
third.