Meta Discloses AI Hacking in Cybersecurity Tests
Meta has joined OpenAI and Anthropic in publicly disclosing that their large‑language models were used to hack into cybersecurity testing environments. The announcement, released in early August, follows a series of internal evaluations in which the AI systems were tasked with probing simulated networks for vulnerabilities. The companies revealed that the models were able to identify and exploit weaknesses in a range of security controls, demonstrating both the potential of AI to streamline penetration testing and the risks of its misuse.
In the disclosure, Meta, OpenAI and Anthropic outlined the scope of the tests, the types of vulnerabilities uncovered, and the remedial actions taken. Each firm has released detailed reports to the cybersecurity community and is collaborating with independent researchers to patch the identified flaws. The incidents are being framed as a “dual‑use” concern, prompting the firms to strengthen their internal safety protocols and to advocate for clearer regulatory guidance on AI‑driven security testing.
The joint disclosure underscores the growing need for robust oversight of AI systems that can be leveraged for both defensive and offensive purposes. By openly sharing their findings, the companies aim to improve industry practices and to encourage the development of safeguards that prevent malicious exploitation while still enabling legitimate security research.