An open, structured dataset of publicly documented AI-agent and LLM-application security incidents, coded by the role AI plays: weapon, target or surface.
Is there a public dataset of AI agent security incidents?
Yes, this one: 88 publicly documented events from 2023-02 to 2026-09 (29 incidents, 48 vulnerability disclosures, 11 threat reports), one schema-validated JSON record each, coded under a written codebook and mapped to the OWASP Top 10 for LLM Applications, the OWASP Top 10 for Agentic Applications and MITRE ATLAS. Download it as JSON or CSV, subscribe to the RSS feed, or read the repository. The data is CC BY 4.0.
How are incidents coded?
Every event is coded from a primary source that was opened and read, on eight fields defined in the codebook: type (incident, vulnerability disclosure or threat report), lens (the role AI plays: weapon, target or surface), vector (how the attack or failure got in), channel_in, authority (what the AI component could do), channel_out (how the effect left the system), adversarial (whether an attack technique is involved) and outcome (the most severe harm realised or demonstrated). Each record also carries mappings to the OWASP LLM Top 10, the OWASP Agentic Top 10 and MITRE ATLAS, given only where the source supports them. The paper behind the seed data reports two further blind codings and their agreement.
Can I use it commercially?
Yes. The data is licensed CC BY 4.0: use, copy, modify and redistribute it, including in commercial products, as long as you credit the dataset (name it and link to the repository) and say if you changed it. The build and validation code is MIT. There is no warranty, and the dataset is a convenience sample of what was made public, so no share computed from it estimates a population share.
How do I add an incident?
One event is one JSON file and one pull request: copy an existing record, give it the next free id, fill every field from a public primary source, run python3 scripts/validate.py, and open the pull request. If you would rather not write JSON, use the issue form. Only events with a public primary source are accepted; this is not a place to disclose new vulnerabilities. The checklist is in CONTRIBUTING.md.
How do I cite the dataset?
Use the citation in CITATION.cff (GitHub shows it under "Cite this repository"): Muhammad Basit Ali, AI Agent Incidents: an open dataset of publicly documented AI-agent and LLM-application security incidents, version 1.0.0, 2026, https://github.com/basitalisandhu/ai-agent-incidents. The coding scheme comes from the paper AI as Weapon, Target, and Surface: A Threat Taxonomy and a Deterministic Control Plane for Securing LLM Agents (Ali, 2026), whose codebook is reproduced in docs/codebook.md; cite both when you use the coding.
What can the dataset not tell you?
It is a convenience sample of events that were made public, so it over-represents what vendors and researchers chose to disclose and says nothing about how common any class of event is in the population. Mappings are the maintainer's reading of each source against the published frameworks; an empty mapping list means no confident mapping, not that none applies. Dates are the month of the public report, not of the event. Each record links its primary source so every claim can be checked.