New Agentic AI Red Team Checklist Expands Attack Surface Testing
A free, comprehensive checklist offers 222 tests across 20 categories to rigorously assess autonomous AI systems, moving beyond traditional prompt injection vulnerabilities.

Security researcher Ravi Rajput has released a free, agentic AI red-team checklist designed to provide a structured approach for testing the security of autonomous AI systems. The checklist enumerates 222 distinct tests organized into 20 attack categories, aiming to identify weaknesses that often go unnoticed during standard security assessments. This resource is intended to help security teams move beyond focusing solely on prompt injection and explore a broader attack surface, including infrastructure, cloud access, tool usage, memory, and inter-agent communications.
The checklist addresses a critical gap in current AI security testing, where teams might spend considerable effort trying to manipulate an AI model directly while overlooking more fundamental vulnerabilities. These could include exposed MLflow servers, accessible cloud metadata endpoints, or inadequate customer isolation filters, any of which could lead to credential exposure or the compromise of private data without requiring a complex attack against the AI model itself. Rajput has modeled the checklist's structure after the widely respected OWASP Web Security Testing Guide, providing a familiar framework for organized security testing.
Each entry in the checklist includes detailed objectives, specific testing steps, suggested tools, expected results, severity ratings, and requirements for evidence. While it is inspired by OWASP, it is an independent resource and not an official OWASP publication. The checklist assigns proposed severities to tests, with 75 rated Critical, 108 High, 30 Medium, and nine Low. These ratings indicate potential risk rather than confirmed vulnerabilities in any specific product.
The 20 attack categories are logically grouped into four phases. The initial phase focuses on mapping the attack surface, encompassing discovery, orchestration, cloud identity, and model supply chains. The second phase examines inputs, including prompt injection, system prompt leaks, and unsafe output handling. The third phase delves into the AI's tools, excessive agency, memory management, agent networks, and Model Context Protocol (MCP) servers. The final phase covers deployment pipelines, privilege escalation, lateral movement, persistence, data theft, resource exhaustion, integrity failures, and multimodal inputs.
This phased approach allows testers to understand the full scope of what an AI agent can access before evaluating the impact of potentially malicious instructions. The checklist highlights specific attack vectors such as cloud credential theft via server-side request forgery (SSRF), remote code execution through unsafe Python pickle loading, and the misuse of tool combinations to enable unauthorized data transfers. It also scrutinizes cross-customer data access and the potential for forged messages between agents.
For memory and retrieval systems, the checklist includes tests designed to uncover boundary violations, such as whether removing a tenant identifier filter inadvertently exposes another customer's documents. This focus on data isolation is crucial for multi-tenant AI applications. Additionally, the checklist addresses the security of connected tools, noting that manipulated server requests can lead to unauthorized actions if safeguards are missing, as demonstrated by research into malicious MCP servers.
Rajput emphasizes the importance of defining a clear written scope before commencing testing, marking any excluded checks, and meticulously recording supporting logs or screenshots for each result. Destructive tests should only be performed with explicit authorization, and staging environments are recommended when necessary. The downloadable spreadsheet also maps tests to the OWASP and MITRE ATLAS frameworks, though it avoids fixed CVE references, prompting testers to verify current identifiers when reporting findings. The primary value lies in its broad coverage and the clear record it provides of tested areas and proven vulnerabilities.