Review an LLM app for security
Reviews an LLM app for prompt injection, data exfiltration through tools, excessive agency and unsafe output handling, mapped to the OWASP LLM Top 10. Use before shipping an agent or RAG feature.
A language model cannot reliably tell instructions from data. Any text that reaches its context, whether a user message, a retrieved document, a web page, an email or a tool result, can steer it. The damage depends on what the model can do next. The dangerous combination is access to private data, exposure to untrusted content, and a way to send data out (an outbound request, a rendered image or link, an email). Controls that only ask the model to behave ("ignore malicious instructions") are not security controls. Real controls sit outside the model: least privilege, human confirmation, output encoding, egress limits, isolation.
Review this LLM application: Only if [PROMPTS] is given: Prompts: Only if [TOOLS] is given: Tools:
- Map trust boundaries: list every source of text entering the model's context and who controls it, every tool and what it can read or change, and every place model output goes (a browser, a database, a shell, another model, an email).
- Check each risk in the OWASP Top 10 for LLM Applications (2025): LLM01 prompt injection (direct and indirect), LLM02 sensitive information disclosure, LLM03 supply chain, LLM04 data and model poisoning, LLM05 improper output handling, LLM06 excessive agency, LLM07 system prompt leakage, LLM08 vector and embedding weaknesses, LLM09 misinformation, LLM10 unbounded consumption.
- Pay special attention to:
- Exfiltration paths: markdown images or links rendered with attacker-chosen URLs, tools that fetch URLs or send messages, and logs visible to others.
- Tool permissions: service-wide credentials where per-user ones are needed, write or delete actions without confirmation, parameters the attacker can influence.
- Retrieval: access control enforced at query time per user and tenant, and poisoned documents.
- Output handling: model output inserted into HTML, SQL, shell commands, file paths or code without encoding or validation.
- Secrets in system prompts (assume the prompt will leak).
- Cost and abuse limits: token, rate and loop limits.
- For each finding, write an attack scenario with a short, harmless example of the injected text and where it would come from, the impact, and a fix enforced outside the model.
- Report only risks that the described architecture actually has. If a component is not described, ask under Tests to add or the verdict rather than assuming the worst.
- Do not offer "tell the model to ignore injections" as a fix. Prompt hardening may be listed only as defence in depth beside a real control.
- Keep injected-text examples benign (for example, exfiltrating a marker string), never working payloads against real services.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
Verdict
One line: ship | ship after fixes | redesign needed, and the single biggest risk.
Trust boundaries
Three short lists: untrusted inputs, capabilities (tools and data), output sinks.
Findings
Numbered, ranked. Each: OWASP LLM id, severity, attack scenario, impact, fix.
Adequate controls
What is already sound.
Tests to add
Red-team cases to automate, each with its input source and the expected safe behaviour.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- Security
- level
- Expert
- made for
- Security engineer, ML / AI engineer, Software engineer, Software architect
- risk
- read-only
- version
- v1.0.0 · incubating
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install review-llm-app-security --target claude-codenpx skills add hermes-hq/hodios-dist --skill review-llm-app-security -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of SecuritySecurity auditor
Reviews code for exploitable weaknesses and reports only issues with a concrete attack path. Use as a reviewer persona or subagent for security-sensitive changes.
security-auditorRespond to a leaked secret
Produces an ordered response plan for an exposed key, token or password - revoke and rotate, audit use, clean up copies, notify and prevent. Use right after a secret is committed, logged or shared.
respond-to-leaked-secretReview an API against the OWASP API Top 10
Reviews an API design or implementation against the OWASP API Security Top 10, from object-level authorization and mass assignment to rate limits and SSRF, with attack paths and fixes.
review-api-securityReview a cloud IAM policy
Reviews AWS, GCP or Azure IAM policies for over-broad permissions, privilege-escalation paths, wildcard resources and missing conditions, and proposes least-privilege versions.
review-cloud-iam-policyReview a pull request for security
Reviews a diff for exploitable vulnerabilities and reports only findings with a concrete attack path. Use before merging changes to input handling, auth, data access or dependencies.
review-pr-for-securityThreat model a feature
Builds a threat model for one feature or change, mapping data flows and trust boundaries to ranked threats and mitigations. Use during design, before the code is written or merged.
threat-model-feature