hermes

Review an LLM app for security

Reviews an LLM app for prompt injection, data exfiltration through tools, excessive agency and unsafe output handling, mapped to the OWASP LLM Top 10. Use before shipping an agent or RAG feature.

context

A language model cannot reliably tell instructions from data. Any text that reaches its context, whether a user message, a retrieved document, a web page, an email or a tool result, can steer it. The damage depends on what the model can do next. The dangerous combination is access to private data, exposure to untrusted content, and a way to send data out (an outbound request, a rendered image or link, an email). Controls that only ask the model to behave ("ignore malicious instructions") are not security controls. Real controls sit outside the model: least privilege, human confirmation, output encoding, egress limits, isolation.

task

Review this LLM application: Only if [PROMPTS] is given: Prompts: Only if [TOOLS] is given: Tools:

  1. Map trust boundaries: list every source of text entering the model's context and who controls it, every tool and what it can read or change, and every place model output goes (a browser, a database, a shell, another model, an email).
  2. Check each risk in the OWASP Top 10 for LLM Applications (2025): LLM01 prompt injection (direct and indirect), LLM02 sensitive information disclosure, LLM03 supply chain, LLM04 data and model poisoning, LLM05 improper output handling, LLM06 excessive agency, LLM07 system prompt leakage, LLM08 vector and embedding weaknesses, LLM09 misinformation, LLM10 unbounded consumption.
  3. Pay special attention to:
  • Exfiltration paths: markdown images or links rendered with attacker-chosen URLs, tools that fetch URLs or send messages, and logs visible to others.
  • Tool permissions: service-wide credentials where per-user ones are needed, write or delete actions without confirmation, parameters the attacker can influence.
  • Retrieval: access control enforced at query time per user and tenant, and poisoned documents.
  • Output handling: model output inserted into HTML, SQL, shell commands, file paths or code without encoding or validation.
  • Secrets in system prompts (assume the prompt will leak).
  • Cost and abuse limits: token, rate and loop limits.
  1. For each finding, write an attack scenario with a short, harmless example of the injected text and where it would come from, the impact, and a fix enforced outside the model.
constraints
  • Report only risks that the described architecture actually has. If a component is not described, ask under Tests to add or the verdict rather than assuming the worst.
  • Do not offer "tell the model to ignore injections" as a fix. Prompt hardening may be listed only as defence in depth beside a real control.
  • Keep injected-text examples benign (for example, exfiltrating a marker string), never working payloads against real services.
  • Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
  • If the information you need is not available, say what is missing and how to get it instead of inventing it.
output format

Verdict

One line: ship | ship after fixes | redesign needed, and the single biggest risk.

Trust boundaries

Three short lists: untrusted inputs, capabilities (tools and data), output sinks.

Findings

Numbered, ranked. Each: OWASP LLM id, severity, attack scenario, impact, fix.

Adequate controls

What is already sound.

Tests to add

Red-team cases to automate, each with its input source and the expected safe behaviour.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
Security
level
Expert
made for
Security engineer, ML / AI engineer, Software engineer, Software architect
risk
read-only
version
v1.0.0 · incubating
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install review-llm-app-security --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill review-llm-app-security -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

pairs well with

All of Security
PersonaSecurity

Security auditor

Reviews code for exploitable weaknesses and reports only issues with a concrete attack path. Use as a reviewer persona or subagent for security-sensitive changes.

security-auditor
PromptSecurity

Respond to a leaked secret

Produces an ordered response plan for an exposed key, token or password - revoke and rotate, audit use, clean up copies, notify and prevent. Use right after a secret is committed, logged or shared.

respond-to-leaked-secret
PromptSecurity

Review an API against the OWASP API Top 10

Reviews an API design or implementation against the OWASP API Security Top 10, from object-level authorization and mass assignment to rate limits and SSRF, with attack paths and fixes.

review-api-security
PromptSecurity

Review a cloud IAM policy

Reviews AWS, GCP or Azure IAM policies for over-broad permissions, privilege-escalation paths, wildcard resources and missing conditions, and proposes least-privilege versions.

review-cloud-iam-policy
PromptSecurity

Review a pull request for security

Reviews a diff for exploitable vulnerabilities and reports only findings with a concrete attack path. Use before merging changes to input handling, auth, data access or dependencies.

review-pr-for-security
PromptSecurity

Threat model a feature

Builds a threat model for one feature or change, mapping data flows and trust boundaries to ranked threats and mitigations. Use during design, before the code is written or merged.

threat-model-feature