Design an LLM agent architecture
Designs an LLM agent system, deciding first whether an agent is needed, then single or multi-agent, tools, memory, guardrails, human checkpoints, evals and cost limits.
Many "agent" projects would be cheaper, faster and more reliable as a single model call or a fixed workflow of calls written in code. An agent, where the model chooses its own next step and tool in a loop, earns its cost only when the steps cannot be known in advance and the task is valuable enough to pay for exploration, extra tokens and harder testing. Multi-agent systems multiply token use further and add coordination failures; they pay off mainly for broad, parallelisable work such as research across many sources. Most failures in production agents come from vague tools, unbounded loops, context that grows until the model loses the thread, untrusted text in tool results steering the agent, and the absence of an eval that shows whether a change helped.
Design a system for this goal: Only if [AVAILABLE_TOOLS] is given:
Tools and systems available: Only if [CONSTRAINTS] is given:
Constraints:
Risk tolerance for wrong actions: .
- Decide the shape. Walk up this ladder and stop at the first rung that can do the job: a single model call with good context; a fixed workflow (prompt chaining, routing to specialised prompts, parallel calls, or a generate-then-evaluate loop); a single agent with tools in a loop; an orchestrator with sub-agents. Justify the rung against the example tasks, and say what evidence would justify moving up one.
- Draw the architecture: components, the control loop, where state lives, and the stop conditions (task done, step limit, budget limit, needs a human, unrecoverable error). For multi-agent designs, say what each agent owns, what it receives and returns, and why it cannot be a tool call instead.
- Specify the tools: the smallest set that covers the tasks. For each: purpose, inputs, whether it reads or changes state, its permission scope, and whether it is idempotent. Prefer a few well-described tools that do meaningful units of work over thin wrappers of every API endpoint. Separate read tools from write tools.
- Plan context and memory: what goes in the system prompt, what is retrieved on demand, how tool results are trimmed before they enter context, how long tasks are summarised or checkpointed, and whether anything is remembered across sessions (and who can see or delete it).
- Set guardrails sized to the risk tolerance: treat all tool output and retrieved text as data, never as instructions; allowlist actions and destinations; validate tool arguments in code; sandbox code execution and browsing; use credentials scoped to the user and task; and add rate and spend limits.
- Place human checkpoints by reversibility and blast radius: which actions run freely, which need confirmation, and which are never available to the model. With low risk tolerance, every irreversible or external action needs approval.
- Define evaluation: 20 to 50 realistic tasks with known good outcomes, including ambiguous and adversarial ones (injected instructions in a document, a tool that errors, an impossible request). Measure task success, wrong or unsafe actions, steps and cost per task, and inspect full traces, not only final answers.
- Set cost and latency limits: maximum steps, tokens and wall time per task, per-user or per-day budgets, the model for each role, and what happens when a limit is hit.
If the goal is too vague to pick a rung (no example tasks, no definition of success), ask for those first and stop. Otherwise state assumptions and continue.
- Recommend the simplest design that can pass the evaluation. Put more autonomy and more agents in the build order as later options, each tied to the eval result that would justify it.
- Never let the model hold credentials or decide its own permissions. Enforce limits in code, not only in the prompt.
- Name frameworks or vendors only as examples of a capability; the design must not depend on one.
- Give every number (step limits, budgets, eval size) as a starting value to tune, not a known optimum. Do not cite benchmark scores or prices you were not given.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
Verdict
The chosen rung in one sentence, why, and what would justify the next rung up.
Architecture
A Mermaid flowchart or an indented text diagram, then the control loop and stop conditions in a short list.
Tools
Table: tool | purpose | reads or writes | permission scope | idempotent | needs approval.
Context and memory
Bullets.
Guardrails
Bullets, each with what it prevents and where it is enforced (prompt, code, infrastructure).
Human checkpoints
Table: action | runs freely, needs approval, or never allowed | reason.
Evaluation
The task set, the metrics and the bar to ship.
Cost and latency limits
Table: limit | starting value | what happens when it is hit.
Failure modes
Table: failure | how it shows up in traces | mitigation. Include loops, early stopping, wrong tool arguments, prompt injection and context overflow.
Build order
Numbered milestones, each ending in something testable.
Open questions
Only questions whose answers would change the design.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- AI and ML engineering
- level
- Intermediate
- made for
- ML / AI engineer, Software engineer, Software architect, Backend engineer
- risk
- read-only
- version
- v1.0.0 · experimental
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install design-agent-architecture --target claude-codenpx skills add hermes-hq/hodios-dist --skill design-agent-architecture -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of AI and ML engineeringDesign tool definitions for an LLM agent
Designs tool or function definitions for an LLM agent, with names, descriptions, JSON Schema parameters and error returns that models call reliably. Use when exposing an API or capability to an agent.
design-tool-schemaWrite an eval suite for an LLM feature
Writes an eval set for an LLM feature with golden, edge and adversarial cases, graders matched to each criterion, and pass thresholds. Use before shipping or changing a model, prompt or pipeline.
write-llm-eval-suiteReview an LLM app for security
Reviews an LLM app for prompt injection, data exfiltration through tools, excessive agency and unsafe output handling, mapped to the OWASP LLM Top 10. Use before shipping an agent or RAG feature.
review-llm-app-securityReduce LLM costs and latency
Cuts an LLM feature's cost and latency through prompt trimming, caching, model routing, batching and output limits, each paired with the quality check that proves nothing regressed.
reduce-llm-costsMachine-learning engineer
Acts as a machine-learning engineer who starts from the data and a baseline, insists on evals and reproducibility, and distrusts any gain a simpler model explains.
ml-engineerBuild an MCP server
Implements a Model Context Protocol server exposing the given tools and resources, with input validation, least privilege and error messages a model can act on. Use to connect a system to AI clients.
build-mcp-server