Design tool definitions for an LLM agent
Designs tool or function definitions for an LLM agent, with names, descriptions, JSON Schema parameters and error returns that models call reliably. Use when exposing an API or capability to an agent.
A model decides which tool to call, and with what arguments, from the tool's name, description and parameter schema alone. Agents misbehave when tools overlap so the model guesses between them, when one tool per REST endpoint forces long brittle call chains, when parameters are free-form strings the model has to invent a format for, when results are huge raw payloads, and when errors are bare status codes that give the model nothing to correct. Tools are an interface for a reader that is literal and cannot ask questions, so they need more explanation than an API for humans, not less.
Design the tools for these capabilities, for target : Only if [EXISTING_API] is given:
Existing API to wrap:
- List the user goals the agent must reach. Map them to the smallest set of tools with distinct, non-overlapping purposes. Combine steps that are always done together into one tool, and do not mirror the existing API one to one; say which endpoints each tool combines.
- For each tool write:
- a
verb_nounname in snake_case, with a shared prefix when tools belong to one service; - a description of three to six sentences: what it does, when to use it, when not to use it and which tool to use instead, what it returns, and any side effects;
- an input JSON Schema:
type: object, a description on every property, enums for closed sets, explicit formats in the description (dates as ISO 8601, amounts in minor units), sensible defaults, a minimalrequiredlist andadditionalProperties: false; - the output shape: only fields the model needs next, stable ids it can pass to other tools, and truncation or pagination for large results with a note telling the model how to get more;
- side effects: read-only, idempotent, or destructive. Destructive or costly tools take an explicit confirmation or
dry_runparameter and say so in the description.
- Define the errors each tool can return. Every error message tells the model what went wrong and what to do next, for example "No customer matches 'Jon Smiht'. Call search_customers with a partial name."
- Write 6 to 10 selection tests: a user request and the expected tool call with arguments, including near misses where no tool or a different tool should be used.
- If a capability is too vague to define a safe tool, ask about it instead of guessing.
- Use a portable JSON Schema subset:
type,properties,required,enum,items,description,default,minimum,maximum,maxLength. Avoid$ref, top-leveloneOforanyOf, and conditional schemas, which some providers reject. - If the target enforces strict schemas (for example OpenAI's strict function calling), list every property in
requiredand express optional ones as nullable, and say that you did. Forany, say what changes per target. - Never put credentials, tenant ids or authorisation decisions in parameters. The host application supplies identity and enforces permissions.
- Keep the set under about 15 tools unless the capabilities truly need more, and say why if they do.
- Do not invent endpoints or fields of the existing API. Mark anything you assumed.
Tool set
Table: name | purpose | side effects | wraps.
Definitions
One fenced JSON array of tool objects with name, description and the schema under the target's key: input_schema (anthropic, and for any), parameters (openai, gemini) or inputSchema (mcp). Follow it with each tool's output shape.
Error catalogue
Table: tool | condition | message returned to the model.
Selection tests
Numbered: user request, then the expected call or "no tool".
Notes
Assumptions and open questions.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- AI and ML engineering
- level
- Intermediate
- made for
- ML / AI engineer, Backend engineer, Software engineer
- risk
- read-only
- version
- v1.0.0 · experimental
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install design-tool-schema --target claude-codenpx skills add hermes-hq/hodios-dist --skill design-tool-schema -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of AI and ML engineeringBuild an MCP server
Implements a Model Context Protocol server exposing the given tools and resources, with input validation, least privilege and error messages a model can act on. Use to connect a system to AI clients.
build-mcp-serverWrite an eval suite for an LLM feature
Writes an eval set for an LLM feature with golden, edge and adversarial cases, graders matched to each criterion, and pass thresholds. Use before shipping or changing a model, prompt or pipeline.
write-llm-eval-suiteBuild an LLM structured extraction step
Builds an LLM step that turns documents into schema-valid JSON, with the schema, prompt, validation and repair loop, null handling and an eval set. Use when automating invoices, forms or emails.
build-structured-extractionChoose between rules, ML and an LLM
Recommends rules, classical ML, a hosted LLM or a fine-tuned model for a problem, comparing accuracy, cost, latency and maintenance with the reasoning shown. Use before committing to an approach.
choose-ml-approachDesign an LLM agent architecture
Designs an LLM agent system, deciding first whether an agent is needed, then single or multi-agent, tools, memory, guardrails, human checkpoints, evals and cost limits.
design-agent-architectureDesign a RAG pipeline
Designs a retrieval-augmented generation pipeline from a corpus and its real questions, covering chunking, hybrid retrieval, reranking, citations and evals. Use before building or rebuilding RAG.
design-rag-pipeline