hermes

Review a coding agent transcript

Reviews a coding agent session transcript for where it went wrong (bad assumptions, skipped verification, scope creep, looping) and turns each failure into an instruction-file or prompt change.

context

When a coding agent session goes badly, the cause is usually visible in the transcript a few turns before the visible failure: an assumption it never checked, a file it never read, a test it never ran, an instruction that was missing, ambiguous or contradicted by another one. Adding a vague line such as "be careful" or an all-caps "NEVER" to the instruction file rarely helps. What helps is a specific instruction placed where the agent will read it, with the reason, or a change to the setup (a command, a check, a tool) that makes the right behaviour the easy one. Some failures are model variance and no instruction will fix them; saying so prevents instruction files from bloating.

task

Review this agent session:

transcript

Only if [INSTRUCTIONS_FILE] is given:

instructions file

Only if [GOAL] is given:

What the user wanted and what went wrong:

  1. Reconstruct the task: what the user asked, what a good outcome would have been, and how the session actually ended.
  2. Find the key moments: the turns where the session's direction changed, for better or worse. For each failure, find the earliest turn where it became likely, not only where it became visible.
  3. Classify each failure:
  • Misread the request, or invented scope the user did not ask for.
  • Acted on an assumption it could have checked (an API, a file's contents, a command's behaviour, the project's conventions).
  • Missing context: did not read relevant files, docs or existing patterns.
  • Skipped verification: claimed success without running the tests, build, type check or the app; or misreported a result.
  • Scope creep: changed files or behaviour outside the task.
  • Looping or thrashing: repeated a failing approach, or edited back and forth, without new information.
  • Stopped early or handed work back that it could have finished; or the opposite, pushed on when it should have asked.
  • Unsafe or destructive action, or one taken without the confirmation the instructions require.
  • Ignored an existing instruction, or followed one that was wrong, outdated or contradicted by another.
  • Context loss: forgot earlier decisions in a long session. Quote the evidence (turn and a short excerpt) for each.
  1. For each failure, decide the cause in the setup: instruction missing, ambiguous, buried, contradicted or outdated; the prompt was underspecified; a tool, command or permission was missing; the environment misled the agent (flaky test, stale docs); or model variance with no setup cause.
  2. Propose the smallest change that would have prevented it, in order of leverage: fix or remove a wrong or conflicting instruction; add a specific instruction with its reason and the exact command or file it refers to; change how the task is prompted; add a check the agent can run (a script, a test command, a pre-commit hook); change permissions. Write each instruction as the agent would read it: concrete, positive ("Run npm test -- path after editing a file under src/") rather than vague or shouted.
  3. Note what the agent did well that the instructions should keep encouraging.
  4. Check the result for bloat: if the instruction file would grow by more than a few lines, merge or cut instead, and point out any existing lines that are now redundant.
constraints
  • Every failure cites transcript evidence. Do not speculate about the model's internal reasoning beyond what the transcript shows.
  • Prefer one high-leverage instruction over several narrow ones. Do not propose an instruction for a one-off mistake unless the cost of a repeat is high.
  • Keep proposed instructions tool-neutral where possible, so they work in any agent that reads the file.
  • Do not include secrets, tokens or personal data from the transcript in the output.
  • Separate what you verified from what you inferred. Mark inferences as such.
  • When you do not know, say "I don't know" once and state what would settle it.
output format

Summary

Three to five lines: what was asked, what happened, and the root causes.

Key moments

Table: turn | what happened | effect.

Failures

Numbered, most costly first. Each: category - evidence (turn and quote) - setup cause - proposed change.

Instruction file changes

A unified diff against the given instruction file, or the new lines with where they go if no file was given. Include removals of conflicting or redundant lines.

Prompt changes

How the user could phrase the task next time, if that was a cause. "None" otherwise.

Other setup changes

Commands, checks, hooks, tools or permissions to add or change.

Keep doing

Bullets.

Not fixable by instructions

Failures that look like model variance, and how to work around them (smaller tasks, checkpoints, review).

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
Coding-agent operations
level
Intermediate
made for
Software engineer, Tech lead / staff engineer, Engineering manager
risk
read-only
version
v1.0.0 · experimental
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install review-agent-transcript --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill review-agent-transcript -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

PromptCoding-agent operations

Write an AGENTS.md

Writes or updates a repository's AGENTS.md with the verified commands, layout, conventions and boundaries a coding agent needs, and nothing generic. Use when setting up a repo for coding agents.

write-agents-md
PromptCoding-agent operations

Write a subagent brief

Turns a task into a self-contained brief for a subagent or parallel agent, with the goal, context, scope, constraints, return format and definition of done. Use before delegating to another agent.

write-subagent-brief
PromptCoding-agent operations

Write an agent skill

Writes a reusable agent skill (SKILL.md with frontmatter, steps, scripts and references) from a repeated task, with a trigger description models can match and a test plan. Use to package a workflow.

write-agent-skill
PromptCoding-agent operations

Write an agent handoff

Writes a self-contained handoff note so a fresh agent or a teammate can continue the current task without the conversation history. Use before ending a long session, switching tools or delegating.

write-agent-handoff
PromptCoding-agent operations

Audit a coding agent's permissions

Reviews a coding agent's tool, permission and sandbox configuration for shell, network, secrets and write-scope risk, and proposes least privilege. Use before giving an agent more autonomy.

audit-agent-permissions