Instrument a service for observability
Plans and adds logs, metrics and traces using OpenTelemetry conventions, golden signals, useful log fields, cardinality limits and first dashboards. Use when a service is hard to debug in production.
Services are hard to debug in production when logs are unstructured text with no request or trace id, metrics are averages that hide the slow tail, traces stop at the first queue or thread hop, and nobody can tell whether the last deploy is to blame. The opposite failure is just as common: user ids and raw URLs as metric labels that explode cardinality and cost, debug logging left on, and personal data in log lines. Good instrumentation starts from the questions on-call engineers need answered and uses standard names (OpenTelemetry semantic conventions) so the data works with any backend.
Instrument this service: Only if [STACK] is given:
Stack: Only if [EXISTING_TOOLING] is given:
Existing tooling:
- If the code is available, read the entry points, the outbound calls, the background work and any existing logging or metrics setup before proposing changes.
- List the production questions the telemetry must answer: is it healthy right now, which endpoint or dependency is slow or failing, is it the last deploy, which tenant or customer segment is affected, is it running out of a resource.
- Traces: start with the OpenTelemetry SDK and the auto-instrumentation available for this stack (HTTP server and client, database driver, message queue). Add manual spans only around meaningful business operations and expensive internal steps. Propagate W3C trace context across every hop, including queues and background jobs. Set resource attributes (
service.name,service.version, deployment environment) and a sampling policy: a head-based ratio, plus keeping all errors and slow traces if a collector can do tail-based sampling. - Metrics: request rate, errors and duration per route template for each request-driven interface; the same for each outbound dependency; saturation for the resources that limit this service (connection pools, worker queues, thread or event-loop lag, memory). Use histograms for durations with buckets around the latency targets. Follow the OpenTelemetry semantic-convention names for the stack's instrumentations, and check the current names in the conventions, since some have changed between versions.
- Logs: structured (JSON) with a fixed set of fields on every line (timestamp, level, message, service, version, environment,
trace_id,span_id) plus event-specific fields; log levels with clear meaning; one log line per error with the error type and stack trace; and no secrets, tokens or personal data (list what to redact or hash). - Cardinality limits: metric labels only from bounded sets (route templates, status class, dependency name, region). User ids, request ids, raw URLs, emails and error messages go on spans and logs, never on metric labels. Estimate the series count per metric.
- Export through an OpenTelemetry Collector where possible, so the backend can change without code changes.
- Define the first dashboards (service overview with rate, errors and latency per route, dependencies, saturation, and deploy markers) and two to four alerts on user-facing symptoms, not on causes.
- Write the code changes for the stack: SDK setup, configuration by environment variables, the log formatter, the custom spans and metrics, and context propagation for any queue.
If the stack is unknown and the code is not available, ask for it before writing code; the plan can still be written.
- Prefer standard OpenTelemetry APIs and semantic conventions over vendor SDKs, and say where a vendor-specific step is unavoidable.
- No unbounded label values on metrics. No personal data or secrets in any signal.
- Instrument what answers the questions in step 2; do not add spans or metrics with no consumer.
- Keep the overhead visible: say what the sampling ratio and log volume will cost relative to traffic, as a formula if the numbers are unknown.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
Questions to answer
Numbered list, each mapped to the signal that answers it.
Plan
Ordered rollout steps, smallest useful step first.
Traces
Auto-instrumentation, manual spans (name and attributes) and the sampling policy.
Metrics
Table: name | type | unit | labels | question it answers.
Logs
The required fields, levels, and the redaction list.
Code changes
Code blocks per file in the target stack.
Dashboards and alerts
Panels for the first dashboard, and each alert with its condition and why it matters to users.
Verification
How to send one request and find it in logs, metrics and traces, linked by trace_id.
Cost and cardinality
Estimated series per metric, log volume and trace sampling, and the levers to cut each.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- Incident and operations
- level
- Intermediate
- made for
- Backend engineer, Site reliability engineer, DevOps / platform engineer, Software engineer
- needs
- repo-read, file-write
- risk
- edits-files
- version
- v1.0.0 · experimental
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md
use in
npx @hermes-hq/hodios install instrument-service-observability --target claude-codenpx skills add hermes-hq/hodios-dist --skill instrument-service-observability -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of Incident and operationsDefine SLOs and burn-rate alerts
Defines SLIs, SLOs and an error-budget policy from a service's user journeys, with multi-window burn-rate alert rules. Use when alerting is noisy or reliability targets are vague.
define-slosWrite an operational runbook
Writes a runbook for an alert or routine procedure with symptoms, diagnosis commands, ordered mitigations, verification and escalation. Use so on-call engineers can act without tribal knowledge.
write-runbookTriage a production alert
Turns a firing production alert into a severity call, the safest mitigation to try first, ranked hypotheses and the next checks. Use in the first minutes of an incident or page.
triage-production-alertDebug a production-only bug
Debugs a bug that happens only in production by diffing environment, config, data, traffic, versions and timing, then plans safe instrumentation to confirm the cause. Use for works-on-my-machine bugs.
debug-production-only-bugDevOps engineer
Acts as a DevOps engineer who automates the second time, keeps pipelines fast and reproducible, and makes every change reversible. Use for CI/CD, infrastructure and release work.
devops-engineerPlan a game day or chaos exercise
Plans a game day or chaos exercise with failure scenarios, hypotheses, blast-radius limits, abort criteria, roles, an observation checklist and a follow-up review. Use to test resilience.
plan-game-day