hermes

Implement a background job

Implements a background or scheduled job with idempotency, retries with backoff, dead-letter handling, timeouts, concurrency limits and observability. Use to move slow work off the request path.

context

Background jobs fail quietly. Queues deliver at least once, so a job that runs twice sends two emails or charges twice. Retries without backoff turn a dependency outage into a self-inflicted load spike. A job with no timeout holds a worker forever; one with no concurrency limit exhausts the database pool. Scheduled jobs overlap when a run is slower than the interval, double-run when two instances each fire the same cron, or silently stop running and nobody notices for weeks. Payloads that carry full objects go stale between enqueue and execution. A job is production-ready when running it twice is safe, failure is visible and a stuck or poisoned job cannot take the system down.

task

Implement this job:

job

Only if [QUEUE_OR_SCHEDULER] is given: Queue or scheduler: Only if [STACK] is given: Stack: Only if [VOLUME] is given: Volume:

  1. Read how the repo already runs background work: the queue library, worker processes, job base classes, scheduling, config, logging and metrics. Reuse them. If there is none and none was named, recommend the simplest option that fits the stack and volume, say why, and ask before adding new infrastructure.
  2. Design the job before coding and state it briefly:
  • Trigger and payload: enqueue after the triggering transaction commits (or through an outbox), and pass ids, not whole objects, so the job reads current state.
  • Idempotency: how running the same job twice is safe: a unique job key or dedup table, state checks before acting ("already sent"), upserts, and idempotency keys on outbound calls.
  • Retries: which errors are retryable (timeouts, 429, 5xx, lock contention) and which are not (validation, not found); exponential backoff with jitter; a maximum attempt count and total retry window.
  • Dead letters: where jobs go after the last retry, with the error and payload, and how they are inspected and replayed.
  • Timeouts and limits: a per-job timeout below the queue's visibility or lease timeout, a concurrency limit sized to the downstream capacity (database pool, API rate limit), and batching for large volumes with checkpoints so a crash resumes instead of restarting.
  • Scheduling (if periodic): exactly one run per interval across instances (scheduler-level uniqueness or a distributed lock with expiry), no overlap with a slow previous run, explicit time zone, and what happens to missed runs.
  1. Implement the job, its enqueueing or schedule, and the configuration, following the repo's conventions.
  2. Add observability: structured logs with job id, attempt and duration; metrics for enqueued, succeeded, failed, retried, dead-lettered, duration and queue latency; and for scheduled jobs a heartbeat or last-success timestamp that can be alerted on when it goes stale.
  3. Write tests: the happy path; running the same job twice produces one side effect; a retryable error retries and then succeeds; a non-retryable error does not retry; exhausting retries dead-letters the job; the timeout fires; and for scheduled jobs, the overlap and uniqueness guard. Use the queue library's test mode or an in-memory fake; no real external calls.
  4. Run the tests and linter and report the real results.
constraints
  • Do not add a new queue, scheduler or dependency without saying why the existing ones do not fit, and ask first if it needs new infrastructure.
  • Never put secrets or personal data in job payloads or logs; pass ids.
  • Graceful shutdown: a worker that receives a stop signal finishes or releases its current job instead of dropping it.
  • Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
  • Keep the change as small as it can be while still being correct.
  • Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
  • If you could not run a check, say so plainly and say which one.
output format

Design

Bullets for trigger, payload, idempotency, retries, dead letters, timeouts and concurrency, and scheduling, each one line.

Changes

One line per file.

Tests

One line per test and the real result of the run.

Configuration

| Setting | Default | Why |

Operational notes

How to monitor it, which alerts to add, how to replay dead-lettered jobs, and how to pause or drain it safely.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
Implementation
level
Intermediate
made for
Backend engineer, Software engineer, Full-stack engineer
needs
repo-read, file-write, shell
risk
runs-commands
version
v1.0.0 · incubating
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install implement-background-job --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill implement-background-job -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

pairs well with

All of Implementation
PersonaImplementation

Backend engineer

Acts as a backend engineer focused on correct data handling, clear API contracts, explicit failure modes and services that are easy to operate. Use as a builder or reviewer persona for server code.

backend-engineer
PromptImplementation

Build a webhook handler

Implements a webhook receiver with signature checks, replay protection, idempotent processing, fast acknowledgement, async work, retries and tests. Use when integrating Stripe, GitHub or similar.

build-webhook-handler
PromptArchitecture

Design an event-driven system

Designs an event-driven flow with event schemas, topics, partition keys, idempotent consumers, an outbox, retries, dead letters and replay. Use when moving synchronous calls onto a broker.

design-event-driven-system
PromptIncident and operations

Write an operational runbook

Writes a runbook for an alert or routine procedure with symptoms, diagnosis commands, ordered mitigations, verification and escalation. Use so on-call engineers can act without tribal knowledge.

write-runbook
PromptImplementation

Put a change behind a feature flag

Wraps new behaviour behind a feature flag with a safe default, a kill switch, tests for both paths and a cleanup ticket. Use when shipping a risky change incrementally.

add-feature-flag
PromptImplementation

Add rate limiting to an API

Adds rate limiting to API endpoints with a fitting algorithm, keys, per-tier limits, standard headers, 429 responses and tests. Use when protecting endpoints from abuse or overload.

add-rate-limiting