Implement a background job
Implements a background or scheduled job with idempotency, retries with backoff, dead-letter handling, timeouts, concurrency limits and observability. Use to move slow work off the request path.
Background jobs fail quietly. Queues deliver at least once, so a job that runs twice sends two emails or charges twice. Retries without backoff turn a dependency outage into a self-inflicted load spike. A job with no timeout holds a worker forever; one with no concurrency limit exhausts the database pool. Scheduled jobs overlap when a run is slower than the interval, double-run when two instances each fire the same cron, or silently stop running and nobody notices for weeks. Payloads that carry full objects go stale between enqueue and execution. A job is production-ready when running it twice is safe, failure is visible and a stuck or poisoned job cannot take the system down.
Implement this job:
Only if [QUEUE_OR_SCHEDULER] is given: Queue or scheduler: Only if [STACK] is given: Stack: Only if [VOLUME] is given: Volume:
- Read how the repo already runs background work: the queue library, worker processes, job base classes, scheduling, config, logging and metrics. Reuse them. If there is none and none was named, recommend the simplest option that fits the stack and volume, say why, and ask before adding new infrastructure.
- Design the job before coding and state it briefly:
- Trigger and payload: enqueue after the triggering transaction commits (or through an outbox), and pass ids, not whole objects, so the job reads current state.
- Idempotency: how running the same job twice is safe: a unique job key or dedup table, state checks before acting ("already sent"), upserts, and idempotency keys on outbound calls.
- Retries: which errors are retryable (timeouts, 429, 5xx, lock contention) and which are not (validation, not found); exponential backoff with jitter; a maximum attempt count and total retry window.
- Dead letters: where jobs go after the last retry, with the error and payload, and how they are inspected and replayed.
- Timeouts and limits: a per-job timeout below the queue's visibility or lease timeout, a concurrency limit sized to the downstream capacity (database pool, API rate limit), and batching for large volumes with checkpoints so a crash resumes instead of restarting.
- Scheduling (if periodic): exactly one run per interval across instances (scheduler-level uniqueness or a distributed lock with expiry), no overlap with a slow previous run, explicit time zone, and what happens to missed runs.
- Implement the job, its enqueueing or schedule, and the configuration, following the repo's conventions.
- Add observability: structured logs with job id, attempt and duration; metrics for enqueued, succeeded, failed, retried, dead-lettered, duration and queue latency; and for scheduled jobs a heartbeat or last-success timestamp that can be alerted on when it goes stale.
- Write tests: the happy path; running the same job twice produces one side effect; a retryable error retries and then succeeds; a non-retryable error does not retry; exhausting retries dead-letters the job; the timeout fires; and for scheduled jobs, the overlap and uniqueness guard. Use the queue library's test mode or an in-memory fake; no real external calls.
- Run the tests and linter and report the real results.
- Do not add a new queue, scheduler or dependency without saying why the existing ones do not fit, and ask first if it needs new infrastructure.
- Never put secrets or personal data in job payloads or logs; pass ids.
- Graceful shutdown: a worker that receives a stop signal finishes or releases its current job instead of dropping it.
- Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
- Keep the change as small as it can be while still being correct.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
Design
Bullets for trigger, payload, idempotency, retries, dead letters, timeouts and concurrency, and scheduling, each one line.
Changes
One line per file.
Tests
One line per test and the real result of the run.
Configuration
| Setting | Default | Why |
Operational notes
How to monitor it, which alerts to add, how to replay dead-lettered jobs, and how to pause or drain it safely.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- Implementation
- level
- Intermediate
- made for
- Backend engineer, Software engineer, Full-stack engineer
- needs
- repo-read, file-write, shell
- risk
- runs-commands
- version
- v1.0.0 · incubating
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md
use in
npx @hermes-hq/hodios install implement-background-job --target claude-codenpx skills add hermes-hq/hodios-dist --skill implement-background-job -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of ImplementationBackend engineer
Acts as a backend engineer focused on correct data handling, clear API contracts, explicit failure modes and services that are easy to operate. Use as a builder or reviewer persona for server code.
backend-engineerBuild a webhook handler
Implements a webhook receiver with signature checks, replay protection, idempotent processing, fast acknowledgement, async work, retries and tests. Use when integrating Stripe, GitHub or similar.
build-webhook-handlerDesign an event-driven system
Designs an event-driven flow with event schemas, topics, partition keys, idempotent consumers, an outbox, retries, dead letters and replay. Use when moving synchronous calls onto a broker.
design-event-driven-systemWrite an operational runbook
Writes a runbook for an alert or routine procedure with symptoms, diagnosis commands, ordered mitigations, verification and escalation. Use so on-call engineers can act without tribal knowledge.
write-runbookPut a change behind a feature flag
Wraps new behaviour behind a feature flag with a safe default, a kill switch, tests for both paths and a cleanup ticket. Use when shipping a risky change incrementally.
add-feature-flagAdd rate limiting to an API
Adds rate limiting to API endpoints with a fitting algorithm, keys, per-tier limits, standard headers, 429 responses and tests. Use when protecting endpoints from abuse or overload.
add-rate-limiting