Infrastructurehard

Durable background job queue

Durable jobs, safe concurrent claiming, retries with backoff, leases for crashed workers, dead letters and graceful shutdown.

Suggested effort
~5h focused work
Window
24 hours
Starts from
An empty repo
Stack
Your choice
Sign in to start this build →

Opens in a new tab. The clock starts when you press Start inside the environment, not before, and your workspace is kept while you step away.

Time window 24 hours. Suggested effort about 5 hours.

Context

A product team keeps doing slow work (emails, image resizing, webhooks) inside HTTP requests. Build them a small, durable job queue with workers, retries and visibility into what is happening.

Core requirements

  • Enqueue jobs with a type, a JSON payload, an optional run-at time and an optional idempotency key.
  • Jobs are stored durably (SQLite or Postgres) and survive a restart of every process.
  • Workers claim jobs safely: two workers never run the same job at the same time (row locking, SKIP LOCKED, or leases with expiry; explain the choice).
  • Failed jobs retry with exponential backoff and jitter up to a max attempts setting, then move to a dead-letter state.
  • A crashed worker's job becomes available again after its lease expires.
  • A small HTTP API or CLI: enqueue, list jobs by state, retry a dead job, and show counts per state.

Acceptance criteria

  • A test runs several workers against one queue with a few hundred jobs and asserts every job ran exactly once (or at-least-once with idempotent handlers; say which guarantee you give and prove it).
  • A test kills or abandons a worker mid-job and shows the job is picked up again after the lease.
  • Backoff timing is computed by a pure function with unit tests.
  • Graceful shutdown: on SIGTERM a worker finishes its current job and stops claiming new ones.

Stretch goals (optional)

  • Priorities and per-type concurrency limits.
  • Scheduled recurring jobs (cron syntax).
  • A small dashboard page.

Constraints

  • Any language. Suggested: Python, Go or TypeScript (Node) with SQLite or Postgres.
  • No hosted queue services (no SQS, Cloud Tasks and so on): the point is to build the mechanism.

Deliverables (every project)

  • Source code committed in this repository (the grader diffs against the first commit).
  • README.md that replaces the stub, with: how to install, run and test it (copy-pasteable commands); the decisions and trade-offs you made; what you would do next with more time; and a short note on how you used the AI agent (what you delegated, what you checked or rewrote).
  • Automated tests that run with a single command (npm test, pytest, go test ./... or cargo test).
  • No secrets in the repository. Anything configurable reads from environment variables with safe defaults.

Ground rules

  • The 24-hour clock is a window, not a workload. Stop at roughly the suggested effort, then write down what you would do next. A small, finished, tested core beats a large unfinished one.
  • Use the AI agent as much or as little as you like: every prompt is recorded and the report shows how it was used. You are judged on the result and on whether you understood and verified what the agent produced.
  • The work is yours. PraxisAI uses it only to produce your assessment report.
Durable background job queue: a 24-hour take-home project | PraxisAI