24-hour take-homes · empty repo · live agent
Build something real. From nothing.
Each project is a brief, a clear set of acceptance criteria and a rubric you can read before you start. You get an empty repository, a sandbox where you can run anything, an AI agent, and a 24-hour window, so you can fit a few hours of focused work around your day. When you submit, your tests are run, your README is checked and the work is reviewed against the rubric.
4 projects · hard
- AI / LLMhard
Ask-the-docs: question answering with citations
Retrieval-augmented answers over Markdown runbooks with real citations, refusals, an offline fake model and an eval set.
- Python
- TypeScript (Node)
~5h focused work · 24h window - Infrastructurehard
Durable background job queue
Durable jobs, safe concurrent claiming, retries with backoff, leases for crashed workers, dead letters and graceful shutdown.
- Python
- Go
- TypeScript (Node)
- +1
~5h focused work · 24h window - Developer toolinghard
Feature flag service and SDK
Targeting rules, sticky percentage rollouts, an admin API with an audit log, and an SDK that evaluates locally and degrades safely.
- TypeScript (Node)
- Go
- SQLite
~5h focused work · 24h window - Real-timehard
Team chat with rooms and presence
Rooms, live messages, presence, typing indicators and history that survives reconnects without duplicates.
- Node + ws or Socket.IO
- React or Svelte
- SQLite or Postgres
~5h focused work · 24h window
Scoped to a few hours
Every brief names a suggested effort, usually three to five hours. The 24-hour clock is a window, not a workload: stop at the suggested effort and write down what you would do next.
Judged on a public rubric
Functionality, code quality, testing, architecture, UX, documentation and how you used AI, each with a weight you can see before you start.
AI expected, and recorded
Use the agent as much as you like. Every prompt is recorded, and your report shows whether you reviewed and tested what it wrote.