Time window 24 hours. Suggested effort about 5 hours.
Context
An engineering team has a folder of Markdown runbooks nobody reads. They want to ask questions in plain English and get an answer that cites the exact runbook sections it came from, and that says "I don't know" when the docs do not cover it.
Core requirements
- Write or generate a corpus of at least 15 short Markdown runbooks in
docs/(your own content: deploys, on-call, incident levels, database restores, and so on). - Ingest: split documents into chunks by heading, keep the source path and heading for each chunk.
- Retrieve: rank chunks for a question. A lexical method (BM25 or TF-IDF) is enough; embeddings are optional.
- Answer: an
LLMProviderinterface with two implementations: a deterministic offlineFakeProvider(used by tests and by default) and a real provider enabled by an environment variable (any vendor). - The answer cites its sources as
[path#heading], and only sources that were actually retrieved. - A CLI or small web UI: ask a question, see the answer and the cited passages.
Acceptance criteria
- Commit
eval/questions.jsonlwith at least 10 questions and the source each answer should come from. Anevalcommand reports retrieval hit-rate at k=3 and runs with the fake provider, offline. - Questions the corpus cannot answer produce a refusal, not an invented answer (covered by the eval set).
- No API key is needed to run, test or evaluate the project. Keys are read from the environment, never committed.
- Tests cover chunking, ranking on a small fixture, citation filtering and the refusal path.
Stretch goals (optional)
- Hybrid retrieval (lexical plus embeddings) with a measured comparison on your eval set.
- Streaming answers in the UI.
- Incremental re-indexing when a document changes.
Constraints
- Any stack. Suggested: Python (FastAPI or a CLI) or TypeScript (Node).
- The real-provider path is optional; the offline path is required.
Deliverables (every project)
- Source code committed in this repository (the grader diffs against the first commit).
README.mdthat replaces the stub, with: how to install, run and test it (copy-pasteable commands); the decisions and trade-offs you made; what you would do next with more time; and a short note on how you used the AI agent (what you delegated, what you checked or rewrote).- Automated tests that run with a single command (
npm test,pytest,go test ./...orcargo test). - No secrets in the repository. Anything configurable reads from environment variables with safe defaults.
Ground rules
- The 24-hour clock is a window, not a workload. Stop at roughly the suggested effort, then write down what you would do next. A small, finished, tested core beats a large unfinished one.
- Use the AI agent as much or as little as you like: every prompt is recorded and the report shows how it was used. You are judged on the result and on whether you understood and verified what the agent produced.
- The work is yours. PraxisAI uses it only to produce your assessment report.