Time window 24 hours. Suggested effort about 4 hours.
Context
A support team receives hundreds of tickets a day and routes them by hand. They want a service that classifies each ticket (category, urgency), extracts key fields, and drafts a first reply, with a human always in the loop.
Core requirements
POST /triagewith{ subject, body, customer_tier }returns{ category, urgency, entities: { order_id?, product?, email? }, draft_reply, confidence, reasons }. Categories: billing, bug, how-to, account, feature-request, other. Urgency: low, normal, high, critical.- A rule-based baseline classifier that works with no model at all.
- An
LLMProviderinterface: a deterministicFakeProviderfor tests, and an optional real provider enabled by an environment variable. Model output must be validated against a schema; invalid output falls back to the baseline. - Never put secrets or full card numbers in a draft reply: redact anything that looks like a card number or API key from the input before it reaches any model.
- A simple review UI or CLI that lists triaged tickets and lets a human accept or correct the label.
Acceptance criteria
- Commit
eval/tickets.jsonl(at least 30 synthetic tickets you write, with expected labels). Anevalcommand prints accuracy and a confusion matrix for the baseline, offline. - "critical" is only assigned when the ticket mentions an outage, data loss or security; this rule is tested.
- Schema-invalid model output (missing fields, unknown category) is handled and tested using the fake provider.
- The redaction function has unit tests with realistic positives and negatives.
Stretch goals (optional)
- Store corrections and show how accuracy changes when they are used as few-shot examples.
- Batch endpoint with concurrency limits and per-request timeouts.
- Cost and latency logging per call.
Constraints
- Any stack. Suggested: Python (FastAPI + pydantic) or TypeScript (Node + zod).
- No API key is required to run or test the project.
Deliverables (every project)
- Source code committed in this repository (the grader diffs against the first commit).
README.mdthat replaces the stub, with: how to install, run and test it (copy-pasteable commands); the decisions and trade-offs you made; what you would do next with more time; and a short note on how you used the AI agent (what you delegated, what you checked or rewrote).- Automated tests that run with a single command (
npm test,pytest,go test ./...orcargo test). - No secrets in the repository. Anything configurable reads from environment variables with safe defaults.
Ground rules
- The 24-hour clock is a window, not a workload. Stop at roughly the suggested effort, then write down what you would do next. A small, finished, tested core beats a large unfinished one.
- Use the AI agent as much or as little as you like: every prompt is recorded and the report shows how it was used. You are judged on the result and on whether you understood and verified what the agent produced.
- The work is yours. PraxisAI uses it only to produce your assessment report.