Pick a real bug. Ship the fix with AI.
Every challenge is a genuine GitHub issue from a major Python project, graded by the hidden tests from the fix that actually merged. Read any of them free; start a round when you are ready.
- Challenges
- 212 challenges
- Repositories
- 11 repositories
- Graded rounds
- 13 graded rounds
sympy.Array([]) fails, while sympy.Matrix([]) works
inspectdb should generate related_name on same relation links.
+300 XPCompete →
Track your progress
Sign in to see what you have solved, keep a streak, earn badges and XP, and climb the leaderboard.
From Start to score in four steps
- 1
Press Start
A fresh container boots with the real repository checked out at the commit the bug was reported against, dependencies installed and tests runnable. VS Code opens in a new tab, usually within a minute.
- 2
Fix it with the agent
An AI coding agent sits in the editor. You have 60 minutes and up to 1M tokens for the round (less if your balance pays for less); usage is charged to your credit as you go, and the agent stops when the budget runs out.
- 3
Everything is recorded
Every prompt, every tool call the agent makes, tokens spent, files edited, each test run, pastes and time away from the tab. Nothing outside the workspace is watched.
- 4
Submit and get scored
Hidden tests from the real fix decide whether it is solved. The report scores you 0 to 100 from six sub-scores, and the round earns XP toward your level and the leaderboard.
How the score is made
Your overall score is a weighted blend of six sub-scores. Solving the issue matters most, but how you drove the agent, checked your work and spent your tokens is what separates a good engineer with AI from a lucky one. Non-empty practice submissions can earn XP. Your best result on each task counts, with harder tasks and passing tests worth more.
See a sample report →- Outcome35%
Do the hidden tests pass, and does nothing that worked before break?
- Prompt quality15%
Specific, contextual instructions beat “fix it”.
- Token economy15%
Solving it without burning the budget.
- Verification15%
Running the tests yourself before you submit.
- Speed10%
Time to a working fix (only counts when it works).
- Recovery10%
Getting out of error loops instead of repeating them.