Real issues - live AI agent

Pick a real bug. Ship the fix with AI.

Every challenge is a genuine GitHub issue from a major Python project, graded by the hidden tests from the fix that actually merged. Read any of them free; start a round when you are ready.

Leaderboard →
Challenges
212 challenges
Repositories
11 repositories
Graded rounds
13 graded rounds
Daily challengenew in 13h 37m

sympy.Array([]) fails, while sympy.Matrix([]) works

hardsympy/sympy
Solve it today for the Daily challenger badge.Take it on →
Weekly tournament · 2026-W41ends in 2d 13h

inspectdb should generate related_name on same relation links.

mediumdjango/django· 0 competing
Compete →

Track your progress

Sign in to see what you have solved, keep a streak, earn badges and XP, and climb the leaderboard.

Sign in for your free rounds →
12 / 212
How a round works

From Start to score in four steps

  1. 1

    Press Start

    A fresh container boots with the real repository checked out at the commit the bug was reported against, dependencies installed and tests runnable. VS Code opens in a new tab, usually within a minute.

  2. 2

    Fix it with the agent

    An AI coding agent sits in the editor. You have 60 minutes and up to 1M tokens for the round (less if your balance pays for less); usage is charged to your credit as you go, and the agent stops when the budget runs out.

  3. 3

    Everything is recorded

    Every prompt, every tool call the agent makes, tokens spent, files edited, each test run, pastes and time away from the tab. Nothing outside the workspace is watched.

  4. 4

    Submit and get scored

    Hidden tests from the real fix decide whether it is solved. The report scores you 0 to 100 from six sub-scores, and the round earns XP toward your level and the leaderboard.

How the score is made

Your overall score is a weighted blend of six sub-scores. Solving the issue matters most, but how you drove the agent, checked your work and spent your tokens is what separates a good engineer with AI from a lucky one. Non-empty practice submissions can earn XP. Your best result on each task counts, with harder tasks and passing tests worth more.

See a sample report →
  • Outcome35%

    Do the hidden tests pass, and does nothing that worked before break?

  • Prompt quality15%

    Specific, contextual instructions beat “fix it”.

  • Token economy15%

    Solving it without burning the budget.

  • Verification15%

    Running the tests yourself before you submit.

  • Speed10%

    Time to a working fix (only counts when it works).

  • Recovery10%

    Getting out of error loops instead of repeating them.