← Practice guides

How to read your coding process report

Updated

Start with what the report observed, then read how the scoring rules summarize it. The outcome, prompt record, test activity, and resource use answer different questions. A single headline number cannot tell you which decision to change next time.

This guide describes this platform's report fields. Free practice shows your score, verdict, grading outcome, and a short summary. Active plan access, including an eligible complimentary plan, unlocks the detailed report with dimension evidence, prompts, timeline, coaching, and replay. Usage credit alone does not unlock this detail. You can inspect the public sample report before choosing a plan. Process scores use deterministic rules and heuristics; they do not predict job performance or capture every part of your reasoning.

Check whether the round was graded

An environment failure can prevent the grader from running its tests. In that case the score is marked not graded and the overall number is withheld. Read the reason before diagnosing your code. A round that couldn't be measured is different from a patch that ran and failed.

Other statuses carry different evidence. No changes means the grader captured no patch. Broken changes means the submitted code did not compile. A graded result has test observations to inspect. A blank or unknown outcome should send you to the status and notes, not to an assumption that every test failed.

Separate the fix from regressions

Required tests, recorded as FAIL_TO_PASS, check behavior that needed to start passing. Existing tests, recorded as PASS_TO_PASS, check behavior that should remain intact. Read both passed and total counts. The regression count covers the evaluated set and may be smaller than the repository's complete suite.

Required tests can all pass while the regression suite fails to run. The grader then leaves the solve uncertified instead of inventing a pass. Check the notes and whether existing tests were verified. Fixing the reported symptom and preserving surrounding behavior are separate claims that each need evidence.

Read each dimension with its evidence

The six dimensions are outcome, prompt quality, token economy, verification, speed, and recovery. The report includes each score's weight and the evidence used to compute it. Prompt quality looks for cues such as file references and stated constraints. Verification uses observed test runs and file exploration. Those signals can prompt useful questions, but neither directly measures the quality of every decision.

Speed, token economy, and recovery are also limited by how much required-test progress the round achieved. Finishing early or spending less cannot earn their full scores when the required behavior didn't land. Read that explanation before comparing two rounds. Check whether cost came from metering or an estimate, and compare tasks, conditions, and outcomes alongside the totals.

Use the record to choose one next step

Use the full report's replay to inspect the prompts, test activity, and timeline behind a summary. Detailed feedback also includes prompt coaching to help you examine a request and its context. Focus-loss windows, large pastes, and repeated failures are observations worth reviewing. A focus or paste flag alone does not establish why the event happened or prove misconduct.

Find one moment where a different action would have helped: a vague request before a broad edit, a repeated error without a new hypothesis, or a submission without a relevant test. Choose a specific next action and try it on another unfamiliar task. Keep the command output or prompt that shows whether you actually changed that behavior. A score can help locate the question. The recorded work is what lets you answer it.