Scoring method

How a round is scored

Deterministic rules, applied to a recording of the round, with the evidence for every number. Scores help a person review the work; they are not meant to be the sole basis for a hiring decision.

The short version

  • Deterministic

    The same recording always gives the same score. Fixed rules compute every number; no language model forms an opinion of the candidate on a bug-fix round.

  • Explainable

    Every dimension ships with the evidence that produced it, down to the prompt or test run, so a reviewer can disagree with a specific point.

  • Assists, never decides

    Scores support human judgment. PraxisAI never rejects anyone automatically, and a report should not be the sole basis for a decision.

Six dimensions

Each dimension is scored 0 to 100 from the round's recording. The overall score is their weighted sum, with a letter grade (A+ from 90, A from 80, B from 70, C from 60, D from 45).

Outcome

weight 35%

Did the fix work? Measured by the task's hidden tests, which the candidate never sees.

  • Share of the hidden failing tests (FAIL_TO_PASS) that now pass: up to 75 points.
  • Share of the existing tests (PASS_TO_PASS) that still pass: up to 25 points.
  • A submission with no change scores near zero; code that does not run scores zero.
  • Take-home builds: automated checks, plus a rubric review when one was run (the report says which).

Prompt quality

weight 15%

How the candidate directed the agent.

  • Each prompt is rated from its text: specific files or identifiers, error context, a clear action, constraints or acceptance criteria.
  • Pasted material (raw errors, code dumps, the issue text) is rated on the words the candidate wrote around it.
  • Points off for vague prompts, re-sending the same request, and work split into too many or too few prompts.
  • Points on for correcting the agent's course.

Verification

weight 15%

Whether the candidate checked the work instead of trusting the agent.

  • Running the tests, and how often.
  • Reading the code before and while changing it.
  • Testing the final version: edits after the last test run cost points.
  • Whether the last local run before submitting was green.

Token economy

weight 15%

What the result cost in agent usage.

  • Tokens used against the round's budget.
  • Context bloat (the conversation growing fast per request).
  • Tokens burned inside failing agent loops.

Held to what the outcome earned: thrift on work that did not land is only partly credited.

Speed

weight 10%

Time used against the time allowed.

  • Finishing early scores higher; running over scores lower.
  • Credited only for work that landed: fast and wrong never scores.

Fully gated on the outcome. A round with nothing attempted scores zero here.

Recovery

weight 10%

What happened after something went wrong.

  • Error loops: three or more failing agent steps in a row without a change of approach.
  • Repeating an identical prompt, or 'fix it'-style nudges while stuck.
  • Getting back to a green test run after a loop.

Held to what the outcome earned, down to a floor that still shows the behavior.

When the hidden tests cannot run (the task's environment failed to build), the round is marked not graded with no score, rather than a misleading number.

How AI use is treated

Using the agent is expected; it is not cheating and it is not penalized. What is measured is how it was used: whether the candidate gave it specific direction, checked what it produced, noticed when it went wrong, and what the result cost. A round solved with heavy, well-directed agent use can score at the top.

The agent runs through our proxy, so every prompt and token is recorded and capped. Help from outside the workspace (another AI tool, another person) cannot be seen directly; attention signals such as long focus losses and large pastes are flagged for a reviewer instead. After the round, candidates can write a short debrief explaining their change; it is shown to the hiring team as written after the round. It is not part of the score; it feeds the ownership and communication behavior in the hireability read.

The written verdict

The headline on a report (Strong hire signal, Promising, Mixed signal, Concerns) is chosen by rules from the same numbers: for example "Strong" needs the hidden tests to pass with nothing broken, an overall score of 75 or more, solid process scores, no high-severity flag, and prompts mostly in the candidate's own words. The strengths, risks and interview questions under it are templates filled from the evidence, each pointing at the moment it refers to. Candidates see their own copy of an assessment without the hiring team's verdict wording or interview questions.

Hireability

At the top of every report, the hiring side's question: from an AI-coding point of view, how strong is this candidate? Seven behaviors are scored 0 to 100 by itemized rules over the same recording, each point tied to a sentence and, where there is one, a moment in the round. The hireability score is their weighted average (75%) blended with the hidden-test outcome (25%). It is a signal for the person deciding, never a decision.

Problem understandingweight 15
Read the code before changing it, reproduced the failure first, framed the task instead of pasting the issue, changed the right file.
Steering the AIweight 20
Prompt quality, plus points for correcting the agent and asking it to prove its work; off for “fix it” nudges, raw error pastes, code dumps and long unsupervised runs.
Verificationweight 20
How often the tests ran, whether the final version was tested and green, and catching the agent's wrong turns.
Debuggingweight 10
How failures were handled: failing loops left running, retries without a diagnosis, getting back to green. Not measured when nothing failed.
Code qualityweight 15
Size and scope of the diff, files unrelated to the fix, regressions, regression tests added, test files edited while failing.
Efficiencyweight 10
Token economy and pace from the score breakdown (60/40), both held to what the outcome earned.
Ownership and communicationweight 10
The debrief and the in-round answers: how much was explained, whether it names the actual change, and whether answers were pasted.

The call: Strong hire from 85, Hire from 70, Lean no from 50, No hire below. Strong hire needs the issue solved and no behavior under 50; any red flag (pasted debrief answers, tests edited while failing, an unsupervised agent, tests never run) holds the call at Hire, and a high-severity one (only test files changed, an integrity flag) at Lean no. A round that could not be graded, or recorded too little, gets no call. A behavior that was not observed is left out of the average rather than scored zero.

Attention flags

Raised from editor and proxy events, shown with timestamps, and never applied to the score automatically except where a dimension above says so.

Long focus loss
The candidate left the workspace tab for more than a minute.
Large paste
More than 500 characters pasted at once.
Paste after focus loss
A paste soon after returning to the tab. Marked low severity when the text came from the round's own terminal.
Paste-heavy prompting
Half or more of the prompts were mostly pasted material.
Repeated prompts
The same prompt sent more than once.
Budget exhausted
The agent stopped answering because the token budget ran out.

Known limitations

The weights are a design choice
35/15/15/15/10/10 reflects what we think matters. It has not yet been validated against later job performance, and we have not yet run an independent bias audit. Treat the overall number as a summary of the evidence, not a measurement of a person.
Prompt rating reads text
Terse prompts score lower. People writing in a second language, or who prefer short instructions, may be rated lower for reasons unrelated to skill. Read the prompts themselves before relying on this dimension.
Public tasks can be memorized
Bug-fix tasks come from SWE-bench, whose issues and merged fixes are public on GitHub and likely in AI models' training data. An agent may reproduce a fix from memory, which inflates the outcome. Hiring teams can use tasks held back from the public practice site, and every task carries a contamination label; neither makes an upstream issue unknown to models.
One harness, one agent
Candidates use the agent built into the workspace, which may not be the tool they use every day. Unfamiliarity can cost time early in a round.
Attention signals are signals
Focus and paste flags have innocent explanations (a second monitor, reading documentation, pasting from the round's own terminal). They are for a person to look at, never an automatic penalty.
Percentiles need volume
A percentile compares against graded rounds so far. With few rounds it is noisy; the report states the cohort size.