Reviewing candidates¶
This guide explains what a reviewer sees after a candidate submits: automated scores, behavioral insights, integrity flags, and how to run a fair manual review.
The evaluation queue¶
Interviewer → Evaluations lists submissions by status:
- Pending — submitted, awaiting review (automated evaluation may still be running).
- Reviewing — a reviewer has it open.
- Reviewed — evaluation complete.
From the queue you can open each submission's full report, trigger re-evaluation, and apply manual adjustments.
The score report¶
Each scored attempt produces a report with:
Overall score and breakdown¶
The overall score combines several areas, each shown separately in the report:
| Area | Based on |
|---|---|
| Correctness | Test results, including hidden tests run at evaluation time |
| Code quality | Structure and readability of the submitted code |
| Test coverage | How well the changed code is exercised by tests |
| Behavior | How the candidate worked during the session |
| Communication | Interaction with the AI customer/TPM, notes quality |
Correctness carries the most weight.
Code review¶
You can see the candidate's submission as a diff, compare it with the reference solution (for PR-derived tasks, the actually-merged PR), and read their notes and chat transcript with the AI customer.
Behavioral insights¶
The behavior area summarizes job-relevant working practices, such as how the candidate approached the problem, tested their work, handled the requirements and managed risk. Committing secrets (credentials or API keys) to code is treated as a red flag. Items that strongly affect the result surface as red flags in the report — always click into these before drawing conclusions.
Interpreting AI-usage flags¶
When a task (or your organization) disallows AI assistance, the platform looks for activity patterns associated with AI-generated code. If it finds them, the submission is marked AI usage flagged and surfaced in your queue.
Treat flags as a reason to investigate, not a verdict:
- Legitimate workflows can look similar — copying from the task's own documentation, moving code between files, standard boilerplate.
- Review the flagged details in the report and where they fall in the session timeline.
- The submission timeline includes periodic snapshots of the evolving code — check whether the progression looks organic.
- The best follow-up is usually a short live conversation asking the candidate to walk through their solution.
If AI is allowed on the task, no flags are raised for AI use — score the work on its merits.
Manual review¶
For any submission you can apply a manual review: a points adjustment with a required reason and optional feedback for the candidate. Use it to:
- Credit approaches the automated tests undervalue (e.g., a better architecture that fails an over-specific test).
- Discount solutions that pass tests but violate the brief.
- Resolve investigated AI flags either way.
Manual adjustments are recorded with your identity and reason, so decisions are auditable.
Reports¶
Interviewer → Reports provides aggregates:
- Task completion reports — pass rates, score distributions, and time-to-complete per task. Useful for calibrating difficulty and time limits.
- Candidate reports — one candidate's performance across tasks.
Organization admins additionally get org-wide analytics (skill distribution, improvement trends, code-quality metrics) under Org Admin → Analytics.
Fairness checklist¶
- Apply the same evidence standard to every candidate; document manual adjustments.
- Confirm the AI policy the candidate actually saw before acting on a flag.
- Remember hidden tests: a candidate who passed all visible tests but scored lower on correctness likely failed hidden ones — check which, and whether they're fair.
- Candidates can see their own behavioral report (where enabled); assume anything in scoring may be discussed with them.