Vendor comparison
Codility approaches technical hiring as a measurement problem: tasks and scores meant to stay consistent from candidate to candidate, with an assessment-science posture behind that consistency. SkillFoundry approaches the same hire as an evidence problem: repository work, a trail reviewers can check, and a score of AI judgment.
Codility has spent years telling employers that a technical assessment is a measurement problem. The public posture is assessment science: tasks and scores designed so a result can be compared across candidates, with attention to whether the instrument measures the skill it names. Teams that buy on that promise are buying consistency and a defined meaning for the number.
SkillFoundry is aimed at the neighboring requirement that shows up once AI is in the editor and once counsel asks for the file. The candidate works a repository task. The score rests on tests of the submitted code and on an evidence trail reviewers can check. AI use is part of the assessment. The result describes the judgment the candidate applied to AI output. A benchmark against the hiring team's own engineers gives the number a local meaning.
Both aims are legitimate. A standardized score and an inspectable record answer different questions in the same hiring meeting. This page sets Codility's assessment-science positioning beside SkillFoundry's evidence trail and AI-judgment scoring, and it stays respectful of the measurement work. It does not restate Codility's research, and it does not invent performance figures.
If assessment science is the reason Codility is on your shortlist, evaluate it from Codility's own methodology materials and from a pilot on your roles. A competitor's summary is the wrong place to learn someone else's validity argument. What this page can say, fairly, is that Codility's identity in the market is that measurement posture, and SkillFoundry's identity is the evidence file plus an explicit treatment of AI judgment.
Ask Codility how the current product treats AI assistance, because that policy determines what a controlled score is allowed to mean once candidates have copilots. Ask SkillFoundry to show a result where AI was used, including the evidence behind the judgment score and the sentence the product prints when evidence is missing. Those two questions keep the bake-off honest.
A measurement-first product asks you to trust the instrument: the task design, the administration, and the claim that two candidates with the same score can be treated as comparable on the skill being measured. That is a serious request, and it deserves a serious reading of the vendor's own materials.
An evidence-first product asks you to trust the file: the repository work, the tests on the submitted code, the session trail, and the rule for missing evidence. Comparability comes from a shared task, a shared rubric, and a benchmark against your own engineers. The committee can open the same materials the score came from.
The rows contrast an assessment-science posture with an evidence-and-AI-judgment posture. They are written to be fair to both. Confirm Codility's current candidate rules and methodology packet before you treat any cell as a contract term.
| Area | Codility | SkillFoundry |
|---|---|---|
| Primary format | Online coding assessments built to be administered the same way across a candidate pool. | Async, repository-based assessments with a recorded working session and a result reviewers can check. |
| AI stance | Assessment integrity, and what a score is allowed to mean, sit at the center of the public story. How AI assistance is permitted is a current-product question; confirm it in Codility's candidate rules. | The AI assistant is available where the organization allows it. The score describes judgment applied to AI output, under the same rule for the cohort. |
| Scoring evidence | A score grounded in a standardized task, which is the artifact an assessment-science program is built to defend. | Trusted tests on the submitted code, plus a session evidence trail. Reviewers verify the result against that trail, including the note printed when evidence is missing. |
| Task style | Structured coding tasks designed so candidates can be compared on the same kind of work. | Ticket-style tasks inside a repository, with acceptance criteria and a human review before publication. |
| Team benchmark | Interpretation is tied to the skill the task is built to measure, on Codility's scoring model. | A benchmark against engineers already on your team, so the result also speaks to your local bar. See pricing for which plans include it. |
| Compliance tooling | Enterprise customers expect a consistent instrument they can describe to candidates. Ask Codility for the methodology and compliance packet that matches your jurisdiction. | Bias-audit tooling, logged human review, and retention controls, with the live-versus-in-progress list published on the Trust page. |
| ATS/API | ATS connectivity and employer APIs are part of how Codility sits in a recruiting stack. Confirm the current connector list with Codility. | Greenhouse, Ashby, and Workable connectors, API keys, and outbound webhooks. Additional connectors remain on the roadmap. |
| Who it fits | Teams that want a standardized coding score and a vendor whose public identity is assessment science. | Teams that want an evidence file and an explicit read on AI judgment, and that will accept a smaller task catalog in exchange for that record. |
| Measurement posture | Assessment science: consistency across candidates, and a defined meaning for what the score claims to measure. | Evidence and auditability: what was observed, what was graded on the submitted code, how AI judgment was described, and what was missing. |
Feature sets vary by plan and change over time. This page does not state Codility fees. Verify Codility's current documentation for procurement.
Codility fits teams that want a standardized coding assessment and a vendor that leads with assessment science. The promise to test is consistency: candidates face a controlled instrument, and the score is meant to be interpretable as a measure of the skill the task targets. Hold Codility to their own published methodology, and confirm the current AI rules in their candidate-facing materials.
SkillFoundry fits teams that want the evidence trail and the AI-judgment score to be what the committee discusses. Repository tasks, tests on the submitted code, a session record, and a team benchmark are the materials in that discussion. Compliance tooling is there because the screen can be an automated step the organization may need to explain.
You can respect the science and still choose the record, or respect the record and still choose the science. Pick the product whose promise matches the sentence your hiring committee has to stand behind. Walk through a SkillFoundry session before you decide. Plan details for SkillFoundry are on the pricing page. Codility's fees stay with Codility.
You want a standardized coding score and a vendor whose public identity is assessment science, consistency, and a defined meaning for the result.
You want an evidence trail reviewers can check and a score that describes AI judgment on repository work, with compliance tooling on the same record.
A few limits, stated the same way on every comparison. Accounts can run an assessment now. Evidence-based results become the primary score after a readiness review. SkillFoundry allows the AI assistant where the organization permits it, and the score describes the judgment a candidate applies to AI output. The product is explicit that assistance can still slip past any single check, and that tooling alone does not remove bias. Bias-audit tooling, logged human review, and retention controls are part of the compliance story. SOC 2 Type II is an examination in progress. Trust & Compliance lists what is live and what is still in progress.
Walk through repository work, the tests on the submitted code, and the record a reviewer can open. SkillFoundry plan details are on the pricing page.