Evidence-based assessment · rolling out to pilot customers

Assessment that shows its work

SkillFoundry describes how an engineer investigated, used AI and verified their work, not just whether the tests passed. Every result is backed by evidence reviewers can check, and nothing is guessed when evidence is missing.

Evidence-backed resultsAI use is not penalizedMissing evidence never counts against anyone

What we look at

See how the work was done, backed by evidence

SkillFoundry looks past whether the tests pass to how a candidate investigated, used AI and verified their work. Reviewers can check every result against the evidence, and missing evidence is never counted against a candidate.

How they investigate

Whether the candidate took the time to understand the problem and the code before changing it.

How they use AI

AI use is not penalized. What matters is the judgment a candidate applies to what the AI produces.

How they verify

Whether the candidate checked that their solution works, alongside the result of grading the code they submitted.

Rolling out to pilot customers, with a readiness review before any organization's results change.

Principles

Built to be fair to candidates and defensible for reviewers

Results reviewers can verify

Every result is backed by the evidence behind it, so reviewers can check it for themselves instead of trusting a number.

Missing evidence never counts against a candidate

When something could not be captured, the result says so. A gap in the record never lowers a candidate's result.

AI use is not penalized

Candidates can use the AI assistant wherever the organization allows it. Judgment is what matters, not whether AI was used.

Graded on the submitted code

Correctness comes from running the code the candidate actually submitted, in a trusted environment.

Tasks reviewed before publication

Every assessment task is checked and approved by a person before candidates see it.

Descriptive feedback for candidates

Candidates get plain-language feedback on how they approached the task, not just a pass or fail.

What candidates see

  • Candidates are told what is captured before a session, and the workspace shows when activity is being recorded.
  • Where evidence-based assessment is the primary result, candidates receive descriptive, plain-language feedback.
  • Using the built-in AI assistant is allowed wherever the organization permits it.

Privacy

What we capture, and how long we keep it

Behavioral evidence is collected for one purpose: describing how the assessment was done.

Behavioral events, not keystrokes

We record behavioral interaction events during an assessment, such as workspace activity, test runs and use of the AI assistant. Never keystroke content, webcam, microphone or screen recording.

Code and transcripts stay protected

Code changes, test reports and AI assistant transcripts are stored as protected artifacts. Secrets are redacted, only authorized reviewers can open them, and access is logged.

24 months for evidence, 12 for artifacts

By default, behavioral evidence is kept for 24 months and protected artifacts for 12 months. Organizations can set shorter or longer periods within fixed bounds.

Evidence deleted within 48 hours

On request, behavioral assessment evidence is deleted from our live systems within 48 hours, unless a legal hold applies. Other account data is deleted or anonymized in line with our retention policy. Backups are generally kept for up to 90 days and roll off; if a backup is ever restored, pending deletions are applied again.

Research use only with opt-in consent

Practice attempts are used to improve our analysis only if the candidate explicitly opts in. Consent is off by default, separate from the terms, limited to people 18 or older, and can be withdrawn in Settings at any time. Organization assessments are used for calibration only where the organization has agreed to it by contract.

Pseudonymized research copies

Research copies are pseudonymized: names, emails, code and transcripts are removed. They are never used to score anyone and never shown to employers.

FAQ

Common questions

Does using AI lower a candidate’s result?
No. Using the AI assistant is not penalized. What matters is the judgment a candidate applies to what the AI produces.
What happens when evidence is missing?
The result says what could not be captured. Missing evidence is never counted against a candidate.
Can reviewers check a result?
Yes. Every result is backed by the evidence behind it, so reviewers can verify it before making a decision.
Is it available to every customer today?
Rolling out to pilot customers, with a readiness review before any organization's results change.

Hire on evidence you can explain

See evidence-based assessment on a real review, and talk to us about joining the pilot.