Skip to content

Bias audits and adverse impact

This is the page regulators scrutinize hardest, so it is the most detailed. If SkillFoundry's scores — including the behavioral score — help decide who advances in a hiring process, the tool may be an Automated Employment Decision Tool (AEDT) and fall under laws such as NYC Local Law 144, general EEOC / Title VII disparate-impact doctrine, and the EU AI Act (which classifies employment-selection AI as high-risk).

This obligation is primarily the employer's

SkillFoundry cannot commission your bias audit or give your candidates notice for you. Those are duties of the employer/employment agency that uses the tool. What SkillFoundry provides is the transparency and the data an independent auditor needs, plus a documented account of what the score measures. Read this page with your employment counsel.

Is SkillFoundry an AEDT?

Under NYC Local Law 144, an AEDT is a computational process that issues a simplified output (a score, classification, or ranking) and substantially assists or replaces discretionary employment decisions. Whether SkillFoundry is an AEDT in your process depends on how you use it:

How you use SkillFoundry Likely AEDT?
The score/behavior profile auto-advances or auto-rejects candidates Yes
The score is weighted heavily and reviewers rarely override it Likely
Results are one input a human meaningfully weighs alongside others Possibly — depends on how "substantial" the reliance is
Used only for practice / learning, not selection No

Because the behavioral score produces exactly the kind of opaque, simplified output regulators focus on, assume you are in scope unless counsel concludes otherwise, and design your process accordingly.

What NYC Local Law 144 requires (of the employer)

  1. An independent bias audit within the prior 12 months, performed by an auditor who is not involved in using or developing the tool.
  2. A publicly available summary of the audit results (on the employment section of your website), including the audit date, the distribution date of the tool, and the selection/scoring rates and impact ratios by category.
  3. Candidate notice at least 10 business days before use, stating that an AEDT will be used, the job qualifications and characteristics it assesses, and how to request an alternative process or accommodation. On request, disclose the data collected, its source, and the retention policy.

The audit must compute, at minimum:

  • Selection rate per category = (candidates selected in that category) ÷ (candidates who applied in that category).
  • Scoring rate for a scoring tool = the rate at which candidates in a category score above the sample's median score.
  • Impact ratio = the selection (or scoring) rate of a category ÷ the rate of the most-selected / highest-scoring category.

…across sex categories, race/ethnicity categories, and intersectional sex × race/ethnicity categories.

The EEOC / four-fifths rule

The federal touchstone for adverse (disparate) impact is the four-fifths (80%) rule: if the selection rate for any protected group is less than 80% of the rate for the highest-scoring group, that is evidence of adverse impact requiring investigation and, potentially, validation of the tool as job-related and consistent with business necessity (or adoption of a less-discriminatory alternative). The impact ratio computed for LL144 is the same statistic; the 0.80 threshold is the flag.

In-product methodology (SkillFoundry backend)

SkillFoundry applies the following rules when computing adverse-impact indicators for benchmark and workforce comparison views:

  • Assessment mode: When at least five subjects have employer-supplied hiring decisions and those decisions cover at least half of scored outcomes, the service uses selection rates (advance/hire). Otherwise it falls back to scoring rates (share above the sample median score).
  • Minimum category sample: Per-category impact ratios are computed only when a demographic category has at least five joined subjects. Smaller categories are reported as insufficient data rather than producing a misleading ratio.
  • Suppression: Team benchmark and skill-map comparisons are withheld when any sufficient category falls below the four-fifths threshold.

What the behavioral score is — and is not

Defensibility starts with knowing what you are auditing. SkillFoundry's evaluation combines automated correctness and code-quality checks with a behavioral assessment of how the candidate worked during the session. Reviewers see an overall result with a breakdown by area (correctness, code quality, testing, working process, communication). See Reviewing candidates and Monitoring and data collection.

For an auditor and for counsel, the important properties are:

  • It measures work process, not identity. The behavioral assessment looks at job-relevant working practices, such as how the candidate approached the problem and tested their work; committed secrets are a hard red flag. Results decompose into documented components rather than a black box, and the scoring documentation is available to auditors under NDA.
  • Any working-style descriptor is not a ranking — SkillFoundry's own guidance instructs reviewers not to filter on it.
  • Behavioral measures can still correlate with protected characteristics. Speed, tooling familiarity, and workflow norms can track disability, age, socioeconomic access, or neurodivergence. That correlation is precisely why an adverse-impact audit is necessary rather than optional — a "behavioral" measure is not automatically neutral.
  • Human-in-the-loop is designed in. Integrity flags trigger human review and never auto-reject; manual score adjustments are recorded with reviewer identity and reason. Keeping a meaningful human decision in the loop is both good practice and relevant to whether the tool "substantially replaces" discretion.

Data needed for a bias audit

An auditor computing impact ratios needs, per candidate in the assessment population:

Data Source
Assessment outcome — overall score, subscores, pass/fail, advance/reject SkillFoundry (exportable)
Behavioral results and their components SkillFoundry (exportable)
Selection decision (did the candidate advance?) Employer's ATS — the hiring decision lives with you
Demographic data — sex, race/ethnicity Employer, via voluntary self-identification (SkillFoundry does not collect protected-class data)
Time window / tool version SkillFoundry + employer records

SkillFoundry does not collect protected-class attributes

Race, ethnicity, and sex are not part of a candidate's SkillFoundry profile or telemetry (telemetry deliberately records activity, not content). Demographic data for the audit must come from your own voluntary, EEO-compliant self-identification process and be joined to SkillFoundry outcomes by the auditor, ideally on pseudonymized identifiers. This keeps sensitive data minimization intact while still enabling the audit.

What SkillFoundry provides today vs. roadmap

Capability Status
Export of per-candidate outcomes (scores, breakdowns, behavioral results) for a task or population In place — via organization data export and the API (see API reference)
Documented, decomposable scoring model an auditor can inspect In place — scoring documentation is provided to auditors under NDA
Human-in-the-loop controls (flags → review, audited manual adjustments) In place
Retention of historical outcomes to support the "prior 12 months" audit window (anonymized after the retention period) In place
In-product adverse-impact / four-fifths report (upload or join demographics → impact ratios by category and intersection) Roadmap
Candidate AEDT-notice template and delivery workflow (10-business-day notice, characteristics assessed, alternative-process request) Roadmap — today, deliver notice through your ATS/careers site; SkillFoundry supplies the "characteristics assessed" language below
Published-summary generator (LL144 website summary from audit output) Roadmap

Language you can reuse for candidate notice

For the "job qualifications and characteristics the AEDT assesses" disclosure, SkillFoundry assesses, from a candidate's work on a realistic engineering task:

  • Task correctness — whether the solution meets the acceptance criteria and passes automated tests (including tests revealed only at evaluation).
  • Code quality — structure and readability of the code.
  • Test coverage — how well the changed code is exercised by tests.
  • Working process ("behavioral") — how the candidate approached the problem, tested and verified their work, followed the requirements, managed risk (e.g. avoiding insecure patterns and committed secrets) and worked with simulated stakeholders.
  • Integrity indicators (decision-support only) — patterns in session activity that may warrant a closer look. These indicators are reviewed by a human reviewer and never auto-reject a candidate. Candidates may request a verified live round or other alternative process (see Candidate rights and appeals).
  • Communication — quality of interaction with the simulated customer/technical program manager and of the candidate's notes.

It does not assess protected characteristics, and does not use webcam, microphone, screen recording, keystroke logging, or clipboard contents (see Monitoring and data collection).

A practical compliance checklist for the employer

  • [ ] Determine, with counsel, whether your use makes SkillFoundry an AEDT.
  • [ ] Run a voluntary, EEO-compliant demographic self-identification process separate from the assessment.
  • [ ] Export SkillFoundry outcomes for the relevant population and window.
  • [ ] Engage an independent auditor to compute selection/scoring rates and impact ratios (sex, race/ethnicity, and intersectional) and apply the four-fifths flag.
  • [ ] If adverse impact appears, investigate: validate job-relatedness, adjust weighting/thresholds, or adopt a less-discriminatory alternative; keep humans meaningfully in the loop.
  • [ ] Publish the required summary on your website (LL144).
  • [ ] Give candidates the ≥10-business-day notice with the characteristics assessed and an alternative-process/accommodation path.
  • [ ] Re-audit at least every 12 months and whenever the tool materially changes.
  • [ ] Under the EU AI Act, if you deploy this in the EU for selection, layer on high-risk obligations (human oversight, logging, transparency, and — for providers — conformity assessment).