Hiring teams keep returning to a question that felt strict in 2022 and now measures the wrong thing: did the candidate use AI? The useful question is whether they can direct a model, verify what it returns, and own the result that ships. SkillFoundry scores that judgment. The instrument is called Copilot Command, and it runs on repo-based tickets rather than on a puzzle a model can finish in one pass.
A ban feels like rigor. A locked browser, a proctoring flag, a written rule that says no copilot: these are easy to repeat in a debrief, and they are weak as a prediction of the job. The engineers you want will use AI on Monday morning. The engineers who cannot tell a confident wrong answer from a correct one will use it too. A screen that only measures abstinence does not separate those people. It separates people who complied with a constraint the workplace will not keep.
The ban is a proxy for a harder question
A ban answers a compliance instinct. If we cannot see the work, forbid the tool. That instinct is understandable, and it is aimed at the wrong risk. The risk in an AI-assisted hire is not that a model touched the keyboard. The risk is that nobody checked the model. A candidate who pastes a ticket into a chat window, accepts the first diff, and cannot explain the failure mode is not demonstrating seniority. A candidate who scopes the change, rejects a patch that misses the invariant, and writes the test that would have caught the bug is demonstrating the job.
Those two people can produce similar-looking code. A pass or fail on hidden tests will often score them the same. A ban tries to admit only the second person, and it fails at that too, because a determined candidate can use a model off screen. You punish the honest candidate and miss the careless one. The signal you wanted, judgment, never enters the record. The debrief then fills the gap with memory: who seemed senior, who hesitated, who explained with fluency. That reconstructed impression is the process an evidence trail is meant to retire.
Regulation is moving toward explainable use, not toward a fantasy of unaided coding. Regulation (EU) 2024/1689, the EU Artificial Intelligence Act, treats certain AI systems intended for recruitment and candidate evaluation as high-risk. The text is published on EUR-Lex. A hiring process that cannot explain how a score was produced is a weak place to stand, whether or not a given screen is in scope. This article is an argument about what to measure. It is not legal advice, and it is not a reading of your obligations under that regulation.
What directing, verifying, and owning look like
Directing means the candidate decides the shape of the work before the model does. On a real ticket that looks like a short plan: which module owns the bug, which invariant must hold, which test is the proof. A vague prompt, such as fix the flaky payment retry, is not direction. A prompt that names the idempotency key, the duplicate-charge failure, and the test that should go red and then green is direction. You can read the difference in the session. You cannot reliably read it from a final file alone, because a finished patch hides whether anyone steered it.
Verifying means the candidate does not treat a fluent suggestion as a merge. They run the tests. They read the diff against the ticket, not against their hope that the task is done. They reject a patch that papers over a race, or that changes a public contract the ticket did not authorize. Verification is a behavior you can record. It shows up as a rejected proposal, a test command, a revision that names what was wrong. It is not a vibe a reviewer reconstructs after the call, and it is not the same thing as a green check mark on a hidden suite the candidate never saw.
Owning means the candidate can say what shipped and why, tied to the code that was submitted. If the model proposed three approaches, they can say which one they refused and what evidence they used. Ownership is the opposite of claiming the AI wrote it and stopping there. The name on the change is still a person, and that person can walk a teammate through the decision without inventing a story the session does not support. None of this requires the candidate to code without help. It requires them to be the engineer in the loop.
That is the bar most teams already claim they hire for. Job descriptions ask for people who can use AI well. Interview loops then try to catch people using AI at all. The contradiction is expensive. Senior candidates notice it, and the ones with options decline a process that treats the tool they use every day as misconduct. You keep a ritual. You lose the measurement. The format has to change before the scorecard will mean what the hiring manager thinks it means.
Why detection-first screens fail the job
Detection-first products try to answer whether AI was used, with keystroke anomalies, browser locks, and similarity checks. Some of those signals are useful for a different question: was this the person who sat down, or did someone else complete the task? Authorship and AI judgment are not the same problem. Conflating them produces two bad outcomes. The first is false comfort. A candidate can clear a lockdown and still paste unreviewed code the moment they have a job laptop. You measured the constraint, not the habit. The second is a flag that outruns the evidence. Editing rhythm changes when someone is thinking, when a build is slow, and when they switch from reading to typing. A flag without a trail becomes a story the reviewer cannot check.
SkillFoundry keeps those questions apart on purpose. Copilot Command assumes the pair engineer is available and scores how it was used. Authenticity is a separate record a reviewer can open. Missing evidence does not silently lower a score. We do not claim to catch every form of outside help, and we do not claim a score eliminates bias. We claim a narrower thing, and the product is written so a reviewer can see what was measured. A candidate should be able to see the same evidence, rather than a reconstructed impression of how the session felt. Honesty about that limit is part of the assessment, not a footnote for the security review.
How Copilot Command scores judgment on a repo ticket
Copilot Command places a controlled AI pair engineer next to a repo-based ticket. The candidate is expected to use it. Using AI is not penalized. The score looks at whether they directed the pair, whether they checked what came back, and whether they took responsibility for the result. Reviewers get a session replay. They do not get a single number with nothing behind it. When a decision is later questioned, the evidence is the session, the submitted code, and the tests that ran on that code. AI-orchestration scoring is the product description of that instrument, without treating AI use itself as a defect.
The task is a ticket in a repository, scoped like work a team ships: a bug with a failing case, a small design constraint, tests that mean something. That setting matters. A blank algorithm prompt rewards people who have memorized patterns, or who can prompt a model to recite them. A ticket in a codebase rewards people who can locate the relevant module, respect an existing contract, and prove the fix. The model can help with all three. It can also invent a helper that breaks a caller the candidate never opened. The difference shows up when someone reads the diff before they accept it. That reading is the work. The score is a record of whether it happened.
The same rubric applies to every candidate on the task, and the attempt stays tied to the scoring profile that produced it. That consistency is what makes a comparison across a slate possible. It is also what makes a later dispute boring, in the way engineering leaders and counsel want disputes to be boring. Here is the ticket. Here is the code. Here is the replay. Here is the score. There is no side channel where one candidate was graded on charm and another on a hidden test nobody can rerun. Demographic attributes are not inputs to the score. The work product and the session are.
What to ask instead of whether they cheated
Replace the ban with three questions you can put on a scorecard. Did they direct the tool? Look for a plan that names the constraint, not a transcript of hopeful prompts. Did they verify the output? Look for a rejected suggestion, a test they ran before accepting, and a diff that matches the ticket. Did they own the result? Look for an explanation a teammate could review, tied to the submitted code, not to a memory of the call. Those questions are the job. They are also the only version of AI in the interview that still means something when every candidate has a model in the next tab.
Banning the tool measures obedience to a rule the workplace will not enforce. Scoring the judgment measures the skill you will actually manage. If the role uses an AI pair, the assessment should use an AI pair, under conditions you can explain to a candidate, a hiring manager, and a reviewer who was not in the room. The point of the record is that the reviewer does not have to have been in the room. The session holds what happened. The score names how it was judged. The human who advances or rejects the candidate does so against that file, and the file remains after the meeting ends.
If you want to see how SkillFoundry runs that screen, the SkillFoundry home page describes the assessment, and a sample report shows the kind of evidence a reviewer keeps. Read both before you write another policy that treats the model as contraband. The question was never whether a serious engineer will touch AI. The question is whether they can direct it, verify it, and own the result. That is the screen worth giving.
See the judgment, not the ban
Copilot Command scores how a candidate directs, verifies, and owns AI-assisted work on a repo ticket.