Skip to content

Copilot Command: the AI-native assessment

Copilot Command measures the skill that defines modern engineering: directing an AI agent well. Instead of asking whether a candidate used AI, the assessment hands them an AI pair engineer and scores how well they supervise it.

For candidates

Inside the browser workspace, the AI pair engineer tab connects you to an agent that does real work on the task:

  1. Delegate a unit of the ticket. Clear, specific instructions get better proposals than "do the task".
  2. Review the proposal. The agent returns a concrete code change — it can span multiple workspace files when the work calls for it — with an explanation and per-file diffs. Read it. Run the tests.
  3. Decide. Accept & apply, reject, or request changes with feedback. You own everything that ships: the agent is confident, but it is not always right.

There is no penalty for using the agent heavily or barely at all — what's scored is the quality of your supervision, not the quantity of delegation.

Verify before you decide

You own what ships. Check the agent's work before you accept it.

For hiring teams

Every Copilot Command session produces an orchestration score alongside the standard results. It reflects how well the candidate supervised the agent — the quality of their instructions, whether they checked the agent's work, whether they caught its mistakes, and the soundness of what they chose to ship — with a verdict label that summarizes the result. The score appears on the evaluation dashboard and contributes to the overall result when a session exists; attempts that never used the agent are scored exactly as before.

Why the scores are defensible

  • Comparable sessions. The agent's behavior is controlled so that every candidate on the same task faces the same challenge, and what is scored is the candidate's response to it rather than the free-form behavior of the model.
  • Reproducible. Each session records the model and scenario configuration it ran against, so historical scores stay auditable as the platform evolves.
  • Guardrails. The agent is prevented from giving away the challenge to the candidate.
  • Traceable. Every delegation, proposal, and decision is recorded in the session's evidence trail, so any number in the report can be traced back to specific, replayable events — the same auditable-score standard as the rest of the platform.

Reviewers open a submission from the evaluation dashboard to see the full orchestration panel and a session replay timeline of delegations, proposals and decisions.

A public walkthrough is available at /demo/copilot-command with a canned replay.

Token metering

Each assessment session tracks the LLM tokens consumed by the agent. Subscription plans cap tokens per assessment; exceeding the cap returns an upgrade prompt in the workspace. Organizations with sales demo mode enabled use scripted proposals and do not consume tokens.

Supported stacks

Scenarios are available for Python, JavaScript/TypeScript and Go tasks.

VS Code extension

The SkillFoundry VS Code plugin includes an AI pair engineer tab with the same delegate → review → decide flow as the browser workspace, including workspace snapshots and test-run tracking.

Availability and controls

Copilot Command is on by default for organizations that allow AI assistance. Two opt-outs exist:

  • Organization level — org admins can turn off Enable Copilot Command in organization settings without disabling AI assistance entirely. Disabling AI assistance disables Copilot Command as well.
  • Task level — individual tasks can opt out via the task's pair_agent_enabled flag, overriding the organization default for that task only.

Policy and privacy

  • The evidence trail records interactions and decisions, consistent with the platform's monitoring and data collection principles.
  • Scenario details are visible to reviewers only and are never exposed to candidates during the assessment.