Short answer

Company size should change the staffing and integration plan for an AI hiring pilot, not the standard of evidence. Choose one workflow, name the human who owns its decisions, verify the data and permissions, then measure accepted work, errors and candidate experience against the existing process. Expand only after reviewing the actual results.

This is a proposed operating guide based on linked public sources checked on September 9, 2026. The size categories are planning heuristics, not research-derived thresholds. The worksheet is a template, not a record of customer deployments.

Start with the work that needs to change

An early-stage company and a multinational may both struggle with interview scheduling. They need different implementation support, but neither benefits from buying a broad promise of autonomous recruitment without defining the immediate task.

Before a demo, write down where work currently stops. Is someone waiting for a hiring manager to review an application? Does the team need help organizing interview notes? Does a candidate need an accessible way to complete an assessment? Those are different problems and may require different tools.

Specify the input, the allowed action and the person who accepts the output. For a scheduling assistant, acceptance might mean that the right people receive a mutually available time without exposing private calendars. For application review, it includes whether the reviewer can inspect the evidence behind a suggested match.

Ashby’s application-review documentation provides a concrete vendor example: job criteria and editing permissions are part of the configured workflow. A feature description is not a demonstration that those criteria predict job performance. The buyer still has to assess its intended use.

Assign work according to capacity, not a headcount formula

The table below is a planning framework created for this guide. It does not claim that a particular employee count predicts implementation success.

Organization contextA reasonable starting scopeOwnership to nameExpansion condition
Small team with limited specialist supportOne administrative task with human approvalA hiring owner and a backupThe team can review exceptions and disable the tool
Growing company with a recruiting teamOne role family or a bounded process stageRecruiting operations, a hiring manager and the relevant data ownerThe pilot exposes integration failures and reviewer workload
Enterprise with multiple systems or jurisdictionsOne business unit with explicit data and permission boundariesBusiness sponsor, recruiting operations, security, privacy and appropriate legal reviewersResults and controls are verified for that deployment, not inferred from another unit

A small company may need more help with procurement and security than this table suggests. A large company may already have a simple, well-defined workflow. Change the plan when the actual operating conditions differ.

Avoid naming an owner who cannot see failures or stop the system. Accountability has little practical value if the responsible person receives only a monthly success summary.

Put the risk questions beside the implementation work

The NIST AI Risk Management Framework Playbook offers voluntary guidance around Govern, Map, Measure and Manage. It is a source of questions for a risk-management process, not a certification that a product or this worksheet satisfies legal obligations.

For this pilot, identify whose data enters the system, where it travels and which decisions it can influence. Record access, retention and deletion arrangements. Ask what changes when a vendor switches a model, adds a subprocessor or enables a new feature.

A published bias audit is relevant evidence, but its coverage matters. Match the audited feature, version, population and decision process to the deployment under review. The site’s AI hiring evidence register keeps those distinctions beside the vendor sources.

Applicable employment obligations need qualified review. For example, New York City’s DCWP page on automated employment decision tools describes requirements concerning a bias audit, public information and notices for covered uses. It does not mean that every hiring tool is covered, or that one audit resolves every possible discrimination concern.

The EEOC’s worker guidance on AI and employment discrimination also explains that existing protections remain relevant when an employer uses AI. This article is not legal advice and does not determine which rules apply to a particular employer.

Define a pilot that can fail

Document the baseline and the stop conditions before the pilot. Without them, a team can keep extending an unsuccessful trial without identifying what would change its decision.

Choose authorized data and a scope the team can inspect. Where practical, begin with offline evaluation or a shadow workflow that does not independently reject, rank or contact candidates. Review whether the test itself requires permissions, notices or other safeguards.

Record the current process using the same task boundaries. If the existing process includes review and corrections, the AI-assisted comparison must include them too. Do not compare end-to-end manual work with only the model’s generation time.

Set acceptance criteria before running the evaluation. The proposed worksheet below leaves the result column blank deliberately.

Include the candidate’s route to a human when information is missing or an automated assessment is inaccessible. Verify who receives that request and how it affects the hiring process. A team should not interpret a candidate’s inability to complete a task as evidence that the person lacks the skill being assessed.

MeasureRecord for both workflowsPilot result
Completed workEligible tasks and outputs accepted by the responsible reviewerTo be measured
Human effortSetup, normal review, correction and exception timeTo be measured
ErrorsMissed information, incorrect outputs and consequential actionsTo be measured
Candidate experienceCompletion, accessibility problems, complaints and available alternativesTo be measured
Data handlingInputs, permissions, retention and deletion behaviorTo be verified
Total costSubscription, usage, integration and ongoing staff effortTo be measured
Stop conditionsFailures that require pausing or disabling the systemTo be agreed

Report small samples as a limitation. A handful of successful cases cannot establish that rare harms are absent. Measures involving demographic groups need appropriate permissions, expertise and interpretation. This guide does not prescribe a universal fairness threshold.

Count correction work before calling it a saving

A suggestion can look efficient while moving effort from one person to another. A recruiter may spend less time assembling a shortlist while a hiring manager spends more time checking ambiguous recommendations.

Keep setup costs separate from repeat operating costs, but include both in the purchase decision. Record abandoned tasks and work completed manually after a failure. Do not count time saved twice when several people participate in the same task.

A simple comparison can use total authorized workflow cost divided by accepted outputs. That is a proposed accounting method, not a promised ROI. It requires consistent definitions and enough observations to explain the variation.

The budget discussion should include the fallback process. If an outage or model change makes the tool unsuitable, who performs the work, and can the team recover its records?

Expand only the part that was tested

Moving from one role family to another can change the relevant skills, language, documents and candidate population. Moving from one business unit to another can change permissions and legal requirements. Treat those as new evaluation questions, not administrative copies of an earlier approval.

Keep the configuration and review date with the results. Revisit the pilot after material changes, and let a named owner suspend the workflow when the evidence no longer describes it.

Before renewal, ask the person who handles exceptions to show the last unresolved case. That often gives the buying team a more useful next question than another demonstration of the easiest task.

Correction and disclosure

The September 9, 2026 revision withdraws the earlier claims that this guide was based on more than 50 implementations and interviews across four continents. Verifiable study records were not provided with the publication. Unsupported case-specific savings, vendor rankings and outcome figures have been removed; the guidance above is explicitly a proposed framework based on public sources.

Gene Dai is a co-founder of Metix AI, a recruitment technology company. That relationship is relevant to this analysis. The original URL and publication date remain unchanged.