# How Should Employers Protect Candidate Privacy When Using AI Hiring Tests?

psychprofile.io · September 30, 2026

> What AI hiring test privacy actually means AI hiring test privacy means controlling what information an employer collects, how an assessment system...

## What AI hiring test privacy actually means

AI hiring test privacy means controlling what information an employer collects, how an assessment system analyzes it, who can see the results, and how long those results are retained. An AI-assisted hiring system may combine test responses, résumé data, interview notes, employment history, audio or video recordings, device information, and inferences about personality, motivation, or likely job performance. Those outputs are not automatically objective facts: statistical patterns can reproduce bias present in training data, and an incorrect inference can affect a real applicant. Privacy therefore overlaps with accuracy, accessibility, discrimination, cybersecurity, and employment law; it is not merely a promise to keep an applicant database private.

**Also worth reading:** [What Is the 2026 AI Hiring Compliance Guide for Employers Using Screening Tools?](https://psychprofile.io/knowledge/what_is_the_2026_ai_hiring_compliance_guide_for_employers_using_screening_tools.php) · [What Should Employers Include in an AI Hiring Audit Checklist in 2026?](https://psychprofile.io/knowledge/what_should_employers_include_in_an_ai_hiring_audit_checklist_in_2026-3.php) · [How Does AI Hiring Bias Affect Candidates, and What Can Employers Do About It?](https://psychprofile.io/knowledge/how_does_ai_hiring_bias_affect_candidates_and_what_can_employers_do_about_it.php)

A useful privacy review asks whether the employer can explain the purpose of every data element and reject data that is not necessary for the job. It must also determine whether applicants receive meaningful notice, can correct inaccurate information, and can request deletion where applicable. As of September 30, 2026, there is still no single American federal rule that governs every psychological AI hiring tool, so employers must account for federal agencies, states, cities, and the countries in which applicants live. The safest default is to collect less, infer less, retain less, and give human reviewers—not the model—the final hiring decision.

## Why applicants and employers are exposed

AI hiring systems create risk partly because applicants often cannot see the hidden variables behind a result. A score may be based partly on facial features, voice characteristics, response timing, writing patterns, or correlations learned from a historical workforce. Those variables can be proxies for disability, race, sex, age, national origin, or other characteristics protected by law. The National Law Review has described employment AI as raising bias, privacy, and compliance concerns, while the Center for Democracy & Technology has examined how third-party AI assessments can make accountability harder when vendors control the model, data, and validation methods.

Employers remain responsible for selecting a tool, configuring it, interpreting its output, and acting on that output. A contract stating that a vendor is “AI compliant” does not transfer the employer’s legal obligations to the provider. Vendors can disappear, change model versions, retain logs, reuse information, or combine datasets unless the agreement expressly restricts those practices. A 2024 White & Case report on emerging AI regulation also illustrates why a patchwork of laws matters, although the exact legal effect on a particular system depends on the applicant’s location and the tool’s function.

Security incidents add a second category of exposure. Applicants may assume that a rejected candidate’s data disappears immediately, but operational copies can remain in databases, backups, analytics systems, support tickets, and subprocessors. Under principles associated with the European General Data Protection Regulation, a candidate inquiry could concern inaccurate data, profiling, retention, access, correction, restriction, or objection to automated decision-making. GDPR does not apply merely because a company calls a service “AI,” and a psychological test is not always legally “special category” data by itself; inferred health or biometric attributes can nevertheless raise additional restrictions depending on how they are produced and used.

## Psychological profiles and inferred traits

An AI psychological profile is an estimate of traits, preferences, or predicted behavior. It may describe communication style, conscientiousness, stress sensitivity, leadership potential, or risk of attrition, but no workplace test can observe personality directly. A test measures responses under particular conditions, and the interpretation depends on the questions, scale, comparison group, language, culture, setting, and model design. Research published in Nature has explored the role of AI in analyzing behavior and predicting traits or disorders, but experimental accuracy should not be confused with permission to deploy the same model in selection.

The privacy stakes are higher when an inference is medically suggestive. A system may label an applicant anxious, neurodivergent, depressive, unstable, or unsuitable for high-pressure work even though the employer lacks a valid basis to make that judgment. Some tools use voice, face, keystrokes, or eye tracking, which may qualify as biometric information under laws such as Illinois’s Biometric Information Privacy Act. BIPA’s private right of action can create substantial liability when a company collects biometric identifiers without the required written notice and consent, and its restrictions cannot be avoided simply by claiming the data was collected by a vendor.

Candidates should receive clear notice before a test, including the categories of data collected, the intended job-related use, the main factors influencing scores, the retention period, the vendor, and whether human review occurs. “AI may be used” is not adequate notice if the applicant cannot determine whether recordings are generated, how inferences are made, or what remedy is available. Notices should be understandable at an appropriate reading level and available in accessible formats, especially because the use of AI must not become a disguised barrier for candidates with disabilities.

## Legal duties that apply before deployment

In the United States, the Equal Employment Opportunity Commission has warned that Title VII and other federal laws prohibit discrimination based on race, color, religion, sex, national origin, age, disability, and genetic information. Algorithmic scoring is still a hiring practice; adverse impact can occur even when no programmer intended to exclude a protected group. Employers should test whether selection rates differ materially across groups, investigate unexplained gaps, and avoid adopting a threshold that produces or worsens disparity unless a documented, lawful job-related reason supports it.

New York City Local Law 144 is one of the clearest examples of jurisdiction-specific requirements. For covered employers using an automated employment decision tool to make a substantial selection decision, it requires a bias audit, advance notice to candidates, access to the underlying data and explanation of the tool’s role, and an opportunity for human review of adverse decisions. The law does not make AI illegal, and not every recruiting tool is covered, so an employer must first examine the tool’s purpose, decision thresholds, employment context, and New York City nexus. “Human in the loop” is not a magic phrase if a reviewer merely rubber-stamps the model without information, time, and authority to disagree.

Accessibility law can apply even when privacy law does not. A candidate may need a screen reader, extra time, a different response mode, or an alternative assessment. The Job Accommodation and ADA framework should be applied to the assessment process, not postponed until an accommodation is denied. AI, test, privacy, and vendor clauses must be reviewed together: a legally designed alternative may still expose personal data if it is sent to a system not approved for that purpose.

## A practical privacy-control framework

The first step is to document the entire selection workflow, including résumé parsers, chat assistants, scheduling tools, test platforms, scoring, and video-analysis products. Create a data map that records each input, system of record, processor, location, retention period, access role, and downstream recipient. Ask whether an item genuinely predicts performance or merely makes the process feel technological. Where evidence is weak, eliminate the feature; collecting more data can increase legal exposure without improving hiring quality.

Before a pilot, conduct a job-relatedness review and examine subgroup effects. Establish pass scores and adverse-decision thresholds in advance rather than tuning them to a desired hiring outcome. Keep the assessment short, use a validated methodology, and ask an independent reviewer with relevant statistical and subject-matter expertise to examine differential accuracy and false-positive rates. Do not use a protected trait to “correct” scores manually, and do not treat small sample noise as proof of fairness—or dismiss persistent disparity merely because a vendor calls the model unbiased.

Vendor terms should limit data resale, model training on applicant information, advertising use, remote access, and secondary model training. They should require encryption, access logging, incident notice, defined breach procedures, deletion, export, audit cooperation, and clear controls for subprocessors. Set a practical incident-notification period, such as 24 to 72 hours for the vendor to notify the employer, while recognizing that the employer may need even less time to meet its own legal reporting deadlines. Confirm whether service logs include prompts, answers, audio, video, resumes, keystrokes, or identifiers that the intended assessment does not need.

## Comparison of assessment and privacy approaches

There is no single privacy model that fits every job. A structured, validated inventory test can collect less information than a live video interview, but validation and bias checks are still required. A human-only interview can reduce opaque automated scoring, although it may also produce inconsistent notes and interviewer bias. The right choice depends on job necessity, evidence, applicant population, cost, accessibility, and the organization’s ability to govern the tool.

| Feature | Validated non-AI or low-data assessment | AI-assisted psychological assessment | Social-media or open-web screening |
| --- | --- | --- | --- |
| Data collected | Structured job questions and work samples | Test answers plus possible voice, video, timing, or behavior signals | Public posts and information gathered across websites |
| Main privacy issue | Excess retention or access to applicant records | Opaque inference, recordings, profiling, and third-party reuse | Collection without context or applicant knowledge |
| Explanability | Usually easier when scoring rules are documented | Depends on whether meaningful factors are disclosed | Often difficult because relevance and source quality vary |
| Bias risk | Human judgment and test construction can be biased | Historical data and model design can reproduce proxy bias | Unequal access, cultural differences, and selective interpretation |
| Legal exposure | Privacy, ADA, and testing claims | Privacy, discrimination, biometric, consumer-protection, and accuracy claims | Privacy, reliability, discrimination, and off-duty conduct claims |
| Best use | Lower-stakes structured screening when validated | Carefully controlled assessments with strong notice and audits | Rarely justified; generally higher risk than conventional methods |
| Typical cost direction | Moderate per candidate, offset by simpler administration | Higher software, integration, validation, and governance costs | Potentially low initial price, but high review and litigation exposure |

A practical threshold is not a universal dollar figure but a governance test: if the tool cannot be explained, audited, corrected, and securely deleted, it is not ready for a consequential employment decision. The Center for Democracy & Technology’s work on third-party assessment stresses that assessment must be understood as a whole rather than treating a vendor certificate as proof of fairness. Psychological models may offer useful information, but their value cannot justify surveillance or uninformed inference.

## Common mistakes that make privacy worse

One common error is collecting social-media content because it is publicly visible. Public availability does not automatically make every post relevant to a current job, and an employer can learn personal information that has no legitimate connection to performance. In the United States, the EEOC’s 2022 guidance about social-media screening and religious beliefs illustrates how automated or targeted searches may identify protected activity even without religious intent. Employers should avoid protected-trait searches and use narrowly defined, job-related criteria rather than asking a language model to judge overall character.

Another mistake is promising that the system is “objective” or “bias-free.” No model is bias-free, and claiming complete neutrality can weaken the evidence needed to correct defects. Companies also make the mistake of deploying a pilot without a holdout comparison, using one vendor metric as proof of job performance, or ignoring applicants who withdrew before a final decision. Retention defaults, such as keeping unsuccessful applicants indefinitely, should be replaced by defensible schedules—potentially 12 to 24 months for routine recruiting records, depending on legal needs—while security incidents, litigation holds, or other justified exceptions are documented.

Finally, a human review stage can become a privacy illusion if recruiters cannot access the explanation, underlying data, or relevant test factors. Applicants should be able to challenge an error, request correction, and obtain human reconsideration before rejection. If a score affects a job offer, the employer should define a threshold in advance—for example, requiring a minimum validated score plus a documented work-sample result—not automatically selecting the highest algorithm rating. The model may summarize evidence, but it should not manufacture missing facts or infer sensitive conditions.

## When to act and what it may cost

Employers should act before the first pilot, not after candidates complain or a regulator asks for records. Privacy-by-design begins with a written purpose, a data inventory, and a decision to avoid features that are not needed. Organizations should also act immediately if a vendor announces a model update, changes its data-retention policy, begins using applicant data for training, or cannot explain a new score. A suspected discrimination pattern, inaccessible test, inaccurate inference, or data breach warrants pausing affected decisions while the facts are reviewed.

Pricing varies sharply. Conventional applicant-tracking or test-license fees may range from free to several hundred dollars per month for basic recruiting tools, while advanced video, facial, voice, or behavioral-analysis systems can cost tens of thousands to hundreds of thousands of dollars annually, sometimes plus per-candidate fees, integrations, validation studies, and legal review. These figures are market ranges rather than quoted prices, and buyers should obtain written pricing covering storage, support, audits, accessibility features, API use, and deletion. An inexpensive platform can become costly if it creates adverse-impact findings, manual accommodation work, incident response, or discrimination claims.

The strongest business case is proportional: use a psychological tool only when the job analysis identifies a decision that the tool can improve and the employer can test that benefit. Measure quality and privacy outcomes together, including subgroup selection rates, false positives, candidate complaints, override rates, time to delete records, and incident frequency. Review results at least annually and after any material model, data, or legal change. As of September 30, 2026, evolving federal, state, and local AI rules make continuous legal review more important than assuming a one-time vendor check is sufficient.

## What responsible use looks like

A defensible program tells candidates what is assessed, why, and how. It uses validated questions and work samples, collects the minimum necessary information, explains material scoring factors, and offers a meaningful route to human review. It tests fairness before deployment, monitors outcomes after deployment, deletes records on schedule, and suspends use when evidence suggests the tool is inaccurate or inaccessible. These controls do not guarantee that every hiring decision is correct, but they create a traceable process in which errors can be found and corrected.

Employers should also remember that privacy protects the applicant, while procedural fairness protects the integrity of the selection process. The two goals reinforce each other: applicants are more likely to provide honest, complete information when they know how it will be used, and recruiters make better decisions when they receive understandable evidence rather than an unexplained personality label. The correct question is therefore not simply whether AI can analyze human behavior. It is whether the employer has a lawful, job-related, proportionate reason to do so and can explain the consequences to the person being assessed.

## Quick answers

### Does an AI hiring test have to disclose the algorithm it uses?

It depends on the jurisdiction and the tool’s function, but responsible practice requires a meaningful explanation rather than a vendor name alone. New York City Local Law 144 can require candidates to receive notice and certain information about covered automated employment decision tools, while other laws may impose different disclosure duties. Candidates should at least be told that AI is involved, what data is used, and how to request human review.

### Can an employer use ChatGPT or another generative AI to review resumes?

Technically, a controlled tool may be used for some recruiting tasks, but use can create discrimination, confidentiality, accuracy, and data-transfer risks. Employers should prohibit uploading sensitive applicant information to unapproved accounts, restrict prompts and retention, test errors and subgroup outcomes, and require a qualified human to verify consequential decisions. A free consumer account is generally not an appropriate governance system for employment records.

### Are psychological test answers always confidential medical information?

No. They are not automatically medical records or protected health information under every privacy law, but a system may infer health-related or biometric characteristics from answers, voice, video, or behavior. Inferences can still be sensitive personal information, and biometric laws such as Illinois BIPA may apply when a tool collects defined identifiers. Employers need a specific legal analysis rather than relying on a broad confidentiality label.

### How long should employers retain unsuccessful applicant data?

There is no universal period, and the answer depends on jurisdiction, litigation holds, recruiting needs, and the type of information. A documented schedule—such as 12 to 24 months for ordinary unsuccessful applications—can be reasonable when justified, but raw video, biometric templates, and inferred profiles may need much shorter retention. Employers should separate necessary application records from optional AI inputs and delete both backup and vendor copies according to the approved schedule.

### Does a human reviewer eliminate AI hiring discrimination?

No. Human involvement helps provide accountability, but a recruiter who simply follows the highest score may reproduce the model’s error or bias. A meaningful review requires access to the relevant evidence, authority to disagree, training, and enough time to investigate. Employers should also measure override rates and outcomes by relevant applicant groups instead of assuming the presence of a recruiter proves fairness.

Canonical: https://psychprofile.io/knowledge/how_should_employers_protect_candidate_privacy_when_using_ai_hiring_tests.php
Markdown: https://psychprofile.io/knowledge/how_should_employers_protect_candidate_privacy_when_using_ai_hiring_tests.php/index.md
