What the research says about AI bias in employee performance reviews

Managers should not treat AI as the final judge of employee performance. It can help organize evidence, draft summaries, identify missing information, and flag inconsistencies, but the manager should remain responsible for the rating, the explanation, and the employment consequence. As of 24 September 2026, the main concern is not that every AI system is intentionally biased; it is that systems can reproduce bias from past ratings, job duties, language, workplace visibility, and the assumptions built into their training data. A recent Newsweek report said that half of companies could already be using AI in performance reviews, although that figure was presented as an estimate rather than a complete census. The practical question is therefore not whether AI can review performance, but whether the organization can show how the system reached its result and how a person challenged it.

Also worth reading: What are the key AI employee surveillance ethics laws and regulations that companies need to know about in 2026? · What is algorithmic bias in hiring audits and how do companies detect it? · How Can Healthcare Organizations Systematically Reduce Algorithmic Bias in Clinical Machine Learning Models?

The available research gives mixed messages. Some vendors argue that standardized algorithms can reduce inconsistent manager judgments, while HBR coverage on middle managers suggests that adoption often succeeds or fails according to whether managers trust and use the tool correctly. Other reporting, including work from FindLaw and Observer, describes legal and ethical risks when AI-driven employment decisions are opaque or involve employee surveillance. A lawsuit reported by FindLaw involving alleged medical and family-leave bias in AI-related layoff decisions shows why leave-related data deserves special scrutiny, even though a complaint is not the same as a court ruling. The safest position is to use AI for assistance under human control, document the evidence, and test whether the tool changes ratings in ways that reflect job performance rather than personal similarity to a past group of successful employees.

How bias enters an AI performance-review system

Bias can enter before the model is used. Historical performance data may reflect biased promotion decisions, uneven assignments, different standards for speaking up, or a tendency to reward employees who work in highly visible roles. If the model learns from those records, it may treat the old pattern as a reliable rule. For example, an algorithm trained mostly on employees who received high ratings after working long hours could learn that long hours predict strong performance, even when the organization does not intend to reward overwork. This is a form of proxy bias: the model may use a neutral-looking feature that stands in for a protected or socially influenced characteristic.

The data used for an individual review can also be incomplete. Email activity, ticket volume, calendar events, project deadlines, customer feedback, and manager comments each measure only part of work. Employees who spend time on mentoring, accessibility support, caregiving, safety work, or internal collaboration may appear less productive in a system that counts messages or visible output. A model can also respond differently to writing styles, accents, names, gender markers, and language choices. Research cited by Communications of the ACM on cognitive and AI biases, along with Anthropic work showing that model behavior can vary by model version and language, supports the view that apparent neutrality is not enough. A system needs testing for different groups, different work contexts, and different types of legitimate performance.

Feedback loops create another problem. Once an AI-generated rating affects development opportunities, compensation, promotion, or termination, that result may become a new input to the next review cycle. Employees who receive fewer opportunities then have less visible evidence of success, which can make them appear weaker in later ratings. The system can therefore convert a small initial difference into a larger one over several review periods. Companies should monitor the full chain from evidence to recommendation, rather than checking only whether the final score matches a manager's opinion.

What AI can and cannot do fairly in a review

AI is well suited to tasks that require sorting or summarizing large amounts of existing evidence. It can group comments by theme, convert free-text notes into a draft summary, compare stated goals with documented outcomes, and list missing evidence before a meeting. These tasks can save manager time and make reviews more consistent when the source material is clear. They are less suitable for deciding intent, judging character, diagnosing a mental condition, or deciding that one employee deserves a lower rating than another. Nature's coverage of research on AI analysis of human behavior and prediction of personality traits should not be read as proof that a model can reliably infer a person's character from workplace traces.

A fair system should separate three questions: what happened, how was it measured, and what decision will follow. AI can help with the first two, but it cannot remove uncertainty from the third. A low number of completed tickets, for example, might reflect a difficult quarter, a change in role, a broken data feed, or a genuinely weaker result. The model should present the observation and the possible explanations, not silently convert the observation into a judgment. If a worker asks for an accommodation, discusses a health matter, or takes protected leave, the system should route that information to a trained human rather than treat it as a performance signal.

The strongest performance systems also avoid pretending that one score can represent an entire year of work. A review should combine multiple evidence types, state confidence levels, and identify where human judgment is needed. Any output that changes a rating by 10 percent or more should trigger a documented check, especially when the change relies on a single metric. That threshold is an internal control rather than a universal legal standard, but it gives reviewers a concrete point at which to pause and investigate.

A safer operating model for managers and HR teams

The first rule should be that AI cannot make the final employment decision. A manager may accept a model-generated summary, but the manager must verify each claim against source records, add relevant context, and explain the rating in ordinary language. HR should set this rule in policy and in the tool's permissions, not only in training slides. The system should show the evidence behind every recommendation and provide an easy way for an employee to correct inaccurate data. A correction process without a real review is only a form, not a safeguard.

The second rule is that inputs must match the job. A system should not use social media activity, off-hours messaging, personality labels, or unrelated personal information to score performance. If the model uses communication data, the company should define which communications relate to the role and exclude private, protected, or legally restricted material. It should also record what data was missing. Missing evidence should lower confidence in the recommendation, not automatically count as poor performance.

Calibration meetings should compare AI recommendations across departments, levels, and demographic groups. Reviewers should examine at least 20 decisions each quarter, or a smaller sample if the organization is small, and look for cases where the model consistently rates comparable work differently. HR can calculate rating distributions, false positives, false negatives, and the share of recommendations changed after human review. Middle managers need training on how to question a recommendation, not just how to operate the software. HBR's point about middle managers is relevant because they control the daily use of the tool and determine whether employees believe the process is fair.

Psychological profiles are a separate and higher-risk input

A psychological profile should not be treated as a hidden component of a performance score. Personality can affect communication preferences or collaboration styles, but the connection between a trait and future job performance is often weaker than vendors imply. A profile generated from chat messages, emails, or observed behavior can mistake temporary stress, cultural differences, disability-related communication patterns, or unfamiliar language style for a stable trait. It should not be used to infer a mental-health condition, to rank an employee for promotion, or to predict who will fit a manager's preferred personality.

If a psychological product is used, its purpose should be voluntary self-reflection or team education, with consent and data minimization. The employee should know what was collected, how long it is retained, who can see it, and whether declining affects the review. A self-reported profile should be kept separate from manager ratings unless the employer has credible, job-related evidence and a lawful reason to use it. Psychometric tools used in employment generally require validation for the role and population; generic personality or mental-health predictions are not equivalent to a validated selection test. The burden of proof should sit with the employer, not the employee.

This distinction matters for products associated with psychprofile.io or similar services. A profile can prompt a conversation about work style, but it should not become an invisible input that changes compensation, promotion, or layoffs. If a profile conflicts with a person's actual performance record, the performance record and documented job evidence should take priority. A useful rule is to ask whether the information improves a work conversation or merely gives a system more data to make a personnel decision. When the answer is unclear, keep the profile out of the employment workflow.

Comparison of review models and alternatives

The best alternative depends on the task being purchased. A manual structured process is slower but easier to explain, while a controlled AI assistant can reduce administrative work if it has clear evidence and human approval. Fully automated ranking offers speed but creates the greatest risk when the model is opaque and the stakes are high.

FeatureHuman-led structured reviewAI-assisted review with controlsAutomated AI ranking
Main strengthContextual judgment and accountabilityFaster summaries and evidence organizationHigh-volume processing and apparent consistency
Main weaknessManager fatigue and inconsistent standardsModel error, input bias, and overrelianceOpaque reasoning and weak employee recourse
ExplainabilityUsually high if notes are keptMedium when sources and confidence are shownOften low
Best useFinal ratings and development decisionsPreparation, drafting, and gap detectionLow-risk administrative triage only
Required safeguardCalibration and documented evidenceHuman approval, audits, and appeal pathNo use as a sole basis for employment decisions
Psychological profilesOptional and separateExcluded unless validated, consented to, and job-relatedShould not be used
The table shows why replacing managers with a score is usually a poor trade. AI assistance is most defensible when it reduces clerical work while leaving accountability intact. A fully automated ranking may be acceptable for sorting non-sensitive requests, such as grouping duplicate records, but not for deciding who deserves a raise, a promotion, or a termination. Companies with strong existing rubrics may gain less from automation than companies struggling with inconsistent manager records, although weak historical data can make automation more risky in either case.

A practical 90-day plan for reducing bias

During the first 30 days, define the review purpose and the prohibited uses before selecting a vendor. Write a plain-language policy that says AI may draft or organize but may not make final employment decisions. Identify protected characteristics, leave records, accommodations, and sensitive complaints that must be excluded from ordinary performance inputs. Map every field the system uses to a job-related reason, and remove features that merely measure visibility, personality, or personal similarity to past high performers. Keep a record of the data sources, model version, and settings so that a later decision can be reproduced.

From days 31 to 60, run a retrospective test on a sample of prior reviews rather than testing only new cases. Compare AI output with the original rating, manager comments, role, level, and documented outcomes. Check whether groups with similar performance evidence receive different recommendations, and whether leave, disability, language, caregiving, or part-time work is acting as a negative proxy. Set an investigation threshold for any 10 percent rating swing, a major ranking change, or a recommendation based only on missing data. Ask a manager outside the employee's reporting line to review a sample of at least 20 cases if the workforce is large enough.

From days 61 to 90, pilot the system with a limited group and a human comparison group where possible. Train managers to challenge outputs, ask employees for correction, and record the reason for accepting or rejecting each recommendation. Measure time saved separately from rating quality, because faster reviews are not necessarily better reviews. Review error rates monthly at first, then quarterly once the process is stable. If a person can overturn the system without penalty, that is a sign the control is real; if the system consistently wins the argument, the organization has created an unaccountable decision-maker.

Common mistakes, costs, and reasons to act now

The most common mistake is confusing consistency with fairness. A model may produce the same score for many employees because it is applying the same biased rule to everyone. Another mistake is assuming that vendor claims of bias reduction count as proof. A vendor should provide validation data, subgroup testing, error definitions, retention rules, and an explanation of what happens when the source data is wrong. Companies also err by allowing employees to discover the system after ratings are finalized, by treating a correction button as meaningful review, and by using AI scores for discipline without an independent human process.

Pricing for performance-review AI is rarely comparable because some products charge per user, others per review volume, and enterprise contracts bundle data storage, integrations, support, and consulting. A small pilot may use existing software and staff time, while a larger deployment can require legal review, data engineering, security testing, training, and ongoing audits. The hidden cost is often governance: someone must investigate exceptions, update rubrics, handle appeals, and retire features that create legal risk. A tool that saves two hours per manager but adds 20 hours of dispute resolution each quarter may not be economical.

Act immediately when a system already influences promotion, compensation, termination, leave, or surveillance, when employees cannot see or challenge its inputs, or when a complaint or regulator inquiry alleges discrimination. In those situations, pause the affected use, preserve records, and obtain advice from employment counsel and the relevant works council or union. Do not wait for a perfect audit before removing an opaque score from a high-stakes decision. Conversely, a low-risk internal drafting tool can proceed with a 90-day trial, clear rules, and a human review requirement. The deciding threshold is not how advanced the AI appears; it is whether the organization can explain, test, and correct the employment decision when the model is wrong.