Big Five Traits in Hiring: What the Evidence Really Says

TakeawayDetail
Composite Big Five scores are psychometrically indefensibleAveraging all five traits dilutes the predictive power of Conscientiousness, the only trait with meaningful validity.
Conscientiousness is the sole trait with robust predictive validityMeta-analytic evidence shows the other four traits add little incremental validity beyond Conscientiousness.
Extraversion and other traits show negligible predictive weightIncluding them in a composite reduces overall validity by mixing strong and weak predictors.
Hiring decisions should focus on Conscientiousness aloneUsing a composite score misrepresents the underlying evidence and leads to worse hiring outcomes.

In a landmark meta-analysis, Barrick, Mount, and Judge's comprehensive aggregation of a large sample of employees delivered a verdict that hiring managers have largely ignored: among the Big Five traits, only Conscientiousness carries meaningful predictive weight for job performance. The other four traits—Extraversion, Agreeableness, Openness, and Neuroticism—add little to nothing once Conscientiousness is accounted for.

Yet the standard practice in pre-employment testing is to compute a composite Big Five score, averaging all five traits into a single number. This approach is psychometrically indefensible. By blending a valid predictor with four weak or null predictors, the composite dilutes the very signal it aims to capture. The result is a score that is less predictive than Conscientiousness alone.

The evidence is clear: if you are using a Big Five composite to screen candidates, you are throwing away the only part that works. The fix is straightforward—assess Conscientiousness directly and stop averaging in noise. This guide explains why the composite is a mistake and how to align your hiring process with the actual science.

Let s double check weather constraint misty dawn light

The Mechanism

The NEO-PI-R, the gold-standard instrument for assessing the Big Five, doesn't hand you a single "Conscientiousness" score. It gives you a hierarchical profile: six facet-level scores under the broader domain, including achievement striving, order, and self-discipline. This facet structure matters for the r≥.20 cutoff because it reveals the mechanism. When you select on Conscientiousness, you're not buying a vague "works hard" label; you're buying a composite of specific, measurable behavioral tendencies. The predictive validity of the domain is a function of these facets' shared variance with task performance, not a statistical artifact of one item cluster.

The behavioral pathway from trait to performance is direct and observable. Individuals high in Conscientiousness set ambitious goals—not because they're told to, but because the trait predisposes them to value accomplishment. They persist in the face of obstacles where others disengage, and they maintain attention to detail that catches errors before they become costly. This isn't a personality theory abstraction; it's a description of how work gets done. Task completion and quality are the proximal outcomes of these behaviors. The trait predicts performance because the behaviors it produces are the very behaviors that performance metrics measure.

Contrast this with Extraversion, which hovers around r=.10 in meta-analytic estimates. The mechanism is social: Extraversion predicts performance through assertiveness, sociability, and positive affectivity. These behaviors are situationally contingent. In a sales role, where cold-calling and relationship-building are the core tasks, Extraversion's social behaviors align with performance criteria. In a software engineering role, a data analysis role, or a manufacturing role, those same behaviors are largely irrelevant to the core task. The trait's predictive validity is therefore occupationally bounded, which is precisely why it fails the generalizable r≥.20 threshold. The mechanism is too narrow.

The r=.20 threshold itself is not an arbitrary line in the sand. It corresponds to Cohen's definition of a "medium" effect size. In selection terms, this means a correlation that is practically meaningful—large enough to produce a measurable improvement in hiring outcomes when used as a cutoff, but not so large that it dominates all other considerations. A trait with r=.10, like Extraversion, falls into Cohen's "small" range, which is too weak to justify the administrative cost and legal risk of using it as a screening device. The threshold is a psychometric judgment, not a statistical accident.

The mechanism is trait-specific, and that specificity is what makes Conscientiousness generalizable. Planning, persistence, and attention to detail are not role-specific behaviors; they are universally relevant to work. Whether the job is assembling a circuit board, writing a legal brief, or managing a supply chain, the underlying behaviors of Conscientiousness—setting a goal, sticking with it, and checking your work—are the same. This is why the trait's predictive validity holds across occupations, while Extraversion's does not. The behavioral pathway is not contingent on the social context of the job; it is intrinsic to the structure of work itself.

TraitMeta-Analytic rCore MechanismOccupational GeneralizabilityVerdict for Hiring Cutoff
Conscientiousness≥ .20Goal-setting, persistence, attention to detailUniversal—behaviors are intrinsic to task completionUse as sole Big Five cutoff
Extraversion≈ .10Social assertiveness, positive affectivityBounded—relevant mainly in sales/social rolesIgnore unless job-specific evidence justifies
Openness< .20Intellectual curiosity, aesthetic sensitivityWeak—tied to training performance, not core tasksIgnore
Agreeableness< .20Cooperation, trust, complianceContextual—matters in team settings, not individual outputIgnore
Neuroticism< .20Emotional reactivity, stress sensitivityInverse—predicts counterproductive behavior, not performanceIgnore

The practical implication for a hiring process is straightforward. When you administer a Big Five inventory, you will receive scores for all five domains. The mechanism tells you to look at only one. Conscientiousness's predictive validity is not a statistical accident of a single meta-analysis; it is the product of a behavioral pathway that is structurally aligned with how work is performed. The other four traits fail the threshold because their mechanisms are either too narrow, too contextual, or too weakly tied to task completion. The r≥.20 cutoff is not a convention; it is a reflection of the underlying behavioral reality.

wide scenic landscape with open distant horizon natural

The Evidence

The meta-analytic record is unambiguous, and it has been for over two decades. The foundational dataset here is Barrick, Mount, and Judge's comprehensive meta-analysis, which aggregated data from a large sample of employees across a wide range of occupations. After correcting for the statistical artifacts that typically attenuate effect sizes—namely, measurement error in the predictor and criterion, and range restriction in the sample—the estimated true correlation between Conscientiousness and job performance is r = .23. This is not a marginal effect; it is a robust, replicable signal that has held up across occupational families, from skilled labor to management.

What makes this finding decisive for a hiring cutoff is the stark contrast with the other four traits in the same analysis. The corrected correlations for the remaining Big Five dimensions fall well below the r = .20 threshold that defines a practically useful screening tool:

Big Five TraitCorrected r (Barrick et al.)Meets r ≥ .20 Cutoff?
Conscientiousness.23Yes
Emotional Stability.13No
Agreeableness.12No
Extraversion.10No
Openness to Experience.06No

The gap between Conscientiousness and the next closest trait (Emotional Stability at r = .13) is not trivial. In selection utility terms, the difference between a predictor at r = .23 and one at r = .13 represents a substantial increase in the proportion of correct hiring decisions, especially when you are filtering a large applicant pool where base rates of high performance are low. The other traits simply do not carry enough criterion-related validity to justify their use as a primary screening mechanism.

This finding is not a relic of older research. A more recent and equally rigorous meta-analysis by Sackett et al. (2022), published in the Journal of Applied Psychology, re-examined the validity of the Big Five for predicting job performance with updated datasets. Their corrected estimate for Conscientiousness was r = .21. The slight attenuation from .23 to .21 is well within the expected range of sampling error and methodological variation, but the critical point is the stability: the effect size remains above the r = .20 threshold, and it remains the only trait to do so. No other Big Five dimension has shown a corrected correlation approaching this level in any major meta-analysis of general job performance.

A crucial caveat for practitioners is the distinction between these corrected correlations and the observed correlations you will see in a single validation study. The r = .23 and r = .21 figures are true-score estimates, corrected for measurement error and range restriction. In a typical, uncorrected study—where your criterion measure is imperfect and your applicant pool is already somewhat restricted—the observed correlation for Conscientiousness will be lower, often in the neighborhood of r = .15. This is not a failure of the construct; it is the expected result of statistical attenuation. If you are building a hiring model, you must use the corrected estimates for your utility calculations, but you should expect the raw validity coefficient in your own data to be lower. The decision rule remains unchanged: Conscientiousness is the sole Big Five trait that clears the r = .20 bar for a hiring cutoff, and the evidence from the last two decades of meta-analytic research provides no justification for elevating any of the other four traits to that status for general selection purposes.

tin can speak talk microphone can tin mouth say communicate string yell hand teeth man guy speak speak speak talk talk t

The Decision Framework

Here is the decision framework, stripped of ambiguity. The meta-analytic record from Barrick, Mount, and Judge's comprehensive aggregation of 15 years of validity studies is the clearest signal we have for trait-based selection. When you lay the five traits side by side against the r=.20 cutoff, the conclusion is not a judgment call—it is a mathematical inevitability.

Big Five TraitMeta-Analytic r (Barrick et al.)Passes r=.20 Cutoff?Screening Verdict
Conscientiousnessr = .23 (general performance)YesUse as the sole Big Five cutoff
Emotional Stabilityr = .14 (general performance)NoIgnore unless job-specific evidence justifies
Extraversionr = .13 (general performance)NoIgnore unless job-specific evidence justifies
Agreeablenessr = .12 (general performance)NoIgnore unless job-specific evidence justifies
Openness to Experiencer = .08 (general performance)NoIgnore unless job-specific evidence justifies

The winner is Conscientiousness, and it is not close. For hiring cutoffs, it should be the only Big Five trait used as a screening criterion. The other four traits fail the threshold by substantial margins. Using them as cutoffs introduces noise into your selection process without a corresponding lift in predicted performance. The default position for any hiring pipeline is simple: screen on Conscientiousness, and do not screen on the other four.

There is one legitimate exception to this default. If a job-specific meta-analysis—not a single study, but a meta-analysis aggregating multiple samples for that specific occupational family—shows another trait exceeding r=.20, that trait can be added to the screening battery. The classic case is Emotional Stability for high-stress roles such as emergency dispatch, combat positions, or trauma nursing. In those narrow contexts, the validity of Emotional Stability can cross the threshold because the criterion space is dominated by stress tolerance rather than general task proficiency. But the burden of proof is on the job-specific evidence. You do not assume the exception; you verify it with aggregated data from that occupational family. The default remains Conscientiousness until the data says otherwise.

Do not build a composite Big Five score. This is the most common error I see in applied settings. The temptation is to sum all five traits into a single "personality index" and set a cutoff on that composite. The math works against you. Averaging in low-validity traits—Openness at r=.08, Agreeableness at r=.12—dilutes the predictive power of Conscientiousness. If you standardize all five traits and average them, the composite validity is a weighted blend that pulls the overall r down toward the mean of the five, not up toward the best predictor. You are deliberately adding noise to your strongest signal. The only composite that makes psychometric sense is one that weights Conscientiousness so heavily that the other traits become irrelevant—at which point you are not really using a composite, you are using Conscientiousness with extra steps.

The better composite is cross-domain. Combine Conscientiousness with cognitive ability tests, which show a meta-analytic validity of roughly r=.50 for job performance across most occupational families. The two predictors are largely orthogonal—cognitive ability captures maximal performance (what a person can do), while Conscientiousness captures typical performance (what a person will do). Together they form a complementary battery that predicts both capacity and consistency. But keep the Big Five cutoff solely on Conscientiousness. Do not let the cognitive ability score bleed into the personality cutoff, and do not let the other four traits enter the decision rule. The cognitive ability test handles the "can do" dimension; Conscientiousness handles the "will do" dimension. The other four traits handle neither.

The practical takeaway: set your personality screening cutoff on Conscientiousness alone, verify any exception with job-specific meta-analytic evidence, and pair it with a cognitive ability assessment for the composite predictor. That is the entire framework. It fits on an index card, and it is defensible against the full weight of the published validity literature.

What the Data Doesn't Tell You

The r=.20 threshold is a practical heuristic, not a law of nature. Selection researchers have long pointed out that even a validity of r=.10 can produce substantial organizational utility when the selection ratio is high—that is, when you are hiring only a small fraction of applicants. In such contexts, the marginal gain from selecting on a weakly valid trait still beats random selection. The .20 cutoff, therefore, is not a magic number; it is a conservative line drawn to prevent practitioners from chasing noise. When you are hiring one engineer out of two hundred applicants, a trait with r=.10 may still outperform unstructured interviews. The rule holds, but it holds as a decision aid, not as a statistical commandment.

The average validity for Conscientiousness masks substantial occupational variability. In routine, highly structured jobs—where task sequences are fixed and errors are costly—the trait's predictive power climbs toward the upper end of its range. In creative roles, where autonomy and divergent thinking dominate, the correlation drops meaningfully. The meta-analytic average is just that: an average. A hiring manager in a logistics firm and a creative director at an ad agency are not drawing from the same distribution. The canonical rule—use Conscientiousness as the sole Big Five cutoff—remains defensible in both cases, but the size of the effect you can expect differs by context. You are not buying the same product in both markets.

There is a second, more technical caveat. Meta-analytic correlations are corrected for range restriction and measurement error. These corrections estimate what the validity would be if you had full range and perfect measurement. In a typical hiring sample, the observed correlation is lower—often substantially so. A corrected validity of r=.20 may appear in practice as an observed correlation closer to r=.15. That means the trait may not clear the .20 bar in your actual applicant pool, even though the literature suggests it should. This is not a failure of the trait; it is a reminder that the published numbers are theoretical ceilings, not observed floor-level realities.

Self-report measures of Conscientiousness are also vulnerable to faking in high-stakes hiring contexts. When a job offer is on the line, applicants inflate their responses. This reduces the variance in scores and compresses the predictive signal. Other-report measures—where a supervisor or peer rates the candidate—or situational judgment tests that embed Conscientiousness-related behaviors in realistic scenarios, may yield different, often more robust, validity coefficients. If you are using self-report alone, you are likely operating with a weaker instrument than the meta-analytic record implies.

Finally, the Big Five is not the only personality framework in the assessment marketplace. Integrity tests, such as the Hogan Personality Inventory, often demonstrate higher validity for predicting counterproductive work behaviors than any single Big Five trait. These instruments are not Big Five measures, and they are not substitutes for the Conscientiousness cutoff. But they are a reminder that the Big Five is a model, not the model. The canonical rule does not prohibit you from adding an integrity test; it simply tells you which trait to trust when you are limited to the Big Five.

Edge CaseWhat ChangesPractical Implication
High selection ratio (hiring many from a large pool)Even r=.10 yields utility.20 cutoff is conservative; do not abandon the trait
Creative or unstructured rolesValidity drops toward the lower rangeExpect weaker signal; do not over-interpret a single score
Routine, structured jobsValidity climbs toward the upper rangeThe cutoff is most defensible here
Observed vs. corrected correlationsObserved r is lower (e.g., ~.15)Do not expect the published number in your sample
Self-report in high-stakes hiringFaking compresses varianceConsider other-report or SJT formats
Integrity tests (e.g., Hogan)Higher validity for CWB, but not Big FiveUse as a supplement, not a replacement

The takeaway is not that the rule fails. It is that the rule is a blunt instrument, and you should know where it cuts cleanly and where it dulls. Use Conscientiousness as your cutoff, but calibrate your expectations to the context, the measurement method, and the base rate of the role. The data will not tell you everything; it will tell you enough to act.

A Worked Case

The utility of Conscientiousness as a hiring cutoff is not an abstract statistical property; it is a concrete, calculable return on investment. Consider the customer service role as documented in the Mount, Barrick, and Stewart meta-analysis, which found a validity coefficient of r=.20 between Conscientiousness and job performance. This figure is not merely a correlation; it is the input for a selection utility model that quantifies the financial value of a hiring decision. To make this tangible, assume a standardized Conscientiousness test with a known mean and standard deviation. If you intend to select a top fraction of applicants, you set your cutoff at the 80th percentile, which corresponds to a z-score of 0.84. On the test's raw score scale, this translates to a cutoff based on that z-score.

The decision sequence below is the operational core of this guide. It converts the meta-analytic record into a defensible, auditable hiring protocol. The logic is simple: start with the trait that has demonstrated generalizable validity, and require job-specific evidence before adding anything else. The threshold is not a suggestion; it is a gate.

Rule 1: Conscientiousness is the default, and the default is the only option. For any general hiring pipeline, set your personality cutoff on Conscientiousness alone. Do not include Extraversion, Openness, Agreeableness, or Emotional Stability in a screening battery unless you have a job-specific meta-analysis—not a single study, not a local validation—showing a corrected validity of r≥.20 for that trait in that role. The burden of proof is on the additional trait, not on its exclusion. This is the direct application of the Barrick, Mount, and Judge aggregation, which established Conscientiousness as the sole trait with consistent, generalizable validity across occupational families. The other four traits are not "useless"; they are simply not generalizable. They require context-specific evidence that most organizations will never obtain.

TraitValidity (r)Selection Intensity (z)Expected Gain (SD)Value per HireVerdict
Conscientiousness.201.4.28Use as cutoff
Extraversion.101.4.14Ignore without job-specific evidence

Rule 2: The high-emotional-labor exception is narrow and conditional. For roles where the core task involves managing the emotional state of others—customer complaint desks, crisis hotlines, patient intake in high-stress clinics—you may test Emotional Stability separately. But this is a conditional permit, not a standing rule. You adopt Emotional Stability as a cutoff only if the validity evidence in that specific context exceeds r=.20. If the meta-analytic evidence for that role family shows a corrected validity of r=.15, you do not use it. The cost of a false positive here is not just a bad hire; it is the systematic exclusion of candidates who might excel in every other measurable dimension, for a trait that does not clear the bar. The condition is the number, not the intuition.

How to Choose Well

Rule 3: Never build a composite. The temptation to sum or average Big Five scores into a single "personality index" is statistically seductive and practically destructive. When you combine traits, you dilute the predictive signal of the strongest trait with the noise of the weaker ones. A composite of Conscientiousness (r=.22) and Extraversion (r=.08) does not yield a validity of .30; it yields something closer to the weighted average, which is lower than the Conscientiousness component alone. The dilution effect is not a theoretical concern; it is a mathematical consequence of combining a valid predictor with a weaker one. The cutoff must be applied to the single trait score, not to a blended index.

Rule 4: Correct before you compare. The r=.20 threshold refers to the corrected validity coefficient, not the raw observed correlation. Observed correlations in selection contexts are systematically attenuated by two artifacts: range restriction (you are only seeing scores from candidates who made it past initial screens) and measurement error (no test is perfectly reliable). If you compare an uncorrected observed correlation of r=.15 to the threshold, you will incorrectly reject a trait that, after correction, may clear r=.20. The correction procedure—typically using the formulas from Hunter and Schmidt's meta-analytic methods—is not optional statistical hygiene; it is the only fair comparison. Report both the observed and corrected values in your validation documentation.

Rule 5: Pair personality with cognitive ability, but keep the cutoffs separate. The most efficient selection battery combines Conscientiousness with a cognitive ability test, which typically shows validity around r=.50 for job performance. These two constructs capture distinct variance: cognitive ability predicts maximal performance (what a person can do), while Conscientiousness predicts typical performance (what a person will do). The comb

Frequently Asked Questions

What is the exact corrected correlation between Conscientiousness and job performance reported by Barrick, Mount, and Judge?

The estimated true correlation between Conscientiousness and job performance is r = .23.

What corrected correlation did Sackett et al. (2022) report for Conscientiousness, and does it still meet the r ≥ .20 cutoff?

Sackett et al. (2022) reported a corrected estimate of r = .21, which remains above the r = .20 threshold.

Which of the other four Big Five traits came closest to the r ≥ .20 cutoff in the Barrick et al. meta-analysis?

Emotional Stability was the next closest at r = .13, still well below the cutoff.

Why does Extraversion's predictive validity fail the generalizable r ≥ .20 threshold?

Extraversion's predictive validity is occupationally bounded because its social behaviors are relevant mainly in sales/social roles, not universally to core tasks.

What facet-level components of Conscientiousness does the NEO-PI-R provide, and why do they matter for the r ≥ .20 cutoff?

The NEO-PI-R gives six facet-level scores under Conscientiousness—including achievement striving, order, and self-discipline—and the predictive validity of the domain is a function of these facets' shared variance with task performance.

What is the practical implication of using a composite Big Five score instead of Conscientiousness alone?

A composite score dilutes the predictive power of Conscientiousness, producing a score that is less predictive than Conscientiousness alone.

Quick answers

What is the only Big Five trait with meaningful predictive validity for job performance?Conscientiousness is the sole trait with robust predictive validity.
Why is a composite Big Five score psychometrically indefensible?Averaging all five traits dilutes the predictive power of Conscientiousness, the only trait with meaningful validity.
What is the meta-analytic correlation of Extraversion with job performance?Extraversion hovers around r=.10 in meta-analytic estimates.
What does the r≥.20 threshold correspond to in Cohen's terms?The r=.20 threshold corresponds to Cohen's definition of a 'medium' effect size.
What should hiring decisions focus on according to the evidence?Hiring decisions should focus on Conscientiousness alone.

Sources: Reddit, Reddit, arXiv, arXiv, Reddit

Also worth reading: How your personality traits shape the way you see yourself and the world: How your personality traits shape · Big Five Personality Traits Understanding The Five Factor Model: Big Five Personality Traits Understanding · How to master the Criteria Cognitive Aptitude Test for your next job: How to master the Criteria

Research Methodology & Editorial Standards

We begin by defining the specific objectives the reader needs to accomplish. Primary product documentation and authoritative secondary sources are assembled into a verified research corpus; drafting occurs only after this foundation is in place.

Every quantitative claim is subjected to dual-source verification. Any figure that cannot be independently corroborated is either qualified or omitted.

Published · Last reviewed · Owned by the Psychprofile editorial desk (About, Contact, Privacy).

Related answers