Social Desirability in Hiring: Veto Above 70th Percentile by 2026

TakeawayDetail
Social desirability is a threshold phenomenon, not a continuous trait.Beyond a critical point, the scale suppresses trait variance and measures profile selection.
The correct policy is a pre-registered hard veto at the upper tail.Regression corrections and caution notes fail to address the concentrated inflation.
Applicant-incumbent differences in Conscientiousness are small on average but mask a concentrated inflation.The inflation is driven by candidates producing a profile-selection artifact.
The field's 'consider with caution' approach is insufficient.The point where the scale stops being a response style indicator is the cutoff.

A study titled 'Comparing chatbots to psychometric tests in hiring' reported reduced social desirability bias but lower predictive validity—a paradox that exposes the field's fundamental misreading of social desirability as a continuous trait to be corrected or ignored.

The correct interpretation is that social desirability behaves as a threshold phenomenon. Beyond a critical point, the scale stops measuring response style and starts independently suppressing trait variance. This is the point where the inflation of Conscientiousness is concentrated, and it is exactly where the scale's utility collapses. The same data that show a small average inflation also reveal that the inflation is driven by candidates in the upper tail of the distribution—those who are more likely producing a profile-selection artifact than a personality signal.

Going forward, the only defensible personnel policy is a pre-registered hard veto at that threshold—not a regression correction, not a 'consider with caution' note. The field must recognize that candidates above this point are more likely producing a profile-selection artifact than a personality signal. The top segment of the social desirability distribution is where the scale stops measuring response style and starts suppressing trait variance.

vast glass walled corporate lobby dawn pale gray light

The Faking Curvature

In applicant screening, a social desirability (SD) scale is not a personality scale; it is a count of implausible endorsements. The current editions of the Hogan Personality Inventory (HPI) "Ideal Employee" scale and the SHL OPQ32 impression-management score are calibrated on applicant norms, not general-population norms. That calibration choice matters because it shifts the reference distribution: a score that looks unremarkable against the general public is already extreme within a pool of people actively trying to appear virtuous. The scale is designed to detect the *act* of claiming perfection, not to measure a stable trait.

The mechanism is the absolute-statement design. A candidate must deny universal, minor human flaws—items like "I always admit my mistakes" or "I have never taken credit for someone else's work." An honest person cannot endorse these items consistently because no one is that flawless. To score at or above the cutoff, a candidate must sustain a context-specific endorsement pattern that has no honest baseline. This is not a matter of mild self-enhancement; it requires a systematic refusal to acknowledge any imperfection across dozens of items. The pattern is qualitatively different from the occasional "yes" that an honest candidate might give to a slightly flattering item.

Faking is a nonlinear displacement. Viswesvaran and Ones meta-analyzed faking experiments and reported an effect for Conscientiousness. That effect is enough to move a candidate who is genuinely below average into the upper portion of the distribution. The implication is stark: a candidate who is genuinely below average on Conscientiousness can, with modest effort, present as above average. The displacement is not uniform across the distribution—it compresses at the top, where ceiling effects and item difficulty make further gains harder, but it is most dangerous in the middle range where most hiring decisions are made.

Once a score passes the cutoff, the SD scale acts as a suppressor variable. It absorbs variance from the intended trait, meaning two candidates with equal observed Conscientiousness scores at different SD levels are not equivalent on the underlying trait. The high-SD candidate's observed score is inflated by response distortion, while the low-SD candidate's score reflects genuine trait variance. The common belief—that you can subtract the SD score from each personality dimension to get a "corrected" estimate—fails precisely here. Above the cutoff, the SD scale is no longer a covariate; it is a suppressor that has consumed the trait variance itself. The corrected estimate is still unreliable because the distortion is not additive but interactive.

The cutoff is the point where the faked fraction exceeds the honest fraction in most applicant pools. This is why a higher cutoff only catches extreme fabricators—those who are so blatant that they are likely to be detected anyway—and why a no-cutoff policy leaves more invalid variance than valid trait variance in the ranking model. At the cutoff, the signal-to-noise ratio inverts. Below it, the SD scale correlates modestly with trait scores and can be statistically managed. Above it, the SD scale dominates the shared variance, and any personality score becomes uninterpretable. The only defensible treatment is invalidation.

Cutoff PolicyWhat It CatchesWhat It MissesVerdict
No cutoffNothingAll fakers; invalid variance exceeds trait variance in ranking modelUnacceptable
A higher cutoffExtreme fabricatorsModerate fakers who still distort trait scoresInsufficient
The cutoffAll candidates where faked fraction exceeds honest fractionNone—scores above this are uninterpretableCorrect

The practical consequence for hiring pipelines is that pre-registering the cutoff is not a conservative choice; it is the only choice that preserves the interpretability of the personality assessment. A higher cutoff gives the illusion of rigor while allowing precisely the candidates who are most likely to have distorted their scores to remain in the pool. A no-cutoff policy is worse—it treats the SD scale as noise when it is, in fact, the dominant signal. Pre-registration matters because it prevents post-hoc rationalization when a high-SD candidate is otherwise attractive. The rule is simple: do not hire any candidate whose validated social desirability score reaches or exceeds the cutoff on the test's applicant norm table; code the entire personality assessment as invalid for that candidate.

narrow rain slicked corridor towering office building twilight diffused

Evidence

The psychometric case for a hard veto at the cutoff rests on four independent lines of evidence, each addressing a different failure mode of self-report inventories under applicant conditions. The first is the measurement precision of the social desirability construct itself. Crowne and Marlowe built the Social Desirability Scale and reported high split-half and test-retest reliability. That precision is the floor that makes a fixed cutoff defensible: if the SD scale were noisy, any cutoff would be arbitrary. But with reliability in that range, the scale is stable enough that a score at or above the cutoff on applicant norms is a reliable signal of response distortion, not measurement error. This is the foundation—without it, the entire veto framework collapses into guesswork.

The second line of evidence concerns which component of social desirability you should actually measure. Li and Bagger reliability generalization found Impression Management to be more reliable than Self-Deceptive Enhancement on the Balanced Inventory of Desirable Responding (BIDR). The cutoff should be computed on the more reliable Impression Management component. The logic is straightforward: Impression Management is a deliberate, conscious presentation strategy—it is the part of SD that a candidate controls and deploys during a job application. Self-Deceptive Enhancement, by contrast, is an unconscious positive self-view that is less responsive to situational demands and less reliable as a measure. If you blend the two into an SD index, you dilute the signal with the less reliable component. The veto should therefore be applied to the Impression Management subscale specifically, not to a composite that includes Self-Deceptive Enhancement.

The third line of evidence shows that the mean inflation in applicant samples is modest but strategically concentrated. Birkeland et al. meta-analysis found job applicants score higher than incumbents on Conscientiousness. A naive reader might look at that modest average difference and conclude that faking is a minor nuisance that can be safely ignored or corrected. That conclusion is wrong. The average difference is modest but concentrated among the same high-SD applicants who are caught by the cutoff veto. The distribution is not shifted uniformly; it is skewed by a subset of applicants who inflate their scores substantially. That average is the average of a large group of honest respondents and a small group of heavy fakers. The heavy fakers are precisely the ones who land above the cutoff on SD. The modest mean hides the tail, and the tail is where the damage to selection validity occurs.

The fourth line of evidence is the most consequential for the "correct, don't veto" crowd. Schmit and Ryan showed that applicant samples produce a different factor structure on an omnibus Big Five inventory, with Conscientiousness items grouping with an impression-management factor. This is not a matter of simple mean shift; it is a structural change in the covariance matrix. When Conscientiousness items load onto an impression-management factor in applicant samples, the trait items are no longer measuring the same latent construct they measure in incumbents. This structural shift is the psychometric reason a high SD score cannot be partialled out safely. Partialling assumes the SD scale is a covariate—a variable that shares variance with the trait but does not change the trait's factor structure. But when the factor structure itself changes, SD is no longer a covariate; it is a suppressor. A suppressor variable shares variance with the error component of the trait measure, and using it in a partialling equation can inflate the corrected estimate rather than clean it. The only defensible treatment is invalidation.

Consider the practical implication of the Schmit and Ryan finding. If Conscientiousness items group with an impression-management factor in applicants, then a high SD score is not merely an additive bias on an otherwise valid trait estimate. The trait estimate itself is contaminated at the item level. The items are functioning differently in the applicant sample than they do in the incumbent sample. No statistical correction can undo that item-level contamination because the correction would need to know the applicant-specific factor loadings, which are unknowable at the individual level. The cutoff veto is the only treatment that respects this structural reality.

A further mechanism worth noting: a study probed item social desirability by correlating personality items with the Balanced Inventory of Desirable Responding (BIDR). The finding is that item-level SD correlations vary widely within an inventory—some items are nearly orthogonal to SD, while others are saturated with it. This means that a candidate who scores above the cutoff on SD is not just inflating a few transparent items; they are likely responding to the entire inventory through a distorted lens. The variance in item-level SD correlations is why a scale-level cutoff is preferable to an item-level screening approach. You cannot simply flag the most transparent items and adjust those; the distortion is diffuse across the inventory.

Evidence SourceKey FindingImplication for Cutoff Veto
Crowne & MarloweHigh split-half and test-retest reliabilitySD scale is precise enough to support a fixed cutoff
Li & BaggerIM more reliable than SDECompute cutoff on Impression Management subscale only
Birkeland et al.Applicants score higher on Conscientiousness than incumbentsAverage difference is modest but concentrated in high-SD tail
Schmit & RyanConscientiousness items load with impression-management factor in applicantsStructural shift makes partialling unsafe; invalidation is required

The evidence converges on a single conclusion: above the cutoff, the SD scale carries more response-distortion variance than trait variance, and no personality score remains interpretable. The reliability of the SD scale, the superiority of the Impression Management component, the concentration of inflation in the tail, and the structural factor shift all point to the same treatment. Pre-register the cutoff, enforce it, and disqualify every candidate above it.

media social media apps social network facebook symbols digital twitter network social networking icon communication www intern

Decision Framework: Ignore, Correct, or Veto

The operational decision for any organization using self-report personality inventories in hiring has narrowed to three options: ignore social desirability (SD) variance, correct it with a regression adjustment, or veto the profile at a pre-registered cutoff. Two of those options collapse at the same boundary — the cutoff on applicant norms — leaving one defensible rule.

Ignore. Using raw personality percentiles without an SD check lets the top segment of the SD distribution score directly into the ranking. In a candidate pool, that treats a block of truncated-validity profiles as equal to valid profiles. The distortion is not random noise; it is a response style. Theoretical and empirical research on this pattern, published on Medium, finds that individuals who over-endorse conscientiousness tend toward obsessive-compulsive behavior and fewer favorable individual outcomes. The inflated SD score is not a mild positivity bias — it is a marker of a maladaptive response pattern, and raw ranking preserves every bit of that distortion.

Correct. The standard fix is a uniform regression correction that subtracts the average SD relationship from every personality dimension. That works in the middle of the distribution and fails at the top because the SD–trait relationship is nonlinear above the cutoff. In that upper stratum, SD changes from a covariate to a suppressor: instead of adding predictable bias that you can remove, it alters the inter-trait structure itself. Subtracting an average slope from scores that are not average leaves the tail distorted and gives the false appearance of psychometric hygiene.

Veto. A pre-registered binary rule at the cutoff codes any profile above the cutoff as invalid. It requires a norm-table lookup — on the Hogan Personality Inventory's applicant norm table, a row on a page. It leaves the remaining profiles untouched, with full rank-order information intact. And because the cutoff is fixed before applicant data are seen, it can be documented in the selection protocol and audited after the fact. This is the only option whose failure mode is visible to a reviewer.

CriterionIgnoreCorrectVeto
Censoring efficiencyRemoves none of the problematic tailRemoves most linear bias but leaves the tailRemoves the entire tail by definition
Legal defensibilityNo pre-set cutoff; post-hoc explanation requiredCorrection model must be justified; the tail remains unexplainedWins — cutoff set before applicant data are seen and documented in the selection protocol
VerdictFails both testsFails at the tailTransparent, auditable, and aligned with the psychometric finding — SD is a validity flag above the cutoff

The veto is the explicit winner: transparent, auditable, and aligned with the psychometric finding that above the threshold SD is not a covariate to be subtracted but a validity flag that renders the entire profile uninterpretable. The decision tree for any hiring cycle is therefore short and mechanical:

IFTHEN
You are scoring a candidate pool and have not pre-registered an SD cutoff before viewing applicant dataStop — pre-register the cutoff veto before any scoring
A candidate's SD score is below the cutoff on the test's applicant norm tableProceed — the personality score retains full rank-order information
A candidate's SD score reaches or exceeds the cutoffDo not hire — code the entire personality assessment as invalid
You plan to correct a profile above the cutoff with a regression adjustmentDo not — SD is a suppressor in that stratum; the corrected estimate remains unreliable
An auditor asks how the cutoff was applied in a past cycleShow the pre-registered protocol and the norm-table lookup
laptop iphone workspace hands coffee binders facebook social media social networking working notebook desk work office busines

What the Data Doesn't Tell You

The meta-analysis by Ones, Viswesvaran, and Reiss is the caveat that has to sit on top of any honest defense of the hard veto: social desirability scales did not moderate criterion-related validity in selection research. A candidate above the cutoff can still be a strong performer, and a candidate below the cutoff can still fail. That does not weaken the rule; it identifies what the rule actually is. The veto is a validity policy, not a performance diagnosis. It asserts nothing about whether the person will succeed on the job. It asserts that, for this candidate, the personality scores cannot be trusted as measurements.

Stöber introduces the second limitation, and it is partly a legal one. On social desirability scales, women and older adults score significantly higher, with group differences large enough to produce systematic differences at the cutoff, which means an unadjusted veto at the cutoff of combined applicant norms will screen out more women and older candidates than their base rates in the applicant pool justify. The risk must be monitored locally: the vendor's norm table, the national aggregate, and the actual applicant stream for a given role are never the same distribution.

Taylor and Brown complicate the interpretation of a high-SD profile in the opposite direction. Their research found that positive self-illusions track adaptive mental health, not pathology. So some high-social-desirability responding is not strategic distortion; it is genuine optimistic adjustment. The candidate may truly believe the flattering statements they endorse. This matters because the veto is not an accusation of lying and cannot be defended as one. The instrument cannot cleanly separate deliberate impression management from authentic self-enhancement, so the only defensible decision remains a validity decision.

The cutoff is also a local-norm statement, and this is the edge case most likely to generate false positives. A score at the cutoff of a wide applicant pool may be far less extreme in a high-executive pool, and the reverse can be true for a low-base-rate pool. Global norms flatten those differences. The safeguard is to pre-register the norm table for each job family before examining an applicant's scores, rather than importing an aggregate threshold across every funnel.

Finally, the rule has a hard boundary at the level of instrumentation. Forced-choice and ipsative formats suppress social desirability by design, because the candidate cannot inflate every scale simultaneously — a gain on one dimension forces a loss on another. On those instruments a cutoff veto is redundant, and in formats without an independent SD scale it is impossible to compute. The mandatory invalidation rule is therefore a rule about normative Likert-style self-report inventories, not a universal law of all personality assessment.

LimitationSource / figureOperational consequence
Criterion validity unaffected by SDOnes, Viswesvaran & ReissVeto is a data-integrity decision, never a performance prediction
Demographic group differencesStöber: group differences on social desirabilityMonitor adverse impact against local applicant norms
Adaptive self-enhancementTaylor & BrownHigh SD may be authentic optimism, not lying — still invalid for trait scoring
Norm-context dependenceCutoff broad pool differs from executive poolPre-register job-family-specific thresholds
Format boundaryForced-choice / ipsative designVeto applies only to normative Likert-style self-reports

None of these limits revives the status-quo compromise — the belief that high social desirability is a mild positive bias that can be arithmetically removed by subtracting the SD scale score from each personality dimension. Once a profile clears the cutoff, the SD scale is not iterating as a covariate; it is acting as a suppressor, absorbing construct variance at a rate that makes any corrected trait estimate uninterpretable. That is why the rule is not adjust, but invalidate.

internet whatsapp smartphone communication phone networking app chat mobile networked global iphone ios make a phone call comm

Worked Case

Nadia’s file is a clean, forced experiment in the difference between a score and an interpretation. As an operations-analyst finalist, she completed the NEO PI-R and a separately administered social desirability scale. According to Costa and McCrae, on the NEO PI-R T-score metric, her Conscientiousness T-score puts her at the upper end of adult norms. That is the kind of result that normally justifies an offer. It is also, under the pre-registered policy, a result that will never be read.

Her social desirability score sits above the cutoff on the test publisher’s applicant norm table. The pre-registered cutoff is the threshold; any score at or above that line invalidates the entire NEO PI-R profile before any dimension is interpreted. So the high Conscientiousness score is not “adjusted” or “flagged” — it is coded invalid. The reason is structural: by the time an applicant’s SD score reaches the cutoff, the validity scale contains more response-distortion variance than trait variance, so no personality dimension can be separated from the distortion. Above that threshold the SD scale is a suppressor, not a covariate; a statistical correction would still leave a trait estimate built on uninterpretable variance.

The cutoff itself is easy to apply. On the T-score metric used by the NEO PI-R, any validity-scale T-score at or above the cutoff triggers the veto. Nadia’s SD score is well past the trigger; the exact T-score corresponding to her score on the applicant norms does not need to be calculated for the decision to be made.

What would have happened under a looser rule is worth making explicit. Had the organization adopted the vendor’s default higher cutoff, Nadia would have stayed in the ranking and her elevated Conscientiousness score would have been treated as real. The gap between the cutoff and a higher cutoff is not a matter of calibration taste. It is the operational difference between a validity veto and a cosmetic check. At the higher cutoff, the SD scale still carries substantial response-distortion variance, but the default rule would let that variance ride inside the trait score as if it were signal.

The hiring manager documented that the interview with Nadia was positive. That documentation has exactly one role in this workflow: it confirms that the decision is being made on the pre-registered rule rather than on post hoc rationalizing. Policy forbids rescuing an invalidated profile. Because her SD score is above the cutoff, no estimate of Nadia’s true Conscientiousness can be given, and the assessment is excluded from all ranking models. The interview cannot substitute for the missing construct; it is a different measurement channel, not a re-run of the personality inventory.

Decision pointValue or ruleConsequenceWinner
Conscientiousness T-scoreHighUpper end of adult norms per Costa and McCraeNo weight; profile invalid
SD score percentileAbove cutoffInvalidates entire profileVeto

Frequently Asked Questions

What is the exact policy for handling candidates who score at or above the social desirability cutoff?

Do not hire any candidate whose validated social desirability score reaches or exceeds the cutoff on the test's applicant norm table, and code the entire personality assessment as invalid for that candidate.

Which component of social desirability should the veto be applied to, and why?

The veto should be applied to the Impression Management subscale specifically, not a composite including Self-Deceptive Enhancement, because Li and Bagger found Impression Management to be more reliable.

What did Birkeland et al.'s meta-analysis find about applicant versus incumbent Conscientiousness scores?

Job applicants score higher than incumbents on Conscientiousness, but the average difference is modest and concentrated among the same high-SD applicants caught by the cutoff veto.

How much can faking shift a candidate's Conscientiousness score according to Viswesvaran and Ones?

The faking effect is enough to move a candidate who is genuinely below average into the upper portion of the distribution.

What happens to the social desirability scale's role once a score passes the cutoff?

Above the cutoff, the SD scale acts as a suppressor variable that absorbs variance from the intended trait, making any personality score uninterpretable.

What reliability evidence supports using a fixed cutoff on the social desirability scale?

Crowne and Marlowe reported high split-half and test-retest reliability for the Social Desirability Scale, providing a stable floor that makes a fixed cutoff defensible.

Quick answers

What is the correct policy for handling social desirability in hiring according to the article?The correct policy is a pre-registered hard veto at the upper tail.
What did the study comparing chatbots to psychometric tests report?It reported reduced social desirability bias but lower predictive validity.
What is the mechanism behind the social desirability scale's design?The mechanism is the absolute-statement design, where a candidate must deny universal, minor human flaws.
What happens when a score passes the cutoff?Once a score passes the cutoff, the SD scale acts as a suppressor variable, absorbing variance from the intended trait.
What is the verdict for a no-cutoff policy?The verdict for a no-cutoff policy is unacceptable, as it catches nothing and leaves invalid variance exceeding trait variance in the ranking model.

Sources: arXiv, arXiv, Reddit, Reddit, Reddit

Also worth reading: Unpacking Conscientiousness How It Shapes Lives: Unpacking Conscientiousness How It Shapes · Facet-Level Conscientiousness: Meta-Analytic Evidence: Facet-Level Conscientiousness: Meta-Analytic Evidence · Alarming 47% Surge in Young Australian Mental Health Cases 2024 Statistical Analysis: Alarming 47% Surge in Young

Research Methodology & Editorial Standards

We begin by defining the specific objectives the reader needs to accomplish. Primary product documentation and authoritative secondary sources are assembled into a verified research corpus; drafting occurs only after this foundation is in place.

Every quantitative claim is subjected to dual-source verification. Any figure that cannot be independently corroborated is either qualified or omitted.

Published · Last reviewed · Owned by the Psychprofile editorial desk (About, Contact, Privacy).

Related answers