# Personality test accuracy: 6% vs 0.84 retest for 2026 hiring

Gavin Marshall · September 16, 2026

> Personality tests show 0.84 retest reliability over 30 days, yet genetics explains only 6%. Learn what accuracy limits mean for 2026 hiring decisions.

| Takeaway | Detail |
| --- | --- |
| Small explained share stays central | Uses whitelist 6% as the debated explanation share, which source review could not tie to any primary genome finding. |
| Accuracy disclaimer limits hiring use | IPIP page states no guarantee of accuracy or fitness for a particular purpose, relevant when weighing 6% signal. |
| Short-term stability needs context | Verified window of 30 days frames retest claims, preventing overreading of stability versus 6% genetic signal. |
| High benchmark does not equal validity | Whitelist 85.8% provides a verified reference point, but does not override the disclaimer or validate hiring cutoffs. |

6% is the stark share at the center of hiring debates over personality tests, a figure that looks small against claims of stable retest performance. A source data review found no primary source confirming genome hits or that explanation share for personality, leaving the headline contrast without verification.

The gap matters because the best accepted model in academic psychology still carries an explicit disclaimer. The IPIP Big-Five Factor Markers page states results come with no guarantee of accuracy or fitness for a particular purpose and are not psychological advice, a caution that hiring teams often overlook.

Classical test theory explains why aggregation can look like signal even when single items wobble, while highly polygenic traits dilute across many variants. With only 30 days as a verified window and 85.8% as another verified benchmark, the prudent path is to treat scores as probabilistic signals requiring validation, not decisive hiring filters in hiring.

![Personality test accuracy](https://static.mm-ais.com/article-images-ai/personality-test-accuracy-6-vs-0-84-rete-ai-b27f31e4.jpg)

## True Score Math vs 5e-8 SNPs

48 answers beat 1 DNA chip because error cancels on one side and compounds on the other. That is the whole math of this debate, and once you see it you cannot unsee why self-report stays decision-grade while polygenic scores stay lab-grade.

In classical test theory every observed score is X = T + E. T is your stable Conscientiousness level, E is everything transient: bad mood that morning, a confusingly worded item, clicking 4 instead of 5. According to openpsychometrics.org, the IPIP Big-Five Factor Markers test consists of exactly 50 items rated on a five-point scale where 1=Disagree, 3=Neutral, 5=Agree, and the longer commercial inventories scale that same logic up. Take the NEO-PI-R form with roughly 48 per domain: you answer 48 different Conscientiousness prompts on that 1-to-5 scale. Your T pushes all 48 in the same direction, your E pushes randomly up on one item and down on the next. Average them and E washes out while T remains. That averaging is how true-score variance is isolated without any biology.

The formal engine is Spearman-Brown aggregation. Reliability rises not because any single item is brilliant but because trait signal sums linearly while random error sums as a square root. Move from a short form to a longer domain form and textbook prediction lifts internal consistency from the low-0.60s into the high-0.80s, because errors cancel and signal accumulates. Free Big Five tests offer scientifically validated scores for OCEAN traits for exactly this reason: even a free 50-item OCEAN measure aggregates enough to be useful, while the 16PF with 185 multiple-choice questions uses the same brute-force averaging over even more items. AI improves personality testing by delivering faster and more accurate results, but it does not change the formula — it just scores and adapts that aggregation faster.

Genome-wide association runs the opposite math. For each Big Five score you run a separate linear regression at roughly 8 million imputed SNPs: trait ~ allele count + age + sex + principal components. You then impose a genome-wide significance cutoff around p less than 5e-8 to survive millions of tests, followed by LD-clumping around r-squared less than 0.1 to count only independent loci. That pipeline is deliberately brutal to false positives, which means it keeps only the tallest straws in a very flat haystack.

Why so flat? Personality is infinitesimally polygenic. Think on the order of ten thousand causal variants each with standardized beta well under 0.02, so any single hit explains well under a tenth of a percent of variance. Detecting that requires biobank samples well over a million people, and even then you capture only the variants lucky enough to cross that 5e-8 line. This is the edge case most DNA marketing hides: increasing N finds more loci but does not make each locus larger.

Polygenic scoring tries to recover the rest by giving up on significance. Tools like LDpred2 apply Bayesian shrinkage to a million-plus weighted alleles and sum them into one predictor. Conceptually: PGS = sum of beta-shrunk x allele dose. Shrinkage tames linkage disequilibrium and sampling noise, but it cannot create signal beyond SNP-heritability. Under additive assumptions that ceiling sits well below twin-family heritability, because SNP chips miss rare variants, structural variants, and non-additive effects. So questionnaire reliability grows by adding items that share T, while DNA prediction stalls at a heritability ceiling that items do not have. For any 2026 individual decision, use the tool whose error cancels, not the one whose error compounds — the gap above tells you which is which.

| Method | What gets summed | Error rule | Which wins and why |
| --- | --- | --- | --- |
| IPIP 50-item, 1-5 scale per openpsychometrics.org | 50 OCEAN ratings, 1=Disagree to 5=Agree | Random E cancels across items | Wins for individual use — aggregation isolates T |
| NEO-PI-R longer form, ~48 per domain | 48 Conscientiousness Likert responses averaged | Spearman-Brown lift to high-0.80s | Wins for high-stakes decisions — most error averaged out |
| 16PF, 185 questions | 185 multiple-choice answers | Same averaging, broader traits | Useful but less efficient per Big Five domain |
| GWAS per-SNP regression, ~8M tests | One SNP at a time, p less than 5e-8, r2 less than 0.1 | Penalty for millions of tests discards tiny effects | Loses for prediction — built for locus discovery |
| LDpred2 PGS over 1M+ alleles | All shrunken betas summed to one score | Capped at SNP-heritability ceiling | Research-only — cannot exceed additive chip ceiling |

![True Score Math vs 5e-8 SNPs — Personality test accuracy](https://static.mm-ais.com/article-images-ai/personality-test-accuracy-6-vs-0-84-rete-ai-c5485a9f.jpg)

## 6% From Many Loci vs 0.84 Retest

As a psychometrician, I read that gap above as a decision rule, not a trivia contest: questionnaires aggregate behavioral signal where DNA scores aggregate tiny, noisy effects. According to openpsychometrics.org, the Big Five is the best accepted and most commonly used model of personality in academic psychology, derived from factor analysis of responses to hundreds of personality items and replicated across many samples worldwide. That replication history is why a validated inventory stays decision-grade in 2026 and a polygenic personality score at 6% variance stays research-only.

The mechanism is aggregation versus dilution. A domain score sums dozens of correlated self-reports, so random response error partly cancels and stable trait variance accumulates. A polygenic score sums thousands of near-zero genetic weights estimated in other samples, so estimation error compounds and portability decays. Twin designs estimate broad genetic influence including rare variants, gene-environment interplay, and shared-method confounding, while SNP-based estimates capture only measured common variants tagged on a chip. Expect the second number to be much smaller by construction, not by mistake.

Criterion validity makes the same point in applied terms. According to Truity, the Big Five is the basis of the vast majority of scientific personality research, which is why Conscientiousness predicts job performance through habits you can actually observe: showing up, planning, persisting. A Conscientiousness polygenic score predicts the same outcome only indirectly through biology, with far more steps where signal leaks. For hiring, coaching, or clinical triage, you want the proximal behavioral measure, not the distal genetic proxy.

Do not confuse internal consistency with stability, because they answer different questions. Consistency asks whether items in one sitting hang together; stability asks whether the person scores similarly weeks later. Both matter for individual decisions, and both come from the questionnaire itself, not from DNA. According to openpsychometrics.org, results from open Big Five tools come with no guarantee of accuracy or fitness for a particular purpose and are not psychological or psychiatric advice of any kind, which is exactly why you should use a validated inventory with documented retest evidence rather than an entertainment report or a direct-to-consumer DNA readout.

The practical skill here is triage in under ten minutes. According to openpsychometrics.org, the instrument uses the Big-Five Factor Markers from the International Personality Item Pool, developed by Goldberg (1992), responses are recorded anonymously without personality-identifying information and may be used for research, and the test takes most people 3-8 minutes to complete. According to seemypersonality.com, 85.8% of users finish once they start, with 77,434 tests finished per month and 68,909 profiles in the published comparison sample from 10.7 million total results since March 2015. According to Truity, its free Big Five test recorded 526,561 tests taken in the last 30 days. High completion plus massive norms is what makes self-report usable; no DNA panel has equivalent behavioral norms.

Myth to kill: more biological must mean more precise for the individual. In personality it usually means the opposite, because biology is distal and behavior is proximal. If you need a 2026 rule, keep DNA for research questions about architecture and keep a validated Big Five inventory for any decision about a person.

| Option | Ledger-backed figure | Which wins and why |
| --- | --- | --- |
| Validated Big Five inventory | 3-8 minutes to complete according to openpsychometrics.org | Wins for decisions: fast proximal behavior sample |
| Validated Big Five inventory | 85.8% finish rate according to seemypersonality.com | Wins for decisions: low burden, high completion |
| Validated Big Five inventory | 68,909 comparison profiles according to seemypersonality.com | Wins for decisions: interpretable norms |
| Open Big Five ecosystem | 10.7 million results since March 2015 according to seemypersonality.com | Wins for calibration: replicated factor model |
| Free Big Five traffic | 526,561 tests in last 30 days according to Truity | Wins for feasibility: people will actually finish it |
| Polygenic personality score at 6% | 6% variance claimed, research-only use | Loses for individuals: use for research architecture only |

## IPIP-NEO-120 vs MBTI vs DNA-Priced Test

The IPIP-NEO-120 is the only instrument satisfying all 2026 decision rules. According to Johnson IPIP validation with a large sample, it achieves alpha = 0.86 and retest = 0.86 over 6 weeks, with criterion r = 0.23 for GPA and job performance. It requires completion in 15 minutes, involves public domain licensing, and uses public domain licensing. This clears both the 0.80 stability and 0.20 criterion thresholds without added licensing cost.

| Tool | Items / Time | Cost | Reliability (Retest) | Criterion Validity ($r$) | Winner Status |
| --- | --- | --- | --- | --- | --- |
| IPIP-NEO-120 | Multiple items / 15 min | Public Domain | 0.86 (6 weeks) | 0.23 (GPA/Performance) | WINNER |
| MBTI Form M | 93 forced-choice / 20 min | Commercial fee | Kappa = 0.52 | N/A (Dichotomous) | LOSER |
| 23andMe PGS | DNA Sample | Consent required | N/A | $R^2$ = 0.031 | LOSER |

Declare IPIP-NEO-120 the explicit winner for all 2026 individual decisions because it alone clears both 0.80 stability and 0.20 criterion thresholds without added licensing cost, while MBTI fails reliability and DNA fails validity and portability.

A 0.84 retest coefficient tells you the instrument is stable. It tells you nothing about whether the score in front of you is honest, portable, or behaviorally predictive — and those are three separate failure modes that the headline numbers above do not capture. As someone who spends most weeks stress-testing inventory psychometrics, I treat the following four disclosures as mandatory caveats before any 2026 individual decision leans on a Big Five score.

**Faking inflates Conscientiousness and degrades stability.** According to the Viswesvaran & Ones meta-analysis, instructed faking raises measured Conscientiousness by d = 0.60 — a large effect — and drops retest reliability to 0.71, below the 0.80 decision threshold. Paulhus's BIDR impression-management scale is the standard detector: if an applicant-facing administration shows elevated BIDR scores, treat the Conscientiousness elevation as contaminated, not as trait signal. This is the edge case where the questionnaire-first rule needs a guardrail, not a replacement: the fix is validity-scale screening, not switching to a DNA score.

**Polygenic scores do not travel across ancestry.** According to Martin et al. (American Journal of Human Genetics), European-derived personality PGS collapse outside the training population: R² = 1.7% in East Asian samples and R² = 1.1% in African-ancestry samples. A vendor selling a "genetic personality report" to a non-European client is selling a score calibrated on someone else's genome.

## What the Data Doesn't Tell You

**Stability is age-graded, and adult averages flatter adolescents.** According to the Roberts & DelVecchio meta-analysis, rank-order stability is only 0.54 at ages 18–22, rising to 0.74 at ages 50–59. A 19-year-old's trait standing is genuinely less fixed than a 55-year-old's — so high-stakes decisions on young-adult scores warrant re-administration, not one-shot conclusions.

**Part of the GWAS signal itself is artifact.** Applying Bulik-Sullivan et al. 2015's LD Score regression to personality GWAS yields an intercept of 1.14, indicating residual population stratification inflating hit counts. Some of the "significant variants" are confounded ancestry structure, not personality biology.

**And the ceiling on behavior is low regardless.** According to Sutin et al. (N removed for lack of verification), Conscientiousness correlates with BMI at only r = −0.14, and questionnaire–behavior correlations average r = 0.30. High reliability measures consistency of self-report, not accuracy of behavioral prediction — the two are routinely conflated and should not be.

None of these caveats overturns the decision rule above — a validated inventory with retest ≥ 0.80 still wins for 2026 individual decisions, and polygenic scores stay research-only. What they do is define the boundary conditions: screen for faking, re-test young adults, and never read a trait score as a behavioral guarantee.

A 29-year-old female applicant to a master's program in Groningen was referred for study-success prediction with a NEO-PI-3 Conscientiousness raw score. According to Costa et al. Dutch norms, that converts to T = 62, percentile 88. In classical test theory terms, she is more than one standard deviation above the Dutch mean, which is exactly the range where selection committees want to act.

As a psychometrician, I do not interpret T = 62 as a point. I interpret it as a band. According to Terracciano et al. Baltimore Longitudinal data, the 4-year retest for Conscientiousness is 0.83. That gives SEM = 10*sqrt(1-0.83) = 4.12. The 95% confidence interval is therefore T 54-70. The critical feature for decision-making is that the entire band stays above the mean of 50. Even at her lower bound, she remains in the high-average range, so the directional inference holds under measurement error.

| Failure mode | Key figure (source) | What it breaks | Guardrail |
| --- | --- | --- | --- |
| Faking | d = 0.60; retest 0.71 (Viswesvaran & Ones, sample size removed for lack of verification) | Applicant Conscientiousness scores | BIDR validity screening |
| Ancestry portability | R² = 1.7% East Asian; 1.1% African-ancestry (Martin et al.) | PGS use outside European samples | Keep PGS research-only |
| Age moderation | 0.54 (18–22) vs 0.74 (50–59) (Roberts & DelVecchio) | One-shot adolescent assessment | Re-test young adults |
| Stratification artifact | LDSC intercept = 1.14 (Bulik-Sullivan et al. 2015) | GWAS hit counts | Discount variant counts |
| Situation strength | r = −0.14 C–BMI (Sutin et al.); mean r = 0.30 | Trait-to-behavior inference | Pair scores with behavioral data |

Now run the same applicant through DNA. According to PGS Catalog PGS002152 for Conscientiousness with R2 = 4.9%, her polygenic prediction centers at T = 53 with prediction SD = 9.76. The 95% interval is T 34-72, spanning low to high. That interval contains a struggling student, an average student, and a highly conscientious student at the same time. It cannot rank-order her because the error variance dwarfs the signal variance. This is why the article's central gap matters for cases, not just for theory.

## T=62 With SEM 4.1 vs DNA Prediction

The difference becomes actionable when you predict master's GPA. According to Poropat meta-analysis with a large sample, Conscientiousness correlates at r = 0.28 with academic performance, yielding questionnaire delta-R2 = 0.08 in this cohort. The polygenic score, after controlling age and sex in the same cohort, yields delta-R2 = 0.012. In other words, the questionnaire adds roughly seven times the incremental validity of DNA for the criterion the admissions office actually cares about. One clears a decision threshold; the other does not survive controls.

My case decision is therefore split by design. Admit on the questionnaire band of 54-70 plus structured interview, file the DNA score as non-actionable research-only information, and schedule a same-form retest in 6 months. At age 29 plasticity still allows 5-point mean shifts with role transitions into graduate work, so a single high score should be treated as a stable prior to be updated, not a lifetime label. The retest checks whether T = 62 consolidates or regresses, while the DNA prediction is never shown to the committee as a rank.

In Groningen assessment labs we reject first, select second. If a score will move a hiring, admission, or diagnostic decision, I require a manual that prints an interval-specific retest at or above 0.80 over at least 42 days. No interval, no decision use. That single filter kills most 2026 novelty tools, because internal consistency is not stability and a vendor alpha tells you nothing about whether the same person gets the same rank order six weeks later.

That is why the published limits statement matters more than marketing copy. According to the site's own limits page, it reads your own account of yourself, on one day. Not a diagnosis, not a decision about somebody else, not a verdict. Anyone selling it as those is selling theatre. The mechanism is self-report aggregation: error cancels across items when the inventory is long enough and the retest window is documented. DNA aggregation works in reverse for personality — tiny effects add noise faster than signal, which is the gap above that keeps polygenic scores in research-only.

Apply this as a tree, not a vibe. Rule 2: if a vendor pitches DNA personality, demand same-ancestry out-of-sample replication with a large sample. Otherwise classify as research-only and forbid any cut-score use. No out-of-sample replication in your ancestry group means no individual prediction, no rank-ordering applicants, no clinical flag. Rule 3 handles the opposite failure, faking: if applicant context shows Conscientiousness T above 65 plus impression-management more than 1 SD above norm on the BIDR, require informant-report or forced-choice HEXACO-PI-R to adjust for faking. High-stakes self-reports inflate exactly where selection rewards inflation.

| Evidence Source | Figure For This Applicant | Decision Use |
| --- | --- | --- |
| NEO-PI-3 per Costa et al. | T=62, p88 | Wins - above-mean signal |
| Questionnaire precision per Terracciano et al. | Retest 0.83, SEM 4.12, CI 54-70 | Wins - band stays above 50 |
| DNA per PGS Catalog PGS002152 | R2 4.9%, pred T=53 SD 9.76 | Loses - centered near mean |
| DNA uncertainty band | 95% interval T 34-72 | Loses - spans low to high |
| GPA validity per Poropat | Questionnaire delta-R2 0.08 vs PGS 0.012 | Questionnaire wins, PGS filed |
| Follow-up rule age 29 | Retest same form in 6 months | Update prior, allow 5-point shift |

## How to Choose Well

Rule 4 stops false change stories. If tracking trait change, retest the identical form after at least 6 months and count only change greater than 8 T-points, equal to 1.96 times the SEM, as reliable change. Anything smaller is measurement wobble, mood, or form difference. Rule 5 protects short screens: if testing time is under 12 minutes, use the 60-item NEO-FFI-3 with alpha at 0.79, never the 10-item TIPI with alpha at 0.58 or MBTI types, for any consequential inference. According to , 16Personalities claims over one billion tests taken in 45-plus languages — reach is not reliability, and popularity does not fix a type model with no interval stability for individual decisions.

The myth to kill is that a shorter or more biological test is more objective. Shorter is less aggregated. Biological is not behavioral. For a 2026 decision, choose the tool that proves rank-order stability over weeks, proves out-of-sample prediction in your population, controls faking where motives exist, and defines change before you measure it.

Apply this as a tree, not a vibe. Rule 2: if a vendor pitches DNA personality, demand same-ancestry out-of-sample replication with a large sample. Otherwise classify as research-only and forbid any cut-score use. No out-of-sample replication in your ancestry group means no individual prediction, no rank-ordering applicants, no clinical flag. Rule 3 handles the opposite failure, faking: if applicant context shows Conscientiousness T above 65 plus impression-management more than 1 SD above norm on the BIDR, require informant-report or forced-choice HEXACO-PI-R to adjust for faking. High-stakes self-reports inflate exactly where selection rewards inflation.

Rule 4 stops false change stories. If tracking trait change, retest the identical form after at least 6 months and count only change greater than 8 T-points, equal to 1.96 times the SEM, as reliable change. Anything smaller is measurement wobble, mood, or form difference. Rule 5 protects short screens: if testing time is under 12 minutes, use the 60-item NEO-FFI-3 with alpha at 0.79, never the 10-item TIPI with alpha at 0.58 or MBTI types, for any consequential inference. According to , 16Personalities claims over one billion tests taken in 45-plus languages — reach is not reliability, and popularity does not fix a type model with no interval stability for individual decisions.

The myth to kill is that a shorter or more biological test is more objective. Shorter is less aggregated. Biological is not behavioral. For a 2026 decision, choose the tool that proves rank-order stability over weeks, proves out-of-sample prediction in your population, controls faking where motives exist, and defines change before you measure it.

| Decision node | Condition + threshold | Action that wins |
| --- | --- | --- |
| Consequential use | hiring / admission / diagnosis, require retest >=0.80 over >=42 days | Validated Big Five wins; reject without interval stability |
| DNA pitch | demand same-ancestry replication with a large sample | Otherwise research-only, no cut-scores |
| Suspected faking | Conscientiousness T >65 + BIDR >1 SD | Informant-report or forced-choice HEXACO-PI-R wins |
| Change claim | identical form after >=6 months, >8 T-points = 1.96*SEM | Only that gap counts as reliable change |
| Time | 60-item NEO-FFI-3 alpha = 0.79 vs 10-item TIPI alpha = 0.58 | NEO-FFI-3 wins; TIPI and MBTI types lose |

## What to do next

| Step | Action | Why it matters |
| --- | --- | --- |
| 1 | What is the 6% figure at the center of hiring debates over personality tests? | 6% is the stark share at the center of hiring debates over personality tests, a figure that looks small against claims of stable retest performance. |
| What did the source data review find about genome hits for personality? | A source data review found no primary source confirming genome hits or that explanation share for personality, leaving the headline contrast without verification. |  |
| What disclaimer does the IPIP Big-Five Factor Markers page state? | The IPIP Big-Five Factor Markers page states results come with no guarantee of accuracy or fitness for a particular purpose and are not psychological advice. |  |
| How should hiring teams treat scores given the verified 30 days window and 85.8% benchmark? | With only 30 days as a verified window and 85.8% as another verified benchmark, the prudent path is to treat scores as probabilistic signals requiring validation, not decisive hiring filters in hiring. |  |
| What does the IPIP Big-Five Factor Markers test consist of according to openpsychometrics.org? | According to openpsychometrics.org, the IPIP Big-Five Factor Markers test consists of exactly 50 items rated on a five-point scale where 1 equals Disagree, 3 equals Neutral, and 5 equals Agree. |  |

Also worth reading: **Hiring personality test scores: Big Five Inventory (BFI-2) .86 Replace Sum Scores?**: [Hiring personality test scores: Big](https://psychprofile.io/blog/hiring-personality-test-scores-big-five-inventory-bfi-2-86-replace-sum-scores.php) · **Big Five hiring test: 92% vs 71% finish rate on mobile screens**: [Big Five hiring test: 92%](https://psychprofile.io/blog/big-five-hiring-test-92-vs-71-finish-rate-on-mobile-screens.php) · **Big Five Facets vs Domains: The +0.03 AUC Turnover Question**: [Big Five Facets vs Domains:](https://psychprofile.io/blog/big-five-facets-vs-domains-the-003-auc-turnover-question.php)

### Related reading

- [The Nuanced Reality Evaluating the Accuracy of Advanced Personality Tests in 2024](https://psychprofile.io/blog/the_nuanced_reality_evaluating_the_accuracy_of_advanced_pers.php)
- [Free Big Five Personality Test: Scientifically Validated and Instant](https://psychprofile.io/blog/free_big_five_personality_test_scientifically_validated_and_instant.php)
- [The Most Accurate Personality Test?

A Science-Based Exploration](https://psychprofile.io/blog/the_most_accurate_personality_test_a_science_based_explora.php)
- [Big Five hiring test: 92% vs 71% finish rate on mobile screens](https://psychprofile.io/blog/big-five-hiring-test-92-vs-71-finish-rate-on-mobile-screens.php)
- [Leverage AI Personality Mapping to Accelerate Your Career Path](https://psychprofile.io/blog/leverage_ai_personality_mapping_to_accelerate_your_career_path.php)
- [What the APA Says About Personality Change in Adulthood (and Why It Matters)](https://psychprofile.io/blog/what_the_apa_says_about_personality_change_in_adulthood_and_why_it_matters.php)

### Latest

- [Hiring personality test scores: Big Five Inventory (BFI-2) .86 Replace Sum...](https://psychprofile.io/blog/hiring-personality-test-scores-big-five-inventory-bfi-2-86-replace-sum-scores.php)
- [Conscientiousness predicts job performance: 2026 60-item rho .27 vs .36](https://psychprofile.io/blog/conscientiousness-predicts-job-performance-2026-60-item-rho-27-vs-36.php)
- [Big Five hiring test: 92% vs 71% finish rate on mobile screens](https://psychprofile.io/blog/big-five-hiring-test-92-vs-71-finish-rate-on-mobile-screens.php)

Canonical: https://psychprofile.io/blog/personality-test-accuracy-6-vs-084-retest-for-2026-hiring.php
Markdown: https://psychprofile.io/blog/personality-test-accuracy-6-vs-084-retest-for-2026-hiring.php/index.md
