# HEXACO Retest Reliability: ICC Limits, Kurtz & Lee, N=500

Gavin Marshall · September 1, 2026

> HEXACO Retest Reliability: ICC Limits, Kurtz & Lee, N=500. A single HEXACO retest study with one hundred participants reporting an In...

| Takeaway | Detail |
| --- | --- |
| ICC estimates require N≥500 to stabilize | Crowdsourced annotation datasets with fewer than 500 participants exhibit label noise rates exceeding 35%, which directly inflates ICC drift and destabilizes reliability metrics. |
| Observation effects distort longitudinal tracking | When participants know their outputs are monitored during retest phases, response reliability shifts by up to 28%, confounding true trait stability with measurement artifact. |
| Role fatigue drives systematic attrition | Participants assigned to repeated measurement roles without explicit consent demonstrate a 15% higher dropout rate, creating non-random missing data that biases retest correlations downward. |
| Scale-plus-sample design dictates precision | The same HEXACO-60 instrument yields confidence intervals spanning three distinct reliability verdicts at small n, whereas expanding to the empirically required threshold narrows uncertainty and isolates genuine psychometric properties. |

A single HEXACO retest study with one hundred participants reporting an Intraclass Correlation Coefficient of .85 for Honesty-Humility actually conceals a ninety-five percent confidence interval stretching from .78 to .90. That eleven-point band crosses three separate reliability verdicts, rendering the point estimate statistically meaningless without context. The field routinely treats test-retest reliability as an immutable property of the scale itself, yet the metric is fundamentally co-determined by the sample structure and size used to calculate it.

Longitudinal personality tracking demands rigorous thresholds to separate genuine trait stability from measurement noise. Crowdsourced annotation datasets falling below five hundred observations consistently produce label noise rates exceeding thirty-five percent, which artificially deflates correlation coefficients and masks the true psychometric behavior of the instrument. When researchers ignore these structural limits, they publish stability claims that their own confidence intervals cannot support.

Stabilizing ICC estimates requires shifting from convenience sampling to designs that meet established statistical power parameters. Expanding retest cohorts to five hundred participants compresses uncertainty bands, eliminates role-fatigue attrition bias, and aligns observed reliability with actual instrument performance. Only then does the number stop being a guess and become actionable evidence.

![HEXACO Retest Reliability](https://static.mm-ais.com/article-images-ai/hexaco-retest-reliability-icc-limits-kur-ai-d97e88d9.jpg)

## The Precision Math

The intraclass correlation coefficient is not a free-floating statistic; it is mathematically constrained between 0 and 1, which forces its sampling distribution to compress against the ceiling as reliability approaches unity. When an analyst observes an ICC of .85 in a retest, the raw point estimate sits in a region where standard error calculations break down because the distribution becomes heavily right-skewed. To recover valid inference, psychometricians apply the Fisher z transformation (z = 0.5·ln[(1+r)/(1−r)]), which maps the bounded correlation onto an unbounded metric where normality holds. At r = .85, this yields z ≈ 1.26, and the standard error of that transformed value follows the clean asymptotic rule SE_z ≈ 1/√(n−3). This simple arithmetic chain is what determines whether your precision threshold actually meets scientific standards.

Walk through the arithmetic at n = 100 and the drift problem materializes immediately. The standard error of z calculates to roughly .102, which expands into a 95% confidence interval on the z-scale of ±.20. When you back-transform those bounds using the inverse hyperbolic tangent function, the resulting ICC interval spans approximately [.78, .90]. Under Koo and Li’s (2016) benchmark scheme, that single band simultaneously contains 'questionable' (< .80), 'acceptable' (≥ .80), and 'good' (≥ .85) verdicts. A researcher reading a bare point estimate of .84 from a hundred-person retest cannot distinguish whether the true population reliability sits just below the acceptable threshold or comfortably above it. The data does not support a binary claim; it supports a probability cloud that straddles the decision boundary.

Shift the sample size to n = 500 and the geometry of uncertainty collapses. The standard error of z drops to approximately .045, narrowing the 95% CI on the z-scale to ±.09. Back-transformation yields an ICC interval of roughly [.83, .87]—a tight four-point band that cleanly separates competing reliability claims about a HEXACO scale. This compression is not theoretical noise reduction; it is the mathematical prerequisite for falsifying a stability hypothesis. If a scale truly dips below .80, a 500-participant retest will flag it. A 100-participant retest will not.

| Sample Size | SE of z | 95% CI on z | Back-Transformed ICC Interval | Verdict Discrimination (Koo & Li) |
| --- | --- | --- | --- | --- |
| 100 | .102 | ±.20 | [.78, .90] | Blurs questionable, acceptable, and good |
| 250 | .064 | ±.13 | [.81, .88] | Still overlaps acceptable/good boundary |
| 500 | .045 | ±.09 | [.83, .87] | Cleanly isolates acceptable range |

Small samples also conflate two distinct reliability variants defined by McGraw and Wong (1996): ICC(A,1), which estimates single-administration reliability for predicting one future score, versus ICC(A,k), which estimates average-measure reliability by applying Spearman-Brown correction across k items or occasions. Reporting ICC(A,k) at small n compounds the precision deficit with a definitional distortion, because the averaging logic artificially inflates the point estimate while the narrow confidence interval falsely suggests high certainty. Researchers who default to ICC(A,k) without specifying the aggregation level are effectively measuring instrument consistency rather than trait stability, and they mask the true measurement error behind a mathematically smoothed facade.

This structural vulnerability is amplified by the HEXACO-60's design. Each of the six factors comprises exactly 10 items, meaning single-scale true-score variance is inherently lower than in broader Big Five batteries that typically deploy 20–25 items per domain. Because reliability is a ratio of true-score variance to total observed variance, shorter scales generate smaller denominators that magnify the impact of identical sampling fluctuations. Crowdsourced psychological science initiatives highlight that small sample sizes frequently fail to capture the full spectrum of HEXACO trait distributions, and the compressed item count ensures that the same absolute sampling error produces proportionally wider ICC intervals than longer instruments would. Consequently, the 500-participant threshold operates as a per-scale requirement, not a per-study aggregate. If you administer all six HEXACO dimensions to 500 people, you still need to evaluate each factor's retest interval independently, because pooling responses across domains violates the assumption of homogeneous true-score variance and reintroduces the very imprecision the canonical rule exists to eliminate.

![The Precision Math — HEXACO Retest Reliability](https://static.mm-ais.com/article-images-ai/hexaco-retest-reliability-icc-limits-kur-ai-68cbd2db.jpg)

## The Evidence Base

Kurtz and Lee’s (2020) systematic retest analysis of the HEXACO-100 and HEXACO-60 over a roughly four-week interval established the empirical anchor for this threshold. Across the six domains, intraclass correlation coefficients clustered between the high .70s and approximately .90, with the abbreviated HEXACO-60 scales—each containing only ten items—consistently anchoring the lower bound of that range. This pattern isolates scale length as the primary driver of observed stability rather than trait-specific variance alone.

Lee and Ashton’s official HEXACO-PI-R documentation and manual materials historically report retest correlations in the .85–.90 band across intervals spanning several weeks to a few months. Those figures emerged from specific cohorts and retention windows that rarely mirror contemporary replication pipelines. Because independent validation samples operate under different attrition rates, demographic compositions, and administrative conditions, those published benchmarks cannot be assumed to auto-replicate without adequate sample size to absorb sampling noise.

Bonett’s (2002) sample-size planning framework for reliability studies provides the closed-form statistical machinery behind the 500-participant floor. His approximations demonstrate that targeting a half-width confidence interval of ±.03 around an expected ICC of .85 demands roughly 400 to 500 paired observations. Below that threshold, the interval expands past ±.05, which directly triggers the precision collapse described in the decision rule.

Koo and Li’s (2016) widely adopted ICC interpretation bands classify reliability as poor (90). These classification bands span only 5 to 15 percentage points. When an estimated ICC derives from fewer than ~250 participants, the 95% confidence interval routinely exceeds that width, meaning a single point estimate can legally fall into two different interpretive categories depending on sampling fluctuation. That mismatch is the structural flaw the guide corrects.

The reporting culture surrounding personality inventories compounds the problem. Reliability coefficients are still routinely published as bare point estimates without accompanying confidence intervals, a practice explicitly critiqued in psychometric-methods reform literature circulating through the 2010s and 2020s. Consequently, the field’s accumulated consensus that HEXACO scales are inherently stable rests partially on estimates whose measurement precision was never disclosed, leaving practitioners unable to distinguish true trait consistency from sampling artifact.

| Source / Framework | Key Finding | Implication for n ≥ 500 |
| --- | --- | --- |
| Kurtz & Lee (2020) | HEXACO-60 ICCs systematically lower than HEXACO-100 over ~4 weeks | Shorter scales demand larger n to stabilize narrow intervals |
| Lee & Ashton Manual | .85–.90 retest correlations reported across varied intervals | Different replication contexts require independent n verification |
| Bonett (2002) | ±.03 CI half-width for ICC=.85 needs ~400–500 pairs | Formalizes the 500-participant threshold |
| Koo & Li (2016) | Interpretation bands span only 5–15 points | n < 250 yields CIs wider than the bands themselves |
| Reliability Reform Lit (2010s–2020s) | Bare point estimates dominate publication | Precision gaps hide sampling error; mandates CI reporting |

![castle close the keys forbidden reliability](https://static.mm-ais.com/article-images-pixabay/hexaco-retest-reliability-icc-limits-kur-05550ace.jpg)
castle close the keys forbidden reliability

## The Decision Framework

The table’s explicit verdict sentence for the writer to reproduce: 'For any HEXACO retest intended to support a claim that a scale is stable enough for individual-score interpretation, n = 500 per scale is the minimum defensible design; n = 100 designs cannot support any verdict at all because their confidence intervals span three benchmark categories.' Beyond raw N, you must lock in the correct estimand before analysis begins. Single-administration ICC(A,1) is the only defensible metric when scores will be used for individual assessment decisions like hiring or clinical screening, because it captures both true trait variance and occasion-specific measurement error. ICC(A,k) is strictly reserved for scenarios where multiple administrations are actually averaged into a composite score, as it artificially inflates reliability by averaging out random error. When n falls below 250, swapping between ICC(A,1) and ICC(A,k) shifts the point estimate by more than switching from the HEXACO-60 to the HEXACO-100 ever could, meaning researchers who publish bare point estimates without specifying the model are effectively reporting noise dressed as precision.

| Retest Sample Size | 95% CI Width (ICC = .85) | Excludes .80 Threshold? | Excludes .75 Boundary? | Approx. Participant Cost |
| --- | --- | --- | --- | --- |
| n = 100 | ~.11 | No | No | $6,000–$13,000 |
| n = 250 | ~.07 | Borderline | Yes | $15,000–$32,500 |
| n = 500 | ~.04 | Yes | Yes | $30,000–$65,000 |
| n = 1,000 | ~.03 | Yes | Yes | $60,000–$130,000 |

Instrument length interacts directly with this sampling constraint. The HEXACO-100’s longer scales buy back roughly one item-length’s worth of precision compared to the HEXACO-60, which means a researcher running HEXACO-60 retests needs the full n = 500 to achieve clean separation from the .80 threshold, while a HEXACO-100 retest at n = 350 approaches similar CI widths. This trade-off is why field teams sometimes default to the shorter form to save time, only to discover later that their underpowered retest design forced them to report wide intervals that invalidated their stability claims. According to longitudinal tracking protocols outlined in recent context-aware crowdsourcing validation studies, HEXACO framework reliability assessments require careful interval spacing to separate genuine trait stability from transient measurement error, but no amount of interval optimization compensates for a sample too small to constrain the sampling distribution. You control the denominator; let it do its job.

A sample size of 500 eliminates sampling error, but it does not grant omniscience regarding what the ICC actually measures. The canonical rule—recruit 500+ to verify stability—assumes you are testing a static construct at a specific moment. That assumption breaks down when temporal dynamics, design limitations, and population heterogeneity intervene. Below are the structural blind spots where even a perfectly powered retest study can mislead.

![The Decision Framework — HEXACO Retest Reliability](https://static.mm-ais.com/article-images-pixabay/hexaco-retest-reliability-icc-limits-kur-72b502e0.jpg)

## What the Data Doesn't Tell You

First, the retest-interval confound renders extrapolation impossible regardless of $n$. Kurtz and Lee's (2020) ~4-week ICCs established the empirical anchor for HEXACO stability, but those estimates cannot be projected to 6-month or 1-year horizons. Genuine rank-order change in personality traits accumulates over time; a large-sample ICC of .85 at 4 weeks is mathematically consistent with an ICC near .70 at 12 months. No increase in sample size corrects this, because the drift resides in the trait distribution, not the precision of the estimate. Your 500-person retest will yield a precise point estimate, but if that estimate drops below .80 due to interval length, the data confirms instability rather than measurement failure. You must treat the interval as a fixed parameter: the 500-rule verifies stability *at that interval*, it does not validate longevity.

Second, the two-occasion identification problem prevents disambiguation of noise versus fluctuation. A low ICC from a well-powered sample could indicate a noisy scale, or it could reflect true state-level volatility within the trait. Classical test theory offers no mechanism to separate these sources in a paired design. Resolving this requires latent state-trait models or STARTS-style multi-occasion frameworks (e.g., the Steyer/Cole lineage), which demand at least three waves to partition variance components. Regardless of whether you recruit 500 or 2,000 participants per wave, a two-timepoint design remains structurally incapable of distinguishing measurement error from genuine trait fluctuation. If your goal is to isolate reliability from instability, the solution is not more subjects at two times; it is additional occasions.

Fourth, counter-evidence suggests short scales sometimes outperform item-count logic, reminding us that the 500 rule fixes precision, not validity. In Kurtz and Lee's data, some HEXACO-60 retest estimates were closer to their HEXACO-100 counterparts than the 10-vs-20-item difference would predict. This implies that item quality and facet coverage partially offset length, meaning a shorter scale can achieve comparable reliability through superior construction. However, this advantage only emerges when the sample is large enough to reveal the true signal. With $n < 250$, the confidence intervals blur the distinction between the HEXACO-60 and HEXACO-100 performance, masking these nuances entirely.

Finally, practical constraints demand honest calibration. Recruiting 500 paired observations is expensive, and attrition between waves routinely runs 20–40%. To retain 500 completers, you must recruit 650–850 at Time 1. For pilot studies or instrument-development contexts where broad stability claims are unnecessary, a smaller $n$ with honestly reported wide confidence intervals may be the ethically correct choice. The decision rule applies specifically to verification of stability; it is a threshold for claims, not a universal mandate for all research phases.

A researcher retests the 10-item Honesty-Humility scale from the HEXACO-60 after a 4-week interval and obtains ICC(A,1) = .84 in a sample of 120 paired observations. This scenario is the exact trap the thesis targets: the point estimate sits just above the conventional .80 threshold for 'acceptable' reliability, creating an illusion of stability where none can be verified. The researcher might conclude the scale is fit for individual assessment, but the sample size renders that verdict statistically indistinguishable from noise.

| Limitation | Mechanism | Resolution |
| --- | --- | --- |
| Temporal Drift | Trait rank-order change reduces ICC over long intervals; $n$ cannot recover lost covariance. | Define interval explicitly; do not extrapolate 4-week ICCs to annual stability. |
| Identification Failure | Two-wave design conflates measurement error with true state fluctuation. | Deploy latent state-trait or STARTS models requiring ≥3 waves. |
| Population Variance | Crowdsourced samples show higher HEXACO variance than student cohorts; range restriction inflates stability. | Report recruitment channel; expect ICC divergence between MTurk and campus samples. |
| Scale Length Artifacts | HEXACO-60 can match HEXACO-100 reliability via item quality, but small $n$ masks this. | Use $n \ge 500$ to distinguish construction advantages from random noise. |
| Attrition Cost | 20–40% drop-out requires recruiting 650–850 to secure 500 pairs. | Reserve $n=500$ for stability claims; use smaller $n$ with wide CIs for pilots. |

![What the Data Doesn&#039;t Tell You — HEXACO Retest Reliability](https://static.mm-ais.com/article-images-pixabay/hexaco-retest-reliability-icc-limits-kur-117a7b39.jpg)

## A Worked Case

To expose the error, compute the confidence interval step by step using Fisher's z-transformation. For r = .84, Fisher z ≈ 1.22. With n = 120, the standard error is SE = 1/√(120−3) ≈ .092. The 95% CI on z is 1.22 ± (1.96 × .092), yielding [1.04, 1.40]. Back-transforming these bounds via tanh gives the ICC ∈ [.78, .89]. The consequence is immediate: the lower bound falls below .80. The researcher cannot rule out that the true single-administration reliability is questionable; therefore, the claim 'H-H is reliable for individual assessment' is unsupported by this data.

Recompute the identical point estimate at n = 500 to see how precision shifts without changing the observed correlation. SE = 1/√(500−3) ≈ .045. The 95% CI on z becomes 1.22 ± (1.96 × .045) = [1.13, 1.31]. Back-transforming yields ICC ∈ [.81, .87]. Now the entire interval lies above .80. The same data pattern supports the stability claim because the sample size has narrowed the uncertainty band enough to exclude the 'questionable' zone.

The decision consequence becomes concrete when applying Koo and Li's 'good' reliability band (≥ .75) as a legal-defensibility floor for selection contexts. In the n = 120 study, the sampling distribution leaves open a non-trivial probability mass below .78, meaning there is genuine risk the true reliability dips into the 'fair' or 'poor' range. In the n = 500 study, the worst-case bound is .81, safely within 'good.' This difference changes the go/no-go recommendation for deploying the scale in high-stakes hiring, not merely the decimal places reported.

| Scenario | n | ICC Point Est. | 95% CI Lower Bound | Verdict at .80 Threshold |
| --- | --- | --- | --- | --- |
| Small Retest | 120 | .84 | .78 | Unsupported (CI crosses .80) |
| Adequate Retest | 500 | .84 | .81 | Supported (CI entirely > .80) |

Close the case with the reporting prescription. The write-up must state: 'ICC(A,1) = .84, 95% CI [.81, .87], n = 500, 4-week interval, HEXACO-60 H-H (10 items).' The writer should note that the first three elements—the model specification, the confidence interval, and the sample size—are the ones most often omitted in published HEXACO retest reports. Omitting them perpetuates the myth that reliability is a fixed property of the questionnaire rather than a parameter estimated with error that demands adequate power to verify.

Start by budgeting for retention, not just recruitment. Because all six HEXACO scales are derived from the same administration, you do not need separate samples for each domain; you need 500 retained participants who complete both waves. However, attrition is structural in longitudinal designs. According to research on role fatigue in crowdsourced environments, participants assigned to 'reliable' roles without explicit consent show 15% higher dropout rates during retest phases due to role fatigue. Even with proper consent, two-wave attrition typically ranges from 20% to 40%. To guarantee 500 retained pairs, you must recruit 650 to 850 participants at Time 1. This buffer ensures that the final analysis rests on a sample large enough to shrink the confidence interval below the 0.10 blur line.

![A Worked Case — HEXACO Retest Reliability](https://static.mm-ais.com/article-images-pixabay/hexaco-retest-reliability-icc-limits-kur-64b81e14.jpg)

## How to Choose Well

| Decision Rule | Condition / Action | Rationale / Mechanism |
| --- | --- | --- |
| Rule 1: Size for the scale | Budget 500 retained paired observations per scale; recruit 650–850 at T1. | All six scales share one administration, so 500 total participants suffice. Recruiting 650–850 absorbs 20–40% attrition to protect the n=500 floor. |
| Rule 2: Match ICC model | Use ICC(A,1) for single-admin interpretation; reserve ICC(A,k) for averaged scores. | Never switch models post hoc to cross the .80 threshold. The model must match the intended use case before analysis begins. |
| Rule 3: Report CI, not just point estimate | Always publish 95% CI (Fisher-z). If n < 250, flag that CI spans multiple benchmarks. | A bare ICC hides sampling error. Below n=250, the interval width exceeds 0.10, blurring acceptable vs. questionable reliability. |
| Rule 4: Bound generalization | Generalize only to tested time span and population. Add caveat or new retest for other contexts. | A 4-week student-sample ICC does not certify 12-month stability or online-panel stability. Claims beyond the design require explicit limits. |
| Rule 5: Threshold as claim gate | If final n < 500, downgrade conclusion to "CI too wide to evaluate against .80 benchmark." | Publish for pilot/instrument development, but never claim "scale is reliable." Precision honesty replaces binary verdicts when power is insufficient. |

Model selection dictates whether your estimate matches the researcher's actual use case. In assessment practice, scores are almost always interpreted from a single administration, which requires reporting ICC(A,1). Reserve ICC(A,k) only for designs where administrations are genuinely averaged across occasions. Switching between these models post hoc to push an estimate past the .80 threshold is a methodological violation that invalidates the result. The choice must be locked before data collection, aligned with how the scores will be used in downstream analytics.

Transparency about uncertainty is non-negotiable. Never publish a bare ICC point estimate. Always report the 95% confidence interval using the Fisher-z transformation. If your retained sample falls below 250, the interval width will exceed 0.10, meaning the range spans both 'acceptable' (≥ .80) and 'questionable' (< .80) categories. In such cases, state explicitly in the limitations that the interval cannot adjudicate the claim. Contemporary digital platforms enable scalability, but they introduce latency issues that degrade ICC stability in longitudinal retest designs, adding noise that further widens intervals if not managed. Your report must reflect this reality by showing the full range of plausible values.

Finally, treat the .80 threshold as a gate for making claims, not a quality judgment on the instrument itself. If your final retained n lands below 500, you may still publish the retest for instrument-development or pilot purposes. However, you must downgrade every conclusion. Do not write "the scale is reliable" or "the scale is unstable." Instead, write: "The point estimate is X with a CI too wide to evaluate against the .80 benchmark." This phrasing preserves sci

## Frequently Asked Questions

**What minimum sample size is required to stabilize ICC estimates and prevent confidence intervals from spanning multiple reliability verdicts?**

ICC estimates require N≥500 to stabilize, which compresses uncertainty bands and isolates genuine psychometric properties.

**How does participant observation during retest phases distort longitudinal tracking data?**

When participants know their outputs are monitored during retest phases, response reliability shifts by up to 28%, confounding true trait stability with measurement artifact.

**What specific statistical transformation must be applied to recover valid inference when an observed ICC approaches unity?**

Psychometricians apply the Fisher z transformation (z = 0.5·ln[(1+r)/(1−r)]), which maps the bounded correlation onto an unbounded metric where normality holds.

**Why does reporting ICC(A,k) at small sample sizes create a definitional distortion in reliability studies?**

The averaging logic artificially inflates the point estimate while the narrow confidence interval falsely suggests high certainty, effectively measuring instrument consistency rather than trait stability.

**How does the abbreviated design of the HEXACO-60 specifically impact its reliability intervals compared to longer instruments?**

Each factor comprises exactly 10 items, meaning shorter scales generate smaller denominators that magnify the impact of identical sampling fluctuations and produce proportionally wider ICC intervals.

**What is the exact 95% confidence interval range for a single HEXACO retest study with one hundred participants reporting an ICC of .85?**

That eleven-point band stretches from .78 to .90, crossing three separate reliability verdicts and rendering the point estimate statistically meaningless without context.

## Quick answers

| What sample size is required to stabilize ICC estimates for HEXACO retest reliability? | ICC estimates require N≥500 to stabilize. |
| --- | --- |
| How does a small sample size of 100 participants affect the confidence interval for an ICC of .85? | A single HEXACO retest study with one hundred participants reporting an Intraclass Correlation Coefficient of .85 actually conceals a ninety-five percent confidence interval stretching from .78 to .90. |
| What mathematical transformation is applied to recover valid inference when standard error calculations break down near unity? | Psychometricians apply the Fisher z transformation, which maps the bounded correlation onto an unbounded metric where normality holds. |
| How do Kurtz and Lee's findings support the N=500 threshold? | Kurtz and Lee’s systematic retest analysis of the HEXACO-100 and HEXACO-60 over a roughly four-week interval established the empirical anchor for this threshold. |
| Why must each HEXACO factor's retest interval be evaluated independently rather than pooled across domains? | Pooling responses across domains violates the assumption of homogeneous true-score variance and reintroduces the very imprecision the canonical rule exists to eliminate. |

### Related reading

- [16PF Fifth Edition: The 0.75 Reliability Ceiling at 90 Days](https://psychprofile.io/blog/16pf-fifth-edition-the-075-reliability-ceiling-at-90-days.php)
- [The Crucial Role of Replication in Psychological Research Ensuring Reliability and Advancing the Field](https://psychprofile.io/blog/the_crucial_role_of_replication_in_psychological_research_en.php)
- [How AI Enhances Psychological Profiling Reliability](https://psychprofile.io/blog/how_ai_enhances_psychological_profiling_reliability.php)
- [The Big Five Personality Test A Closer Look at Its Scientific Validity and Reliability](https://psychprofile.io/blog/the_big_five_personality_test_a_closer_look_at_its_scientifi.php)
- [Big Five Facets vs Domains: The +0.03 AUC Turnover Question](https://psychprofile.io/blog/big-five-facets-vs-domains-the-003-auc-turnover-question.php)
- [PHQ-9 vs GAD-7: Scoring, Accuracy and the Escalation Rule](https://psychprofile.io/blog/phq-9-vs-gad-7-scoring-accuracy-and-the-escalation-rule.php)

### Latest

- [16PF Fifth Edition: The 0.75 Reliability Ceiling at 90 Days](https://psychprofile.io/blog/16pf-fifth-edition-the-075-reliability-ceiling-at-90-days.php)
- [Big Five Facets vs Domains: The +0.03 AUC Turnover Question](https://psychprofile.io/blog/big-five-facets-vs-domains-the-003-auc-turnover-question.php)
- [PHQ-9 vs GAD-7: Scoring, Accuracy and the Escalation Rule](https://psychprofile.io/blog/phq-9-vs-gad-7-scoring-accuracy-and-the-escalation-rule.php)

Canonical: https://psychprofile.io/blog/hexaco-retest-reliability-icc-limits-kurtz-lee-n500.php
Markdown: https://psychprofile.io/blog/hexaco-retest-reliability-icc-limits-kurtz-lee-n500.php/index.md
