# PHQ-9 vs GAD-7: Scoring, Accuracy and the Escalation Rule

Gavin Marshall · August 26, 2026

> PHQ-9 vs GAD-7: Scoring, Accuracy and the Escalation Rule. Nine questions versus seven, under three minutes each — yet at a typical...

| Takeaway | Detail |
| --- | --- |
| A mid-range PHQ-9 score confirms less often than it flags | At a typical clinic base rate, a PHQ-9 of exactly 10 is a true positive only about one time in three — which is why the score starts a conversation instead of ending one. |
| The GAD-7 screens a condition affecting hundreds of millions worldwide | WHO reports 359 million people were living with an anxiety disorder in 2021, including 72 million children and adolescents, with generalized anxiety disorder specifically characterized by excessive worry. |
| Escalation triggers sit outside the numeric score | NHS assessment treats psychosis developing in a few people with severe depression as a severity flag, and requires distinguishing grief — an entirely natural response to loss — from depression, an illness, despite their shared characteristics. |
| The screening workflow itself is now publicly disputed | Kaiser therapists publicly allege the health system's new mental health screening puts patients at higher risk by delaying care — one said 'Thank God they're still alive' — while Kaiser published 'Update on Mental Health Program Progress' and Northwest workforce-expansion news in the same 2026 news cycle. |

Nine questions versus seven, under three minutes each — yet at a typical clinic base rate, a PHQ-9 of exactly 10 is a true positive only about one time in three. That ratio is why the 2026 upgrade refuses to let either number masquerade as a diagnosis: the pair works as triage, and its real signal is the change between administrations, not the first score.

The stakes are not abstract. The World Health Organization counted 359 million people living with an anxiety disorder in 2021 — 72 million of them children and adolescents — characterized by excessive fear and worry severe enough to cause significant distress or impairment. Depression screening walks through NHS-enumerated territory: continuous low mood, hopelessness and helplessness, low self-esteem, tearfulness, guilt. Neither questionnaire hands down a diagnosis; both decide who gets a longer conversation.

Escalation exists because some findings outrank any score: NHS assessment treats psychosis emerging in severe depression as a severity flag, and separates grief — an entirely natural response to loss — from depression, an illness, despite their shared surface. The workflow itself is contested ground in 2026: Kaiser therapists publicly claim its new screening system delays care — 'Thank God they're still alive,' one said — while Kaiser simultaneously published 'Update on Mental Health Program Progress.'

![PHQ-9 vs GAD-7](https://static.mm-ais.com/article-images-ai/phq-9-vs-gad-7-scoring-accuracy-and-the-ai-af3298ef.jpg)

## Two Scoring Engines: How 9 Items Build a 0

Neither instrument weights its items or applies latent-trait scoring — each is a plain sum of ordinal ratings, engineered around a different DSM construct by nearly the same author team fifteen years apart. According to Spitzer, Kroenke, and Williams' 1999 publication, the PHQ-9 began life as the self-report offshoot of the PRIME-MD clinician interview: nine items, each mapping one-to-one onto a DSM-5 major-depression criterion, rated 0–3 over the past two weeks and summed to 0–27. Item 9 is the sole suicidal-ideation probe — it feeds the total like every other item, but clinicians typically review it on its own, because one point there carries signal the sum dilutes.

The GAD-7 arrived in 2006, built by Spitzer, Kroenke, Williams, and Bernd Löwe around the DSM generalized-anxiety core: nervousness, uncontrollable worry, restlessness, irritability. Seven items, same 0–3 anchors, summed to 0–21. According to Löwe's 2008 validation in Medical Care, the scale outgrew its original target — at slightly lower cutoffs, the same seven items also flag panic disorder and social anxiety, which is why a modest GAD-7 total can still justify a structured follow-up.

The severity bands look arbitrary until you trace their origin. The PHQ-9 cuts at 5/10/15/20 (mild, moderate, moderately severe, severe); the GAD-7 at 5/10/15 (mild, moderate, severe). Those breakpoints sit where each instrument's sensitivity and specificity curves cross optimally — empirical trade-off points, not clinical intuition. And the trap this banding breeds deserves killing outright: "moderately severe" on the PHQ-9 does not mean you have moderately severe depression. The label grades the symptoms you reported, not the probability you have the disorder — at population base rates, most positive screens are false alarms, so any score ≥10 functions as a prompt for structured follow-up, never a diagnosis.

Delivery is fully automated inside Kaiser's pipeline. Both instruments load as Epic HealthConnect flowsheets at check-in, auto-sum, file straight to the chart, and surface to the primary care physician and behavioral health. They repeat annually under USPSTF mandates: all-adult depression screening, plus anxiety screening for adults 19–64 — pregnant and postpartum members included — since June 2022. When your 2026 check-in serves both flowsheets, that pairing is deliberate, not redundant.

On reliability, both are sound rulers. Across validation samples, internal-consistency alphas run roughly 0.86–0.89 for the PHQ-9 versus 0.92 for the GAD-7, with test–retest correlations near 0.83–0.84. High alpha signals internal coherence, not interchangeability: the scales share three authors, not an item set, so their totals quantify different constructs.

| Property | PHQ-9 | GAD-7 |
| --- | --- | --- |
| Items × response scale | 9 items × 0–3 | 7 items × 0–3 |
| Authors, year | Spitzer, Kroenke & Williams, 1999 | Spitzer, Kroenke, Williams & Löwe, 2006 |
| Construct anchor | One item per DSM-5 major-depression criterion | DSM generalized-anxiety core symptoms |
| Total range | 0–27 | 0–21 |
| Severity bands | 5 / 10 / 15 / 20 | 5 / 10 / 15 |
| Internal consistency (α) | Roughly 0.86–0.89 | Roughly 0.92 |
| Test–retest correlation | Near 0.83–0.84 | Near 0.83–0.84 |

The asymmetry that matters operationally: each screener takes under three minutes, but the PHQ-9's nine graded points resolve smaller within-person shifts than the GAD-7's seven. Because treatment is evaluated by score change at consistent 4-week intervals — against each instrument's minimal clinically important difference — the denser ruler wins whenever the question is whether anything moved, while the GAD-7 nets broader anxiety pathology the PHQ-9 never touches. Complete both flowsheets at every check-in even if one domain feels clean: the ≥10 escalation fires on either instrument, and a change-based read needs baselines on both.

![Two Scoring Engines: How 9 Items Build a 0 — PHQ-9 vs GAD-7](https://static.mm-ais.com/article-images-ai/phq-9-vs-gad-7-scoring-accuracy-and-the-ai-7a49ed86.jpg)

## The Accuracy Ledger: 88/88 Versus 89/82

Two ratios justify the entire escalation rule: 88/88 and 89/82. When Kroenke, Spitzer, and Williams validated the PHQ-9 in the Journal of General Internal Medicine in 2001, they scored the nine-item sum against structured psychiatric interviews in primary-care samples and obtained 88% sensitivity and 88% specificity for major depression at the ≥10 cutoff. That symmetry is rare in screening psychometrics — most instruments buy sensitivity by spending specificity — and it is why ≥10 became the default trigger in screening protocols rather than a vendor's guess.

The founding numbers then survived a twenty-year stress test. According to Levis et al.'s 2019 meta-analysis in BMJ — 58 studies, 17,357 participants — the PHQ-9 at ≥10 delivers pooled sensitivity of 0.88 and specificity of 0.85. Read that as a psychometrician would: after two decades of translations, different health systems, and very different base rates, sensitivity did not move at all and specificity gave up only three points. Instruments that drift under transportability get recalibrated or retired; this one held its operating point, which is what keeps a fixed ≥10 threshold defensible in the current protocol.

The GAD-7's ledger entry is deliberately asymmetric. The overlapping author team validated it in Archives of Internal Medicine across 2,740 primary-care patients: at ≥10, 89% sensitivity and 82% specificity for generalized anxiety disorder. That trade — one point more sensitivity purchased with six points less specificity — is the statistical signature of an instrument built to net broader anxiety pathology rather than a single diagnostic category. It will flag more people who need a conversation, and more people who don't.

Transportability holds on the anxiety side too. According to Plummer et al.'s 2016 meta-analysis in BMJ Open, GAD-7 accuracy is stable across countries and clinical settings — evidence that the cutoff travels beyond the original US validation sample. Both screeners, in other words, passed the replication standards that retired plenty of older instruments.

Accuracy on paper is worthless if the form never gets administered, which is where the deployment row earns its place. According to Kaiser Permanente Northern California implementation reports, completed PHQ-9 screenings climbed from roughly 43% to about 84% of eligible adult primary-care visits within three years of the universal-screening mandate, with corresponding rises in identified and treated depression. Screening fewer than half of eligible visits is a case-finding failure no matter how good the instrument is; roughly doubling completion is what converts a valid tool into detected cases. Workflow placement moves real-world yield more than a few points of classification accuracy ever will.

Now kill the misreading this ledger invites. None of these percentages describes you. Sensitivity and specificity are properties of the instrument measured against a structured-interview gold standard — they are not the probability that your score reflects a disorder. An 88% sensitivity figure does not mean a positive screen carries an 88% chance of being "real"; at population base rates, most positive screens are false alarms, which is exactly why the protocol treats any score ≥10 as a prompt for structured diagnostic follow-up — never a diagnosis. What the ledger actually supports is narrower: both instruments sit near the practical ceiling for brief self-report screeners, so the choice between them turns less on a few points of accuracy and more on which disorder domain and which job — case-finding versus change-tracking — each performs. That is the argument for administering both.

| Evidence tier | Source | Figure | What it establishes |
| --- | --- | --- | --- |
| PHQ-9, founding | Kroenke, Spitzer & Williams, Journal of General Internal Medicine (2001) | 88% sensitivity / 88% specificity at ≥10 | Cutoff validated against structured psychiatric interviews in primary care |
| PHQ-9, replication | Levis et al., BMJ (58 studies, 17,357 participants) | Pooled 0.88 / 0.85 at ≥10 | Original numbers held for two decades |
| GAD-7, founding | Spitzer et al., Archives of Internal Medicine (2,740 patients) | 89% sensitivity / 82% specificity at ≥10 | Broader anxiety net at a modest specificity cost |
| GAD-7, replication | Plummer et al., BMJ Open (2016) | Stable across countries and clinical settings | Cutoff transports beyond the original US sample |
| Deployment at scale | Kaiser Permanente Northern California implementation reports | Completion ~43% → ~84% of eligible adult visits within three years | Mandate plus workflow raised identified and treated depression |

![The Accuracy Ledger: 88/88 Versus 89/82 — PHQ-9 vs GAD-7](https://static.mm-ais.com/article-images-pixabay/phq-9-vs-gad-7-scoring-accuracy-and-the-d8b84a8c.jpg)

## Side-by-Side Verdict

Nine items versus seven reads like a rounding difference until you notice the two instruments were engineered for different jobs. Side by side, the verdict splits cleanly: the PHQ-9 wins the longitudinal job, the GAD-7 wins the breadth job. The PHQ-9's 0–27 span gives depressive severity finer gradation than any rival brief screen, it carries an established remission marker, and it sits inside the deepest treatment-response literature of any short depression measure — exactly what you want when progress is judged by score change rather than by level. The GAD-7's edge is coverage: one seven-item pass nets generalized anxiety disorder — a condition the World Health Organization characterizes centrally by excessive worry — and still returns usable signals for panic and social anxiety. Two jobs, two winners, no overall champion.

The case for running both is arithmetic, not caution. Roughly half of primary-care patients who meet criteria for depression also meet criteria for an anxiety disorder, so any single-instrument screen structurally misses the second condition in about one such patient in two. Dual administration costs under six minutes total — the cheapest insurance in the entire assessment.

| Dimension | PHQ-9 | GAD-7 | Edge |
| --- | --- | --- | --- |
| Length & time | 9 items; typically a few minutes | 7 items; slightly faster | GAD-7 |
| Score range | 0–27, finer gradation for tracking change | 0–21, sufficient for severity triage | PHQ-9 |
| Severity bands | Five: minimal, mild, moderate, moderately severe, severe | Four: minimal, mild, moderate, severe | PHQ-9 |
| Sensitivity/specificity at ≥10 | Balanced at the cutoff (see "The Accuracy Ledger") | Balanced at the cutoff (see "The Accuracy Ledger") | Tie |
| Target disorders | Major depression, built on DSM criterion symptoms | Generalized anxiety first; usable panic and social-anxiety signals | GAD-7 |
| Minimal clinically important difference | Within-person change benchmark backed by the deepest response literature | Within-person change benchmark, thinner validation base | PHQ-9 |
| Remission threshold | Established marker: return to the minimal band | Less standardized; most teams use return to the minimal band | PHQ-9 |

Escalation is symmetric, and the symmetry is the design. A score of 10 or higher on either instrument triggers the identical next step: a structured diagnostic follow-up — a SCID-style interview or a Kaiser behavioral-health consult. Neither questionnaire diagnoses anything on its own; they route, they don't conclude. The stakes of keeping a human in that loop are live right now: according to CalMatters' reporting, a union alleges Kaiser used an algorithm rather than clinicians to triage mental-health patients, and Kaiser therapists publicly claim the new screening system delays care. The remedy for a flawed workflow is more clinician judgment after the screen — never treating a band label as a verdict. "Moderately severe" on the PHQ-9 does not mean you have moderately severe depression; it labels the severity of the symptoms you reported, and at population base rates most positive screens are false alarms.

The economize branch is narrow. Order the PHQ-9 alone for somatic presentations — sleep, fatigue, appetite — with no worry content; order the GAD-7 alone for pure worry or panic presentations. Run both whenever the picture is ambiguous, and always run both when a prior score sits in the 5–9 gray zone: near the escalation line, the meaningful signal is cross-domain movement, and a lone instrument cannot detect change in the domain it never measured. If your last screen landed there, bring both forms to your next 4-week check and read the deltas, not the levels.

![Side-by-Side Verdict — PHQ-9 vs GAD-7](https://static.mm-ais.com/article-images-pixabay/phq-9-vs-gad-7-scoring-accuracy-and-the-b0204e89.jpg)

## What the Data Doesn't Tell You

Content for What the Data Doesn't Tell You is being prepared.

## The False-Positive Trap

A positive screen on either instrument is, at population base rates, more often wrong than right — and the gap between "elevated score" and "has the disorder" is pure arithmetic, not opinion. Apply Bayes' theorem to the pooled accuracy figures covered above, assume an 8% depression base rate, and a PHQ-9 total of 10 or higher carries a positive predictive value of only about 34% — roughly two of every three flags dissolve under a structured interview. Run the same operation for the GAD-7 at a generalized-anxiety base rate near 3%, and a score of 10 or higher lands around 13%: nearly seven of eight anxiety flags are false alarms. The screener did not fail; the low prior did the damage.

That ratio kills the most seductive misreading in the entire pipeline: a "moderately severe" band on the PHQ-9 labels the severity of symptoms you reported, not the probability that you have moderately severe depression. At an 8% base rate, most members wearing that label are false positives. An assessment — in the words Assessment.com published on August 26, 2026, "a systematic process for collecting, scoring, interpreting, and using evidence about a person… for a defined purpose" — serves its defined purpose only when the purpose is respected, and Kaiser's 2026 screen exists for triage, not classification.

| Quantity | PHQ-9 | GAD-7 |
| --- | --- | --- |
| Escalation threshold | Total ≥ 10 | Total ≥ 10 |
| Base-rate assumption | 8% depression | ~3% generalized anxiety |
| Positive predictive value | About 34% | Near 13% |
| Positives that are false alarms | Roughly two of three | Nearly seven of eight |

The somatic-confound problem stacks on top of the base-rate problem. The PHQ-9 counts fatigue, sleep disruption, and appetite change — precisely the physiology that hypothyroidism, anemia, and obstructive sleep apnea produce — while the GAD-7's restlessness and autonomic items inflate with caffeine, stimulants, and hyperthyroidism. Neither instrument was engineered to separate endocrine or hematologic disease from mood disorder; both were engineered to count symptoms. The defensible sequence therefore runs the medical workup first: a thyroid evaluation, iron studies, and a sleep history can move a total across the ≥10 threshold without touching a single cognition. A score earned on four hours of apnea-fragmented sleep is not evidence of depression; it is evidence of four hours of sleep.

Item 9 deserves its own caution. The PHQ-9 devotes one item to suicidal ideation, and its standalone predictive validity is weak in both directions — a zero does not clear risk, and a one does not establish it. Its correct function is routing: a flag that summons a live risk assessment, never a substitute for one. Cohen's *Psychological Assessment* textbook draws the exact line: testing obtains a numerical gauge, summed by a practically substitutable technician, while assessment answers a referral question through an assessor who selects tools and integrates data. Item 9 is testing. Risk determination is assessment.

Invariance is the quieter threat. Published bifactor and differential-item-functioning analyses of the PHQ-9 show its somatic and affective items behaving differently across sex, age, language, and cultural groups — anyone who validates inventories learns to distrust a raw total that assumes those items mean the same thing everywhere. Identical totals may not encode identical severity across populations, which matters when a system aggregates scores across a diverse membership and compares trend lines between groups.

Then there is what the pair lacks entirely: a validity apparatus. The MMPI-2-RF carries VRIN and TRIN inconsistency scales because self-report invites careless and motivated responding; the PHQ-9 and GAD-7 carry nothing comparable. Kaiser's trend lines inherit whatever response style the member brings that day — minimization from someone dreading a leave-of-absence dispute, amplification from someone whose disability paperwork depends on the number — with no built-in check flagging either.

Finally, bound the trajectory claim. One administration cannot distinguish a new episodic depression from chronic dysthymia or an adjustment reaction; classification requires serial scores across at least two timepoints, read as change at the protocol's consistent 4-week intervals. The NHS marks a parallel boundary between grief and depression — grief is an entirely natural response to loss while depression is an illness, though the two share many characteristics — and no nine-item sum separates them on day one.

Read the threshold as a set of permissions. A score of 10 or higher grants permission to escalate — to a structured diagnostic interview and, where the somatic picture warrants, a medical workup. It grants nothing else: not a diagnosis, not a risk determination, not a trajectory label.

| Observation | It licenses | It forbids |
| --- | --- | --- |
| PHQ-9 ≥ 10 at 8% base rate | Structured diagnostic interview | Reading ~34% PPV as diagnosis |
| GAD-7 ≥ 10 at ~3% base rate | Interview plus medical workup | Reading ~13% PPV as diagnosis |
| Item 9 scored 1 | Live clinician risk assessment | Standalone risk rulings |
| One elevated administration | A baseline timepoint | Episode vs. dysthymia vs. adjustment calls |
| Serial scores at 4-week intervals | Change judged against each MCID | Level-only verdicts |

## Case File

A 34-year-old member completes both screeners at a routine annual visit in 2026 and produces the exact profile the dual-administration protocol exists for: PHQ-9 = 14, moderate band; GAD-7 = 16, severe band; item 9 endorsed at 1 — thoughts of being better off dead on several days. What the Kaiser pathway triggers is deliberately unglamorous: a same-day behavioral-health consult offer plus a brief suicide-risk check. What it does not trigger is a prescription keyed to the number. The score opens a structured conversation; it does not write the order.

Read the 14 honestly and it claims less than it appears to. Fourteen out of 27 is a moderate burden of symptoms self-reported on one day — nothing more. Consistent with the base-rate arithmetic covered above, the working hypothesis stays exactly where the screener leaves it: possible depressive and anxiety disorder, pending structured interview. These are not exotic conditions — according to the WHO, nearly 1 in 7 people worldwide, about 1.1 billion, were living with a mental disorder in 2021, with anxiety and depressive disorders the most common — and prevalent conditions still convert positive screens into false alarms at high rates. The behavioral-health consult supplies the structured interview that upgrades "possible" into either a working diagnosis or a cleared file. The band label grades the symptoms you reported; it does not certify the disorder.

Four weeks later, with CBT underway, the member retakes both instruments at the protocol's fixed interval. The PHQ-9 falls 14 → 8, a six-point drop; the GAD-7 falls 16 → 13, three points. Set against the minimal clinically important differences — 5 points for the PHQ-9 (Löwe et al., 2004) and 4 points for the GAD-7 — the two changes earn opposite verdicts. The depression delta clears its MCID: real signal. The anxiety delta falls a point short: indistinguishable from measurement noise. Two declining lines look like uniform progress; the change-based read shows one confirmed gain and one open question.

Formal criteria sharpen it further. A 14-to-8 drop is a 43% reduction — clinically meaningful, yet short of the conventional ≥50% cutoff that defines "response" in treatment trials. The member is therefore improving (MCID cleared) without having responded (≥50% unmet) and nowhere near remitted (a third threshold entirely). "Improving," "responded," and "remitted" are three distinct verdicts with three distinct thresholds, and collapsing them is how a treatment gets graded a success or a failure at precisely the wrong moment.

The finish line is pre-committed: treatment continues until the PHQ-9 holds at 4 or below across consecutive 4-week checks, while the GAD-7 keeps being scored against its own MCID at the same intervals. The branches run both directions, too. Had the GAD-7 climbed to 18 at week four instead of easing to 13, the identical change-logic would route the member to a step-up review — medication evaluation or intensive-outpatient referral. After day zero, direction and magnitude of change drive every branch in this pathway; the intake number drives none of them.

| Checkpoint | PHQ-9 | GAD-7 | Threshold tested | Branch executed |
| --- | --- | --- | --- | --- |
| Week 0, annual visit 2026 | 14 (moderate band) | 16 (severe band); item 9 = 1 | Escalation at any score ≥10 on either screener | Same-day behavioral-health consult offer plus brief suicide-risk check; no prescription keyed to the number |
| Week 4, after CBT begins | 8 | 13 | MCID: 5 points PHQ-9 (Löwe et al., 2004); 4 points GAD-7 | −6 clears the PHQ-9 MCID (real signal); −3 misses the GAD-7 MCID (noise) — continue CBT unchanged |
| Week 4 response audit | 43% reduction | Roughly 19% reduction | Conventional response definition: ≥50% symptom reduction | "Improving" on both arms; neither reaches formal "response" |
| Ongoing 4-week cycles | Target: hold ≤4 | Rescored each cycle | PHQ-9 remission threshold (≤4 sustained); GAD-7 judged by its 4-point MCID | Treatment continues until the remission-level PHQ-9 holds across consecutive checks |
| Counterfactual week 4 | 8 (course unchanged) | 18 | Direction-and-magnitude review | Step-up review: medication evaluation or intensive-outpatient referral |

## Five Rules for Reading Your 2026 Screen

Content for Five Rules for Reading Your 2026 Screen is being prepared.

## What to do next

| Step | Action | Why it matters |
| --- | --- | --- |
| 1 | Complete both screeners in full at baseline — all 9 PHQ-9 items and all 7 GAD-7 items, summed as plain ordinal ratings with no item weighting, the way Spitzer, Kroenke, and Williams designed them off the PRIME-MD interview. | Neither instrument applies latent-trait scoring, so a skipped item silently corrupts the total; both take under three minutes each, so there is no reason to administer only one. |
| 2 | Treat any score ≥10 on either questionnaire as a prompt for a structured diagnostic follow-up — never write it down as a diagnosis. | At a typical clinic base rate, a PHQ-9 of exactly 10 is a true positive only about one time in three; the score starts a conversation instead of ending one. |
| 3 | Before acting on any total, run the escalation checks that sit outside the numbers: ask whether psychosis has developed alongside severe depression (an NHS-enumerated severity flag) and whether low mood follows a loss — separating grief, a natural response, from depression, an illness. | Some findings outrank any score; a moderate total with a psychosis flag escalates faster than a high total without one. |
| 4 | Re-administer both instruments at consistent 4-week intervals — same questions, same gap — rather than re-screening ad hoc when someone remembers. | The pair works as triage, and its real signal is the change between administrations, not the first score; irregular spacing makes that change unreadable. |
| 5 | Evaluate every treatment decision by score change against each instrument's minimal clinically important difference (MCID), not against the raw baseline number. | A falling total only counts as improvement once it crosses the MCID; otherwise you are reading noise as response. |
| 6 | If screening becomes a gate instead of a doorway — the Kaiser dispute, where therapists allege the new system delays care ("Thank God they're still alive") even as "Update on Mental Health Program Progress" was published — push a flagged score through to an actual clinician visit. | The 2026 news cycle shows the workflow itself is contested ground; a screener that substitutes for care puts patients at higher risk, which defeats the entire point of running it. |

## Frequently Asked Questions

**My PHQ-9 came back exactly 10 — doesn't that mean I have depression?**

At a typical clinic base rate, a PHQ-9 of exactly 10 is a true positive only about one time in three, which is why any score ≥10 functions as a prompt for structured follow-up rather than a diagnosis.

**Question 9 on the PHQ-9 asks about self-harm — why do clinicians look at it separately if it just adds points to the total?**

Item 9 is the sole suicidal-ideation probe — it feeds the total like every other item, but clinicians typically review it on its own because one point there carries signal the sum dilutes.

**Does the GAD-7 only pick up generalized anxiety disorder?**

According to Löwe's 2008 validation in Medical Care, at slightly lower cutoffs the same seven items also flag panic disorder and social anxiety, which is why a modest GAD-7 total can still justify a structured follow-up.

**How accurate is the PHQ-9 at the standard ≥10 cutoff?**

The 2001 validation by Kroenke, Spitzer, and Williams obtained 88% sensitivity and 88% specificity for major depression, and Levis et al.'s 2019 BMJ meta-analysis of 58 studies and 17,357 participants pooled those figures at 0.88 sensitivity and 0.85 specificity.

**How often will I be handed these screeners at check-in?**

Both instruments repeat annually under USPSTF mandates — all-adult depression screening plus anxiety screening for adults 19–64, pregnant and postpartum members included — in effect since June 2022.

**Is there anything that triggers escalation even when my numeric score is low?**

Yes — NHS assessment treats psychosis emerging in a few people with severe depression as a severity flag that outranks any score, and requires distinguishing grief, an entirely natural response to loss, from depression, an illness.

## Quick answers

| What are the total score ranges of the PHQ-9 and GAD-7? | The PHQ-9 sums nine items rated 0–3 over the past two weeks to a total of 0–27, while the GAD-7 sums seven items with the same 0–3 anchors to a total of 0–21. |
| --- | --- |
| How accurate was the PHQ-9 at the ≥10 cutoff in its validation study? | When Kroenke, Spitzer, and Williams validated it in 2001 against structured psychiatric interviews in primary-care samples, the PHQ-9 obtained 88% sensitivity and 88% specificity for major depression at the ≥10 cutoff. |
| Why does a mid-range PHQ-9 score start a conversation instead of ending one? | At a typical clinic base rate, a PHQ-9 of exactly 10 is a true positive only about one time in three, so most positive screens are false alarms and any score ≥10 functions as a prompt for structured follow-up, never a diagnosis. |
| What escalation triggers sit outside the numeric scores? | NHS assessment treats psychosis emerging in a few people with severe depression as a severity flag, and requires distinguishing grief — an entirely natural response to loss — from depression, an illness, despite their shared characteristics. |
| Which screener resolves smaller within-person shifts, and why does that matter? | The PHQ-9's nine graded points resolve smaller within-person shifts than the GAD-7's seven, which matters because treatment is evaluated by score change at consistent 4-week intervals against each instrument's minimal clinically important difference. |

Also worth reading: **Kaiser Permanente's Dual Role as Insurer-Provider A Critical Analysis of Healthcare Quality Impact**: [Kaiser Permanente's Dual Role as](https://psychprofile.io/blog/kaiser_permanente_s_dual_role_as_insurer_provider_a_critical.php) · **Unlocking Personality Insights With AI Assessment for Mental Health**: [Unlocking Personality Insights With AI](https://psychprofile.io/blog/unlocking_personality_insights_with_ai_assessment_for_mental.php) · **Calculating Mental Health Prevalence A 7-Step Guide for Accurate Assessment**: [Calculating Mental Health Prevalence A](https://psychprofile.io/blog/calculating_mental_health_prevalence_a_7_step_guide_for_accu.php)

### Related reading

- [New Diagnostic Tools Enhance Accuracy in Adult Autism Spectrum Disorder Identification](https://psychprofile.io/blog/new_diagnostic_tools_enhance_accuracy_in_adult_autism_spectr.php)
- [Digital Evolution of MMPI Testing A 2025 Analysis of Online Administration Accuracy and Clinical Validity](https://psychprofile.io/blog/digital_evolution_of_mmpi_testing_a_2025_analysis_of_online.php)
- [Uncovering False Positives Examining AI Detection Tools' Accuracy in Academic Integrity](https://psychprofile.io/blog/uncovering_false_positives_examining_ai_detection_tools_acc.php)
- [The Nuanced Reality Evaluating the Accuracy of Advanced Personality Tests in 2024](https://psychprofile.io/blog/the_nuanced_reality_evaluating_the_accuracy_of_advanced_pers.php)
- [Comparing the Big Five and MBTI A 2024 Analysis of Personality Test Accuracy](https://psychprofile.io/blog/comparing_the_big_five_and_mbti_a_2024_analysis_of_personali.php)
- [Understanding the Nuances of Empathy Quotient Test Scoring A 2024 Analysis](https://psychprofile.io/blog/understanding_the_nuances_of_empathy_quotient_test_scoring_a.php)

### Latest

- [2026 Meta-Analysis: Faking Adds 0.5 SD — Do Scales Catch It?](https://psychprofile.io/blog/2026-meta-analysis-faking-adds-05-sd-do-scales-catch-it.php)
- [Wysa vs ChatGPT: d≈0.5, 5–14 PHQ-9 Band, n=129, No RCT](https://psychprofile.io/blog/wysa-vs-chatgpt-d05-514-phq-9-band-n129-no-rct.php)
- [HEXACO-60 BN vs IRT: AUC myths, bias traps, and key factors.](https://psychprofile.io/blog/hexaco-60-bn-vs-irt-auc-myths-bias-traps-and-key-factors.php)

Canonical: https://psychprofile.io/blog/phq-9-vs-gad-7-scoring-accuracy-and-the-escalation-rule.php
Markdown: https://psychprofile.io/blog/phq-9-vs-gad-7-scoring-accuracy-and-the-escalation-rule.php/index.md
