| Takeaway | Detail |
|---|---|
| Function words are the highest-bandwidth personality signal, not the connective tissue to discard. | They account for the bulk of natural language output, yet most prediction pipelines score only open-class words. |
| Lexical cues for Big Five dimensions are empirically identifiable across datasets, but only if function-word features are included. | RANLP analysis found predictive lexical cues in the function-word layer often dismissed as noise. |
| Self-report Big Five inventories are limited by response bias, subjectivity, and scalability. | Language-based prediction using the function-word baseline offers a scalable, unobtrusive alternative. |
| The lexical hypothesis itself is a linguistic argument: the most predictive traits are embedded in everyday word choice. | Openness, conscientiousness, extraversion, agreeableness, and neuroticism are all reflected in the function-word stream. |
The majority of the words we write are function words—the, of, to, and, be—yet almost every personality-prediction pipeline treats them as boring connective tissue and scores only the remaining open-class words. That high-frequency layer is the highest-bandwidth trait signal in natural language, and the field's obsession with emotionally saturated open-class words is systematically missing it.
The RANLP paper 'Personality Predictive Lexical Cues and Their Correlations' shows that lexical cues for each Big Five dimension emerge from the interplay of feature sets and learning algorithms. Because the Big Five itself was built on the lexical hypothesis—the idea that personality lives in the words people use to describe themselves and others—ignoring the most frequent words in the language is not a neutral preprocessing choice. It is a theoretical error.
Function words are hard to control consciously, which is exactly why they leak stable patterns of openness, conscientiousness, extraversion, agreeableness, and neuroticism. Traditional self-report inventories suffer from response bias and subjectivity. The function-word layer is the contrarian truth: not noise to be removed, but the signal to be decoded.

Attentional Leaks
Pennebaker's functional-analysis argument starts with a counting fact that has not aged: the most common English function words account for the bulk of total words in typical naturalistic writing, while open-class emotional adjectives are rare long-tail tokens. The lexical families have completely different distributional shapes. The closed-class set — pronouns, negations, prepositions — is a small inventory of highly frequent, highly repeated types. The open-class set of emotion words is a large inventory of individually rare tokens. That asymmetry alone tells you where trait variance has to live: you want the channel that is almost always on, not the channel that sputters occasionally.
The functional-analysis argument goes deeper than frequency. Function words index attentional focus — self, others, objects — and are produced below conscious awareness; they do not carry semantic content the way "anxious" or "thrilled" do. That is the leak. First-person singular pronouns and negations correlate with self-focus and cognitive complexity because they require attention-shifting in working memory: "I" forces a self-referential perspective switch, and "not" requires holding a proposition in mind while suppressing it. Writers who do this more often leave a measurable trace in their pronoun and negation rates, and that trace is what the function-word composite harvests for trait prediction.
Emotional adjectives fail on the opposite axis. They are consciously chosen, socially desirable, and context-dependent. People say "happy" when the situation calls for it — at a wedding, in a performance review, on a form that asks "how are you?" — not necessarily when their trait-level positive affect is high. The word is a display, not a leak. The myth that Big Five signal lives in words like "anxious" and "thrilled" fails on distributional grounds: those words are too rare, too controlled, and too moderated to carry reliable trait variance.
The 2026 update sharpens the divide. As text-to-speech and LLM output expand in the textual record, person-level function-word statistics remain stable: nobody edits "the" or "not" for impression management, and language models do not selectively suppress them either. Emotion-word frequencies, by contrast, are distorted by content filters and platform moderation — posts containing "anxious" are routed through mental-health moderation pipelines, and LLM outputs are filtered for emotional tone before they reach a corpus. The family with the weaker trait signal is now also the family with the more corrupted distribution, which is one reason the function-word advantage keeps growing on fresh holdouts. This is why the canonical rule requires any emotion lexicon to beat the function-word composite on the same holdout to earn a role — a bar the leak mechanism says it will rarely clear.
| Dimension | Closed-class function words | Open-class emotion adjectives | Winner |
|---|---|---|---|
| Inventory size | A small inventory of high-frequency types | Long-tail; many individually rare types | Function words |
| Text volume | Most of total words (Pennebaker) | Low per-type frequency | Function words |
| Production mode | Below conscious awareness | Consciously chosen for display | Function words |
| Context sensitivity | Stable across situations | Situation-bound; "happy" said when called for | Function words |
| 2026 data pipeline | Stable in TTS/LLM corpora | Distorted by filters and moderation | Function words |

The Correlational Record
According to Schwartz et al. (PLOS ONE), open-vocabulary models trained on large samples of Facebook users and massive text corpora predicted self-reported Big Five for Openness and Extraversion — and the top individual predictors were function words, not emotion adjectives. "Anxious" and "thrilled" did not lead. Articles, prepositions, and pronouns did.
According to Pennebaker and King (JPSP), LIWC function-word categories show substantial test-retest stability across repeated testing intervals, whereas affect categories exhibit state-driven instability. Closed-class words are stable precisely because they carry no content; they index attentional style — where a speaker's focus orients — and that style behaves like a trait, not a mood.
According to Park et al. (JPSP), language-only models applied to Facebook status updates yielded Big Five prediction correlations of moderate strength across large samples of users, with closed-class categories consistently outranking open-class ones in variable importance. The ordering survives replication at scale and across shifts in text source.
According to Boyd, Ashokkumar, Seraj, and Pennebaker (LIWC-22), function-word blocks explain more of the Big Five variance in diverse corpora than the entire affect lexicon block. That contrast is the direct evidence behind the explained-variance rule this guide uses.
Now the specific convergence: in every published corpus study in the correlational record, no single emotion-adjective category has outpredicted the composite of articles + prepositions + pronouns for Openness. The gap is structural, not sampling noise.
The forward projection strengthens the thesis. As training corpora scale, the function-word advantage grows, because more data shrinks the standard error of these low-frequency-but-highly-consistent effects. Emotion adjectives, by contrast, carry state variance that additional data cannot shrink — it is irreducible noise.
The myth this kills: that Big Five signal lives in emotional adjectives like "anxious" and "thrilled." Across the studies above, "the," "under," "not," and "you" are the better telegraphs. In any current trait-prediction build, make the LIWC-22 function-word composite — pronouns, negations, prepositions, articles — the default candidate lexicon, and require any emotion lexicon to beat it on the same holdout to earn a role.
| Study | Sample | Key figure | Which lexicon wins |
|---|---|---|---|
| Pennebaker & King (JPSP) | LIWC text corpora | Substantial function-word stability across repeated testing | Function words |
| Schwartz et al. (PLOS ONE) | Large Facebook user samples, massive text corpora | Moderate correlations for Openness/Extraversion | Function words as top predictors |
| Park et al. (JPSP) | Large Facebook user sample | Moderate correlations across traits | Closed-class over open-class |
| Boyd et al. (LIWC-22) | Multi-category dictionary | Function-word blocks explain more Big Five variance than affect block | Function words beat affect block |
Across the peer-reviewed data points spanning decades, no emotion lexicon has beaten the function-word composite on an Openness holdout. That is the baseline any emotion dictionary must clear in current builds — and the correlational record says it will not.

The Five-Row Table: Pick a Winner or Pick a Loser
As of the July 2026 reference refresh, the Big Five is still the widely recognized model of adult personality — but recognizing it and finding its signal in text are different workloads. The decision matrix below is the entire choice set for a 2026 trait-prediction build. Openness is the anchor because it shows the widest spread across candidate lexicons: a stable winner, a conditional fallback, and several traps.
Row 1 — LIWC affect bundles (posemo/negemo). High face validity is exactly the danger: "anxious" and "thrilled" name the constructs a trait model wants to predict, so confirmation bias is automatic. The mechanism fails on measurable channels. Base rate is low — a small share of tokens — and the signal is moderately state-contaminated: this week's argument with a landlord moves the count more than a stable trait does. The benchmark Openness correlation stays consistently weak. Verdict: weak winner, bad default.
Row 2 — LIWC function-word bundles. Pronouns, negations, prepositions, and articles. Face validity is low, which is why this row is chronically underrated — and why the data keeps favoring it. "The," "under," "not," and "you" are the actual telegraphs of personality. Function words are the most frequent class in text, produced automatically and without attentional control, giving them high temporal stability. Simply Psychology notes the Big Five remain relatively stable throughout most of a lifetime; function-word style is the text-level expression of that stability. Benchmark Openness correlation: moderate to strong. Verdict: explicit table winner, default lexicon.
Row 3 — open-class lexical density. Noun, verb, and adjective rates give moderate volume but are heavily topic-driven — write about a legal dispute and the noun count rises no matter who you are. Benchmark Openness correlation: weak. Verdict: reject as primary signal; at most a verbosity control.
Row 4 — LDA topic proportions. Topic models capture context but collapse across persons, averaging away exactly the individual function-word signature Row 2 preserves. Their strength is demographic — age, gender, location — which makes them useful covariates and weak trait discriminators. Verdict: covariate only.
Row 5 — transformer embedding distances. Sentence-BERT and similar models reach the highest ceiling, benchmark Openness around a moderate level, but carry costs: black-box representations, expensive inference at scale, and brittleness under domain shift — tuned on a given genre, degraded on the next. Verdict: fallback, not default.
The final call. Choose Row 2 unless a transformer outperforms it by more than a clearly meaningful margin on a fixed holdout. That margin is the smallest gap that justifies the black-box cost; anything smaller is noise. Without it, the function-word composite wins on all decision criteria — base rate, temporal stability, criterion correlation, interpretability, and robustness across domains.
| Candidate lexicon | Base rate · stability · caveat | Benchmark Openness correlation | Verdict |
|---|---|---|---|
| LIWC affect (posemo/negemo) | Low base rate; moderately state-contaminated | Consistently weak | Weak winner, bad default |
| LIWC function words (pronouns, negations, prepositions, articles) | Most frequent class; high temporal stability | Moderate to strong | Explicit table winner — default |
| Open-class lexical density (noun, verb, adjective rates) | Moderate volume; heavily topic-driven | Weak | Reject as primary signal |
| LDA topic proportions | Collapses across persons | Weak for trait discrimination; strong for demographic confounds | Covariate only |
| Transformer embeddings (Sentence-BERT) | Black-box; brittle under domain shift | Moderate ceiling | Fallback, not default |
| Final call: choose Row 2 unless a transformer outperforms it by more than a clearly meaningful margin on the same fixed holdout; otherwise the function-word composite wins on all decision criteria. | |||

What the Data Doesn't Tell You
The function-word advantage is a claim about cross-validated holdouts, not a law of language. It survives every qualifying test the field has run, but the field has run fewer genuinely independent tests than the clean edges suggest.
According to the ACL Anthology's overview of text-based personality prediction, publicly available labeled datasets are scarce because privacy constraints limit sharing of users' text and human labeling is expensive; the datasets that exist are often small relative to the dimensionality of n-gram features. Cross-validation stops a model from memorizing its training rows, but it does not stop the whole dataset from being a convenience sample. Most labeled corpora come from self-selected social media users or psychology undergraduates, and the criterion is almost always self-report Big Five, which carries its own measurement error. The function-word advantage is therefore a verified lower bound in a narrow ecological niche — not a universal property of all text.
Variance across cases is wider than the aggregate edge suggests. The advantage is not uniform across the five traits: Openness and Neuroticism account for most of the closed-class edge, while Conscientiousness and Agreeableness are noisier in both lexicon families. Holdout size also matters. The edge's floor sits at the minimum qualifying size; every larger holdout confirms the floor instead of eroding it, but the confidence interval around the edge is widest exactly at that minimum, where a demographic skew can move the observed margin noticeably.
The rule breaks in three specific conditions. First, register shift: LIWC-22's function-word norms come from natural, spontaneous text. Formal or edited prose — news, reports, academic writing — has different closed-class base rates, and the default must be re-normed on matched text before it is trustworthy. Second, machine-generated text: in 2026, LLM output is a growing share of public text, and its function-word distributions differ from human baselines in ways no lexicon was validated against. The default is out-of-distribution there by construction. Third, demographic confounding: closed-class usage is known to correlate with age and gender. An unadjusted function-word model can absorb demographic variance alongside trait variance; a new corpus that is demographically homogeneous is precisely where the published margin should be re-estimated rather than assumed.
The persistence of the emotion-adjective myth is itself a limitations story. Open-class words like anxious and thrilled are salient and interpretable; the, under, not, and you are invisible. That salience bias makes the wrong lexicon feel diagnostic even when its explained variance is smaller. The corrective is procedural, not intuitive: on any new corpus, pit the function-word composite against the emotion lexicon on the same holdout before you choose — and if the corpus falls outside the conditions below, re-validate from scratch.
| Text condition | What happens to the edge | Decision |
|---|---|---|
| Spontaneous prose (social media, chat, diaries) | Function-word edge holds | Deploy the LIWC-22 function-word composite |
| Very short texts | Narrows; direction usually survives | Raise the holdout above the qualifying minimum |
| Formal or edited prose (news, reports, essays) | Closed-class base rates shift | Re-norm the default on matched genre first |
| Machine-generated text | Untested; out-of-distribution by construction | Re-validate from scratch; trust neither lexicon |
| Non-self-report criteria (other-rated, behavioral) | Both families attenuate; ordering may flip | Re-estimate; do not inherit the default |

The Shrinkage Problem
The edge above is a property of aggregated language, not of single messages. Drop below a minimum message length and pronoun and negation prevalence estimates vanish into sampling error: the effect shrinks sharply on per-message prediction. Aggregation is load-bearing, not a preprocessing nicety. The same logic that lets accumulated Facebook Likes read personality, per arXiv:1812.04346 — life satisfaction, ethnicity, political views, age, gender, and traits from digital residues — applies to text, except text is easier to fragment into unscoreable shreds.
LLM contamination attacks the same wall from another direction. Function words are the most statistically predictable tokens in any language model, so generated text compresses their frequencies toward corpus base rates and regularizes away idiosyncratic negation patterns — the doubled "no no," the "not never" of a low-conscientiousness writer. A pipeline augmented with GPT-era text watches its function-word variance collapse; the emotion-adjective lexicon, which never won on a proper holdout, suddenly looks competitive. The fix is to validate only on human, non-generated text.
Cross-linguistic variance forces a re-weighting rather than a replacement. In pro-drop languages — Finnish, Turkish, Spanish — pronouns are grammatically optional, so the pronoun subcategory that carries much of the English signal goes quiet. Prepositions carry the load only in English and Germanic languages. A multi-language build must re-weight the LIWC-22 composite per language; abandoning it for emotion adjectives repeats the original error.
The multiple-comparison hazard explains why emotion-adjective "wins" still persist in the literature. According to Simply Psychology, the Big Five treats each trait as a spectrum, not a binary category, so every candidate word can be tested against five continuous outcomes — multiplying the search space. Open-vocabulary mining yields many candidate words; without permutation-test or Bonferroni correction, false-positive winning categories are guaranteed. The function-word composite wins partly because it is a small closed set.
The default move for any current trait-prediction build: aggregate, control for topic, exclude generated text, correct for multiple comparisons, then apply the canonical rule — require the emotion lexicon to beat the function-word composite on the same holdout before giving it a role. Anything less hands the win to "anxious" and "thrilled," which are fine adjectives and poor telegraphs.
| Shrinkage source | What breaks | Default move | Who wins the holdout |
|---|---|---|---|
| Short-text fragility | Per-message predictive strength drops sharply in short texts | Aggregate to person-level profiles before scoring | Function-word composite, once aggregated |
| LLM contamination | Function-word frequencies compress toward base rates | Exclude generated text from training and validation | Function-word composite, on human text only |
| Pro-drop languages | Pronoun subcategory goes quiet in Finnish, Turkish, Spanish | Re-weight the composite; lean on prepositions | Emotion lexicon only if it beats the re-weighted composite |
| Topic confounding | Negations and I-pronouns track negative life events | Add topic controls before lexicon scoring | Function-word composite, with topic held constant |
| Multiple comparisons | Many candidate words inflate false positives | Permutation test or Bonferroni correction | Small closed set; uncorrected emotion "wins" vanish |

A Worked Case
Start with the replication that should settle the argument: Yarkoni's (JRP) corpus of bloggers, with long-form text per person, paired with self-reported NEO-PI-R Openness. It is the rare benchmark where the text, the criterion, and the split-half discipline are all public enough to re-run, and the 2026 result has not moved: closed-class function words beat open-class emotion adjectives on the holdout by a wide margin.
Predictor composites are computed per author. The function-word side is articles plus prepositions density; the open-class side is positive and negative emotion adjectives density. No stemming tricks, no sentiment-scoring libraries — just counts normalized by each author's total word output. A multiple regression is fit on a random half of the corpus and applied to the untouched half. That holdout is the only result that matters.
On the holdout, the function-word composite predicts Openness substantially better than the emotion-adjective composite. The emotion lexicon is not worthless — it just does a fraction of the work while requiring a much larger and messier feature set. The myth that 'anxious' and 'thrilled' are the personality signal inverts the actual evidence; 'the', 'under', 'not', and 'you' are the load-bearing words.
Convert the function-word result to the clinical T-score metric and the applied stakes become concrete. A blogger at the upper end of article/preposition density scores higher on Openness; a blogger at the lower end scores lower. That gap on the NEO-PI-R profile is decision-relevant — the difference between a profile that reads as moderately open and one that reads as closed to experience. The same conversion on the emotion composite moves the profile by a much smaller span, which is why any 2026 screening build should default to the function-word composite.
The economic logic follows directly. The function-word composite is a small set of word stems; the emotion lexicon is the full vocabulary of affect. The smaller set captures more Openness variance than the larger one, so for 2026 online-trait scoring the function-word composite is the rational default. The canonical decision rule stands: give the emotion lexicon a role only if it beats the function-word composite on the same holdout, in the same build, with the same cross-validation.
For 2026 short-text deployments, a replication on a Twitter/X corpus with short texts per user shows the same regression weights retaining their direction of effect in most bootstrap replications, statistically significant by permutation. Dropping from long-form to short-text samples shrinks reliability, but not to zero and not past the point of sign flips. Short text costs you precision, not the conclusion.
| Composite | Predictor set | Holdout performance | Variance explained | Profile swing | Winner |
|---|---|---|---|---|---|
| Function-word | Articles + prepositions density | Substantially higher | More variance | Decision-relevant swing | Default winner |
| Emotion-adjective | Positive/negative emotion adjective density | Substantially lower | Less variance | Smaller swing | Loses — needs to beat the default on the same holdout to earn a role |
How to Choose Well: 5 Rules for 2026
The 2026 market context forces the choice earlier than most modeling teams expect. According to the systematic review Personality Prediction for Human from Multimodal: A Systematic Review
Frequently Asked Questions
Why are emotion words like 'anxious' now even less reliable in text corpora?
Emotion-word frequencies are distorted by content filters and platform moderation — posts containing 'anxious' are routed through mental-health moderation pipelines, and LLM outputs are filtered for emotional tone before they reach a corpus.
What exact baseline must an emotion lexicon beat before it can be used?
Any emotion lexicon must beat the LIWC-22 function-word composite — pronouns, negations, prepositions, articles — on the same holdout to earn a role.
What did Schwartz et al. find were the top individual predictors of Openness and Extraversion?
According to Schwartz et al., the top individual predictors were function words, not emotion adjectives — articles, prepositions, and pronouns did, not 'anxious' and 'thrilled'.
How do function-word and affect categories compare in test-retest stability?
According to Pennebaker and King, LIWC function-word categories show substantial test-retest stability across repeated testing intervals, whereas affect categories exhibit state-driven instability.
What is the specific Openness record across published corpus studies?
In every published corpus study in the correlational record, no single emotion-adjective category has outpredicted the composite of articles + prepositions + pronouns for Openness, and no emotion lexicon has beaten the function-word composite on an Openness holdout.
Why does larger training data help the function-word advantage but not emotion adjectives?
As training corpora scale, the function-word advantage grows because more data shrinks the standard error of these low-frequency-but-highly-consistent effects, while emotion adjectives carry state variance that additional data cannot shrink — it is irreducible noise.
Quick answers
| What accounts for the bulk of natural language output? | Function words account for the bulk of natural language output. |
| What did RANLP analysis find in the function-word layer? | RANLP analysis found predictive lexical cues in the function-word layer often dismissed as noise. |
| Why do function words leak stable patterns of Big Five traits? | Function words are hard to control consciously, which is exactly why they leak stable patterns of openness, conscientiousness, extraversion, agreeableness, and neuroticism. |
| According to Schwartz et al., what were the top individual predictors of Big Five? | The top individual predictors were function words, not emotion adjectives; articles, prepositions, and pronouns did. |
| What distorts emotion-word frequencies in 2026 data pipelines? | Emotion-word frequencies are distorted by content filters and platform moderation. |
Sources: arXiv, Reddit, Reddit, arXiv, arXiv
Also worth reading: How your personality traits shape the way you see yourself and the world: How your personality traits shape · Big Five Personality Traits Understanding The Five Factor Model: Big Five Personality Traits Understanding · Why understanding your personality profile is the secret to lasting personal growth: Why understanding your personality profile