# Story Ideas for Writers: Up to 8.1% Higher Novelty Ratings—Scaffold Selectively

Gavin Marshall · September 19, 2026

> Explore what AI idea-diversity research actually shows, why the 8.1% story novelty claim is unverified, and how writers can use scaffolds selectively.

| Takeaway | Detail |
| --- | --- |
| Treat the headline percentage as unverified. | The supplied excerpts do not establish an 8.1% improvement in reader-rated story novelty or identify a fiction experiment supporting it. |
| Keep the experimental task in view. | “Prompting Diverse Ideas: Increasing AI Idea Variance” studied product ideas for college students priced under $50, not stories or novels. |
| Separate idea diversity from reader-rated novelty. | In the under-$50 product task, researchers evaluated cosine similarity, unique ideas, and idea-space exhaustion; those measures do not establish reader judgments of fiction. |
| Scaffold selectively without promising originality. | Even a verified 8.1% average increase in novelty ratings would not demonstrate that a particular novel is unprecedented; selective scaffolding remains a practical recommendation, not a tested fiction outcome here. |

Under $50: that is the product-price constraint in the supplied research abstract, not a measure of literary originality. The headline’s 8.1% novelty increase cannot be verified from these excerpts. “Prompting Diverse Ideas: Increasing AI Idea Variance,” by Meincke, Mollick, and Terwiesch, examines AI-generated product ideas for college students—not a story-writing experiment. That distinction changes what a fiction writer can responsibly take away.

From a psychometric perspective, the central question is what the measurement licenses you to conclude. The abstract reports that plausible prompts produced AI idea pools less diverse than human-generated pools, while prompt engineering could substantially improve AI idea diversity. Its measures concern similarity, unique ideas, and exhaustion of the idea space. They do not establish improved reader-rated story novelty. Even a verified average ratings gain would describe responses under study conditions, not certify that your next novel is unprecedented.

For writers, selective scaffolding is therefore a cautious workflow recommendation rather than an experimentally validated fiction strategy. Use generated suggestions as options to challenge, compare, and reshape; retain control over voice, character, and narrative decisions. Those safeguards may help structure your process, but this source set did not test their effects on fiction.

![Story Ideas for Writers](https://static.mm-ais.com/article-images-ai/story-ideas-for-writers-up-to-8-1-higher-ai-869e6b4e.jpg)

## GPT-4 as an Idea Scaffold

GPT-4’s role was to change the writer’s starting conditions, not to demonstrate a change in the writer’s underlying creativity. That distinction matters: supplying a premise can remove a bottleneck without strengthening every capability needed to produce original fiction. An assisted story can therefore earn favorable reader judgments while leaving unanswered whether its author would perform differently on a subsequent, unaided task.

The experiment used a three-arm randomized design: an unaided condition and conditions offering different levels of access to AI-generated starting ideas. GPT-4 supplied those narrative seeds; participants remained responsible for writing the stories. Random assignment made access to ideation assistance the experimental contrast, rather than simply comparing writers who independently chose AI with writers who avoided it. The intervention was assistance before composition, not automated story production.

An external premise can reduce the search burden by supplying a setting, a conflict, or an unexpected causal connection that the writer would otherwise need to discover. Consider an illustrative GPT-4 seed about an archivist finding records of events that have not happened. The suggestion establishes a searchable direction: anticipation, institutional secrecy, or unreliable evidence. It does not establish why this archivist investigates, what an investigation costs, or which discovery could justify the ending. Those unresolved decisions are substantial parts of authorship, not cosmetic additions to an already completed story.

This separates divergent search from elaboration. Divergent search opens possible directions; elaboration develops a direction into motivated action and connected consequences. A supply of suggestions need not itself constitute broad exploration if the suggestions repeatedly occupy similar conceptual territory. Nor does an intriguing seed validate its execution. According to “Collective Intelligence Is Not Brainstorming,” generating possibilities and validating them are distinct activities: brainstorming “does not reliably validate” ideas. In fiction, character motivation, scene-level causality, and a coherent ending remain problems the writer must solve.

Psychometrically, a reader’s novelty judgment is an observed rating, whereas an author’s creative capability is a latent construct—something inferred through measurement rather than directly inspected. The rating concerns a particular finished artifact under particular task and evaluation conditions. It reflects the interaction of writer, available assistance, execution, and evaluator judgment; it is not a direct quantity of creativity residing inside the author. Higher ratings for assisted work therefore do not establish that AI reliably makes an individual writer more original. Predicted ratings provide no shortcut around that distinction.

Use AI as an optional idea scaffold when stuck, then introduce an independent decision stage before drafting. Set the suggestion aside and rebuild the premise around a motivation, causal dependency, and ending you can justify yourself. For the archivist premise, decide why acting on a prediction changes what can happen next, rather than merely renaming the protagonist. Then write the prose yourself. This reconstruction is a practical safeguard for human authorship—not an experimentally established mechanism of improvement, and not a guarantee of originality.

![GPT-4 as an Idea Scaffold — Story Ideas for Writers](https://static.mm-ais.com/article-images-ai/story-ideas-for-writers-up-to-8-1-higher-ai-d52198d7.jpg)

## The 8.1% Result

The 8.1% figure is not a 2026 result. It comes from Anil R. Doshi and Oliver P. Hauser's 2024 paper in *Science Advances*, "Generative AI enhances individual creativity but reduces the collective diversity of novel content" (DOI 10.1126/sciadv.adn5290). Any guide presenting it as a fresh experiment is misdating a two-year-old study — and, more consequentially, misreading what kind of number it is.

Doshi and Hauser analyzed stories written by participants, then had those stories rated by a separate sample of people. That separation is the design's most underrated feature. The people who wrote the stories never rated them; the people who rated them never wrote them. In psychometric terms, producer and rater are disjoint samples, which strips out the self-enhancement bias that inflates most self-reported creativity measures. Every novelty number from this study is an aggregate of third-party judgments, not a writer's assessment of their own work.

Against an unassisted control, the one-idea condition drew novelty ratings roughly 5.4% higher on average. The five-idea condition drew ratings roughly 8.1% higher. Both are relative differences in mean ratings: the control's mean is the denominator. They are not percentage-point gains, and they are not absolute movements on a rating scale. A 5.4% relative lift on a mean rating of 5.0 and a 5.4% relative lift on a mean of 7.0 describe different absolute shifts; the paper reports the relative form, and that form is what you should quote.

"Up to 8.1%" is a condition-level summary. It describes the best-performing arm of the reported design, averaged across the writers assigned to it. It does not mean any individual writer who used five AI ideas scored 8.1% higher, and it does not mean the effect reproduces at that magnitude in a different population, prompt, or genre. Condition means and individual outcomes are different objects. Conflating them is the error that converts a modest group-level shift into a false promise of personal improvement — the exact myth this guide exists to kill.

Novelty ratings are also not a currency. They do not convert into publication probability, sales, or an equivalent percentage increase in writing talent. A rater's novelty judgment is one aggregated signal, not a latent trait score, and nothing in this line of work establishes a mapping from rating points to career outcomes. Treat the percentages as evidence about a distribution, not about you.

| Condition | Comparison | Reported difference | What the number actually is |
| --- | --- | --- | --- |
| Unassisted control | Reference group | 0% (baseline) | Mean novelty rating used as the denominator |
| One AI idea | vs. unassisted control | ~5.4% higher average rating | Relative rating difference, not percentage points |
| Five AI ideas | vs. unassisted control | ~8.1% higher average rating | Condition-level "up to," not an individual guarantee |
| Story producers | Participants | Wrote the stories | Never rated their own output |
| Story evaluators | Separate sample of people | Rated the stories | Disjoint from the writers; third-party judgment |

If you cite these numbers anywhere, cite them with the control attached. A percentage without its comparison group is not a finding — it is a headline. And when you sit down to write, the honest reading of Doshi and Hauser is narrow: an optional idea scaffold shifted average third-party novelty ratings upward in one controlled setting. It says nothing about whether the next story you write will be more original than the last.

![The 8.1% Result — Story Ideas for Writers](https://static.mm-ais.com/article-images-pixabay/story-ideas-for-writers-up-to-8-1-higher-8d3784cb.jpg)

## Choose Selective Scaffolding

Selective scaffolding wins only when premise search is the bottleneck. This is a comparison of workflows for a fiction writer who cannot identify a workable premise—not a ranking of generators or a reason to add assistance to every writing session. The winner below is an editorial recommendation, not a separately tested experimental treatment.

A usable premise means the writer can name a protagonist, a consequential desire, and an obstacle capable of generating events. A setting or mood alone does not qualify. “Mara tends an abandoned lighthouse” supplies atmosphere; “Mara must keep the lighthouse operating to guide her missing sister home, while the town dismantles its power supply” supplies a desire and an obstacle that can produce choices and consequences. When those elements already exist, additional idea generation lacks an established need.

| Workflow | Best-fit condition | Main trade-off | Decision |
| --- | --- | --- | --- |
| Independent ideation | A usable premise already exists | No external starting-point assistance | Prefer when unstuck |
| Unfiltered adoption | Speed matters more than independent premise selection | The first suggestion can become the default story | Do not make this the standard workflow |
| Selective scaffolding | Premise search has stalled | Requires additional human reconstruction | WINNER for a blocked writer |

Selective scaffolding addresses the missing starting point while reserving the governing conflict and resolution for the writer. Treat a suggestion as material to interrogate: whose desire matters, why does resistance force action, and what outcome should the story earn? Independently rebuilding those relationships is different from renaming the suggested protagonist while retaining the generator’s entire causal structure. Then write the story yourself.

According to Medium’s “Using AI for Idea Generation Without Losing Your Voice,” “Guided Prompting” is its first technique for using AI tools without losing creative control. That offers a practical vocabulary for requesting bounded assistance, not experimental validation of this workflow. For the lighthouse premise, an optional request could target an unresolved obstacle rather than solicit a complete plot and ending.

The measurement distinction is between an evaluated output and an inference about its author. Higher predicted novelty ratings are not proof of originality, nor do higher ratings establish that AI reliably makes an individual writer more original. Selective scaffolding is preferable here because it fits the stated bottleneck—not because independent reconstruction has been shown to outperform every other AI-assisted method.

Apply this short decision-tree in order; stop seeking ideas once the premise is usable.

| Condition | Option and decision |
| --- | --- |
| You can name the protagonist, consequential desire, and event-generating obstacle. | Choose independent ideation and draft; no missing starting point justifies assistance. |
| An element is missing, but you can still identify workable alternatives yourself. | Continue independent ideation; an unfinished premise is not necessarily a stalled search. |
| An element is missing and premise search has stalled. | Optionally choose selective scaffolding; request material addressing that gap rather than a finished story. |
| A suggestion is available, but its conflict and ending remain your defaults. | Reject unfiltered adoption as the standard; independently reconstruct the premise before drafting. |
| The rebuilt premise is usable, but predicted ratings favor another suggestion. | Choose independent drafting; ratings do not establish a need to replace your premise or prove originality. |

![filmmakers youtuber script screenwriter writing storyboard story boarding film producer from above in the workplace lapt](https://static.mm-ais.com/article-images-pixabay/story-ideas-for-writers-up-to-8-1-higher-1ed92863.jpg)
filmmakers youtuber script screenwriter writing storyboard story boarding film producer from above in the workplace lapt

## What the Data Doesn't Tell You

Better-rated stories can still make a less diverse collection. According to Anil R. Doshi and Oliver P. Hauser’s Science Advances paper, “Generative AI enhances individual creativity but reduces the collective diversity of novel content,” AI-assisted stories became more similar to one another. Individual evaluation and collective diversity therefore answer different questions: readers may favor a particular story while the group produces a narrower range of stories overall. Shared suggestions could channel writers toward overlapping narrative territory, although the similarity finding alone does not establish that mechanism. A higher individual rating cannot demonstrate that assistance expands the group’s creative range.

The experiment’s task boundary is also an evidence boundary. Brief, constrained fiction does not establish benefits for novel-length plotting, sustained character development, revision over months, or professional editorial selection. These are not simply longer versions of the measured task: they introduce dependencies across scenes, accumulated commitments, and judgments made under different selection criteria. For example, an appealing opening premise might create a conflict that cannot sustain a novel’s character arc. That is a transfer question, not a demonstrated failure of assistance. Treat each untested demand as a separate claim requiring evidence rather than extending the short-fiction result by analogy.

Differences across writers require equally careful interpretation. According to Doshi and Hauser, benefits were greater among participants with lower baseline creativity scores. The psychometric distinction is between performance on an operational measure and an enduring attribute of a person. A task-specific score samples behavior under particular conditions; it is not a validated personality diagnosis or a permanent “uncreative writer” classification. Measurement error and limited coverage of the creativity construct constrain what the score can mean. The defensible inference concerns variation in observed benefits, not a rule for sorting writers into fixed categories—or a promise that assistance reliably makes any particular writer more original.

Perceived novelty is not historical uniqueness. Reader ratings indicate how fresh a story seems to those readers under the evaluation conditions; they do not establish that its premise is absent from published fiction. Textual similarity analysis addresses another bounded question: resemblance captured by the chosen representation and comparison set. Low similarity cannot certify an unprecedented premise, and high similarity does not by itself establish infringement. Neither result is an exhaustive originality or infringement test. When assessing a premise, distinguish “these readers found it novel” from “a specified comparison found little resemblance” and from the much stronger, generally unsupported claim “this has never been done.”

For this guide’s present-day use, temporal transfer remains unresolved. Changes in model behavior, reader familiarity with generated fiction, and writing practices could alter the observed results. Without a directly comparable replication, neither persistence nor disappearance of the effect should be assumed. When evaluating newer evidence, check whether the writing constraints, assistance procedure, participant characteristics, and rating process remain comparable; a newer model alone does not supply that comparison. Optional ideation assistance remains a bounded choice, not a guarantee: independently rebuilding a suggested premise does not turn a higher predicted rating into proof of originality.

![What the Data Doesn&#039;t Tell You — Story Ideas for Writers](https://static.mm-ais.com/article-images-pixabay/story-ideas-for-writers-up-to-8-1-higher-54185ab2.jpg)

## An 8-Sentence Jungle Story

“A guide loses the expedition’s map” is an invented writer’s stalled fragment; “A plant reveals a hidden route” is an invented AI-style seed. The seed supplies an event, not yet a reason to care about this particular guide’s decision. The writer, the character Mara, the seed, and the finished draft below are teaching material, not actual participant data. The eight-sentence jungle adventure matches a story length and topic used in Anil R. Doshi and Oliver P. Hauser’s research in Science Advances.

The useful distinction here is between adding an incident and reconstructing its significance. A revealing plant can solve the logistical problem almost immediately: the expedition follows the route and escapes. But that sequence leaves the writer’s chosen emotional stakes unspecified. Who would hesitate to take the route, and what would that hesitation cost? Those questions expose what the seed has not supplied without assuming that an uncomplicated escape story must be unoriginal.

In this illustration, the writer independently rebuilds the premise by making Mara a former poacher. The revealed route leads to a hidden nesting ground she once exploited and now protects. Rescue therefore threatens to repeat her earlier betrayal: bringing rescuers directly through the habitat would expose it, while avoiding it requires a harder journey with an injured companion. The moral conflict is the writer’s contribution here. The plant remains, but its narrative function changes from convenient solution to temptation.

Mara lost the expedition’s map when the river took her pack. At dusk, luminous vines revealed a path beneath the canopy. She recognized the marks she had cut there years earlier while trapping rare birds. Beyond them lay a nesting ground she had sworn never to expose. Her injured companion asked her to signal the rescue helicopter from its clearing. Instead, Mara fashioned a splint and led him toward a bare ridge. By sunrise, he understood why she had accepted the harder climb. When the helicopter arrived, she gave its pilot their position and kept the nesting ground off his map.

To examine the reconstruction, remove Mara’s poaching history mentally. The luminous path still works as an adventure device, but her refusal loses its specific connection to guilt and restitution. This counterfactual check identifies what the writer added; it does not certify originality. For a stalled draft, the concrete next move is to rebuild the seed around a consequential choice, then write the story independently—not merely decorate the suggested event.

According to Doshi and Hauser’s Science Advances paper, access to multiple suggestions produced a separate 9.0% increase in usefulness ratings. For arithmetic illustration only, an explicitly hypothetical control mean of 5.00 would become 5.45 under that relative increase: 5.00 × 1.09 = 5.45. Neither mean is an observed study value. This calculation cannot predict the illustrative jungle story’s score, and usefulness is not interchangeable with novelty. Neither a higher predicted rating nor this reconstruction demonstrates that AI reliably makes an individual writer more original.

![An 8-Sentence Jungle Story — Story Ideas for Writers](https://static.mm-ais.com/article-images-pixabay/story-ideas-for-writers-up-to-8-1-higher-898c14dc.jpg)

## How to Choose Well

The decision is not whether to use a generator. It is which deficit you are actually trying to close. In psychometric terms, an intervention only moves the construct it targets: premise generation is one narrow facet of fiction writing, and a between-subjects novelty gain — the kind Doshi and Hauser report in *Science Advances* — tells you about the average effect of that intervention on that facet, not about your bottleneck. If your blockage is prosody, research, or revision, an ideation result is a category error dressed as evidence. Diagnose first, then decide.

**Rule 1 — Locate the blockage.** Before opening any generator, name the failure in one sentence. If you can state the premise but cannot make the prose move, the deficit is executional, not ideational. If you know the scene but not the facts inside it, the deficit is research. If the draft exists but sags, the deficit is structural. Ideation assistance addresses none of these, and a higher novelty rating in a controlled comparison is not a license to apply it anyway.

**Rule 2 — Protect nonnegotiable material.** If the premise depends on an identifiable person's confidential experience, do not enter those details into an external generator. Abstract the situation — shift the setting, the relationship, the occupation — until the person is no longer recoverable, or work unassisted. This is a consent constraint, not a quality tradeoff, and no novelty gain offsets it.

**Rule 3 — Reject premise substitution.** Ask whether a suggestion helps you explore the theme you intended or replaces it. A seed that is more immediately marketable but points at a different question is a substitution, and marketability is a separate construct from the one you set out to measure. Discard it.

**Rule 4 — Require causal understanding.** Use a proposed idea only if you can explain, unaided, why its major events follow from the characters' decisions. If you must ask the generator to justify its own premise, you do not own the causal chain, and the draft will stall where the chain breaks — usually at the midpoint. Return to independent premise development.

**Rule 5 — Apply a stopping condition.** Once an accepted idea can support a scene you genuinely want to write, stop requesting alternatives and draft it. Treat this as a bounded workflow rule, not a research-derived optimal stopping threshold; the published data do not identify such a threshold, and pretending otherwise is the myth in miniature — that a group-level rating gain reliably transfers to your individual originality.

| Blockage | Diagnostic question | Correct remedy | Does ideation help? |
| --- | --- | --- | --- |
| Premise search | Can you state the story in one sentence? | Optional scaffold, then rebuild | Yes, conditionally |
| Sentence rhythm | Does the premise exist but the prose stall? | Line-level revision, read aloud | No |
| Factual research | Do you lack domain facts, not ideas? | Primary sources, named experts | No |
| Revision and structure | Does the draft sag after the midpoint? | Outline audit, causal re-threading | No |
| Confidential material | Is a real person recoverable from the details? | Abstract the situation or stay unassisted | No |
| Theme substitution | Does the seed replace your intended question? | Discard the seed | No |
| Causal gap | Can you explain the events without the generator? | Independent premise development | No |

Concrete next action: write your one-sentence premise and the opening scene before requesting a second batch of alternatives. If the scene moves, the scaffold did its job and the workflow ends there.

## What to do next

| Step | Action | Why it matters |  |
| --- | --- | --- | --- |
| 1 | When you are stuck on a premise, prompt GPT-4 for several distinct starting ideas and treat the output strictly as an optional scaffold — not as a draft, outline, or finished premise. | The three-arm randomized design in the supplied research used GPT-4 to change the writer's starting conditions, not to demonstrate a change in the writer's underlying creativity. |  |
| 2 | Before repeating the headline's 8.1% figure, open "Prompting Diverse Ideas: Increasing AI Idea Variance" (Meincke, Mollick, Terwiesch) and confirm its actual task: product ideas for college students priced under $50. | The excerpts do not establish an 8.1% improvement in reader-rated story novelty, an Frequently Asked Questions How did getting one AI idea compare with getting five? In Doshi and Hauser’s 2024 study, the one-idea condition received roughly 5.4% higher average novelty ratings than the unassisted control, compared with roughly 8.1% higher for the five-idea condition. Does the 8.1% figure mean an increase of 8.1 percentage points? The 8.1% figure is a relative difference in mean novelty ratings using the unassisted control’s mean as the denominator, not a percentage-point gain. Can I expect my own story to score 8.1% higher if I use five AI ideas? The roughly 8.1% increase is a condition-level average, not an individual guarantee or evidence that the same magnitude will reproduce in another population, prompt, or genre. Did the writers rate their own stories for novelty? A separate sample of people rated the stories, and the story writers never rated their own output. Did GPT-4 write the stories or just suggest starting points? GPT-4 supplied narrative starting ideas before composition, while participants remained responsible for writing the stories. Do higher novelty ratings mean better publication or sales prospects? Nothing in this line of work establishes a mapping from novelty ratings to publication probability, sales, or an equivalent percentage increase in writing talent. Quick answers Can the headline’s 8.1% novelty increase be verified from the supplied excerpts? | The headline’s 8.1% novelty increase cannot be verified from these excerpts. |
| What did “Prompting Diverse Ideas: Increasing AI Idea Variance” study? | “Prompting Diverse Ideas: Increasing AI Idea Variance” studied product ideas for college students priced under $50, not stories or novels. |  |  |
| What measures were used in the under-$50 product task? | In the under-$50 product task, researchers evaluated cosine similarity, unique ideas, and idea-space exhaustion; those measures do not establish reader judgments of fiction. |  |  |
| How should writers use generated suggestions? | Use generated suggestions as options to challenge, compare, and reshape; retain control over voice, character, and narrative decisions. |  |  |
| What should writers do before drafting from an AI-generated premise? | Set the suggestion aside and rebuild the premise around a motivation, causal dependency, and ending you can justify yourself. |  |  |

### Related reading

- [Personality test accuracy: 6% vs 0.84 retest for 2026 hiring](https://psychprofile.io/blog/personality-test-accuracy-6-vs-084-retest-for-2026-hiring.php)
- [Conscientiousness predicts job performance: 2026 60-item rho .27 vs .36](https://psychprofile.io/blog/conscientiousness-predicts-job-performance-2026-60-item-rho-27-vs-36.php)
- [Big Five hiring test: 92% vs 71% finish rate on mobile screens](https://psychprofile.io/blog/big-five-hiring-test-92-vs-71-finish-rate-on-mobile-screens.php)
- [Affect-to-Spend Circuit: Why 0.19 vs 0.12 R² Fails to Hold](https://psychprofile.io/blog/affect-to-spend-circuit-why-019-vs-012-r-fails-to-hold.php)
- [BFI-2 Conscientiousness r=.22: Hiring Cutoff vs Feedback](https://psychprofile.io/blog/bfi-2-conscientiousness-r22-hiring-cutoff-vs-feedback.php)
- [HEXACO Retest Reliability: ICC Limits, Kurtz & Lee, N=500](https://psychprofile.io/blog/hexaco-retest-reliability-icc-limits-kurtz-lee-n500.php)

### Latest

- [Personality test accuracy: 6% vs 0.84 retest for 2026 hiring](https://psychprofile.io/blog/personality-test-accuracy-6-vs-084-retest-for-2026-hiring.php)
- [Hiring personality test scores: Big Five Inventory (BFI-2) .86 Replace Sum...](https://psychprofile.io/blog/hiring-personality-test-scores-big-five-inventory-bfi-2-86-replace-sum-scores.php)
- [Conscientiousness predicts job performance: 2026 60-item rho .27 vs .36](https://psychprofile.io/blog/conscientiousness-predicts-job-performance-2026-60-item-rho-27-vs-36.php)

Canonical: https://psychprofile.io/blog/story-ideas-for-writers-up-to-81-higher-novelty-ratingsscaffold-selectively.php
Markdown: https://psychprofile.io/blog/story-ideas-for-writers-up-to-81-higher-novelty-ratingsscaffold-selectively.php/index.md
