I was scrolling through Amazon's self-published listings a few weeks back, looking for a quick read for a flight, and I noticed something weird. Dozens of books had nearly identical cover styles, blurbs read like they were generated from the same template, and publication dates were stacked on top of each other like an assembly line. One author — or what I assume is one author — had released eleven books in a single month. Eleven.
I brushed it off as a curiosity. Then I read a paper that dropped on arXiv this week, and it turns out that little observation is a symptom of something much bigger and much more consequential than I thought.
The paper is called "Generative AI Floods and Dilutes the Market for Books." It's the first large-scale empirical study of what AI-generated fiction is actually doing — not to literature or quality, but to the market itself. To the money, to the rankings, and to the human authors trying to earn a living in the same digital shelf space.
The headline finding hit me harder than I expected: AI doesn't need to write good books to hurt human authors. It just needs to write a lot of them.
The Study: What They Actually Did
Let me walk through the methodology, because the scale here is what sets this paper apart from the usual "AI wrote a novel, isn't that neat" discourse.
The research team — Tuhin Chakrabarty from Stony Brook, Jane Ginsburg from Columbia Law, Paramveer Dhillon from Michigan, and Xinyue Liu — assembled a dataset of 14,419 self-published genre-fiction e-books released on Amazon between January 2023 and March 2026. They tracked daily sales through June 2026 using a proprietary sales panel from a major "Big Five" publisher, which covers roughly 95% of daily e-book unit volume on Amazon.
They didn't just skim blurbs or judge books by their covers. They ran full-text AI detection on every single title. They obtained texts by emailing authors for ePubs, borrowing through library apps like Libby and Hoopla, or purchasing them directly from Amazon. Then, they classified each book into three buckets:
- No AI text — zero detected AI-generated content
- Light AI text — some AI content, but 25% or less of the text
- Substantial AI text — more than 25% of the text flagged as AI-generated
They matched all of this to Kindle Unlimited status, genre, author metadata, and daily sales rank data across eight genre clusters: General Romance, Speculative/Adventure Romance, Crime/Suspense Romance, Sports Romance, Fantasy/Supernatural/Horror, Mystery/Thriller/Crime, Science Fiction/Adventure/Dystopian, and General/Contemporary Fiction.
This isn't a vibes-based commentary. This is a market autopsy.
The Flood: AI Books Are Everywhere Now
Here's the first thing that jumped out at me. Books with substantial AI text have spread into every single genre the researchers tracked. The share of new releases containing substantial AI text rose steadily from near zero in early 2023 to roughly 20% of the entire catalog by 2026.
However, there's a nuance that a lot of the "AI slop" discourse misses: these books don't sell as well on average. Books with substantial AI text make up 20% of all books in the sample but earn only 12.1% of sales and 11.3% of revenue. Meanwhile, books with no detected AI text account for 62.9% of books but pull in 71.7% of sales and 72.5% of revenue.
So, the average AI-heavy book underperforms. If you stopped reading the paper here, you might think, "Great, the market self-corrects, readers can tell the difference, and human authors are fine."
You'd be wrong. This is where the paper gets really interesting.
The Dilution: It's Not About Quality. It's About Volume.
The researchers introduce a concept they call dilution, and I think it's the single most important idea in this study.
Here's the math that should make every self-published author uncomfortable:
- The cumulative catalog of released titles grew 38.3 times between Q1 2023 and Q1 2026.
- The number of books actually selling in a given quarter grew 19.2 times.
- Quarterly revenue grew only 8.9 times.
Read that again. The number of books competing for reader attention grew nearly four times faster than the money available to pay for them.
Revenue per selling book fell, and it didn't just fall for AI books. It fell for human-written ones, too.
This is the critical finding, and the authors go out of their way to prove it isn't a statistical illusion. The worry would be a "mix shift" — where average revenue drops only because a bunch of low-earning AI books were added to the denominator, while human books do fine individually. The researchers ran a shift-share decomposition and showed that's not what's happening. Revenue per selling title for books with no AI text fell from $23,877 in the 2023 cohort to $19,739 in the 2025 cohort. That's a 17.3% drop for books that don't contain a single word of AI text.
The decline showed up in seven out of eight genres for no-AI-text books. The one exception? Fantasy/Supernatural/Horror — the genre AI reached last and least. In that genre, revenue per book for no-AI-text titles actually rose 35%.
That contrast is hard to explain away as a general market trend. When the one genre AI hasn't flooded yet is the only genre where human authors are still earning more, the pattern points in a clear direction.
The Top Ranks Are Shifting Too
I'll admit, I assumed AI books would stay stuck at the bottom of the market. The "slop" framing suggests they'd just pile up harmlessly in the digital equivalent of a bargain bin.
The data says otherwise.
The share of constructed Top-25 rank slots held by books with substantial AI text grew from near zero in early 2023 to roughly 31% of new Top-25 entrants by 2026. Combined with light AI text, books containing AI content reached about 40% of top-rank slots by mid-2026.
The top of the market is turning over faster. Retention of no-AI-text books in the Top 25 fell to a trough near 28% before recovering unevenly to about 62%. The top ranks aren't a fortress anymore; they're contested ground.
Losses concentrate exactly where you'd expect: in genres with the highest AI exposure, the share of top slots held by no-AI-text books drops from about 88% in the least-exposed genres to about 63% in the most-exposed ones.
Kindle Unlimited Makes It Worse
This part surprised me. The researchers found that the dilution effect is significantly amplified in genres where more titles are available on Kindle Unlimited.
The logic makes sense once you break it down. Kindle Unlimited operates as a subscription pool. Readers don't pay per book; they borrow from a shared catalog. Every new AI-generated book that enters that pool competes for the same finite number of borrows. The marginal cost to the reader of trying another book is zero, so they are more likely to grab an AI title that looks "good enough" rather than hunt for a human-written option.
In high-KU genres, the lead that no-AI-text books hold over substantial-AI-text books is 8.5 percentage points smaller in sales share and 8.4 points smaller in revenue share compared to low-KU genres.
If you're a self-published author relying on Kindle Unlimited for income, this should be on your radar.
The Authors Who Adopted AI Didn't Just Write a Little Faster
The supply-side findings tell their own story. The researchers tracked 824 author bylines that released at least one book with substantial AI text. Of the 385 who went on to release more AI-heavy books after their first, 287 (74.5%) increased their monthly output after adoption.
This wasn't confined to a few power users. The expansion was broad, though uneven. The top quartile of producers accounted for 60.5% of post-adoption AI books (Gini coefficient of 0.48). While there is concentration, it isn't driven by a handful of bots. It's a wide base of people who found a tool that allowed them to publish faster and kept using it.
At the top end, the numbers are extreme. The most prolific author identities were releasing dozens of books per year, almost entirely containing substantial AI text. The top author identity generated $1.7 million in gross consumer revenue across eight titles. The single highest-earning AI-heavy book pulled in $643,000 on 80,431 sales.
While these are gross consumer spending numbers rather than author royalties after Amazon's cut, they represent substantial commercial figures. It is real money that, in a pre-AI market, would have gone to writers spending months crafting books by hand.
The Textual Fingerprint: AI Books That Succeed Lean on Existing Books
This finding is particularly fascinating, and it carries significant implications for copyright law.
The researchers measured how much of a book's text is covered by rare expressions — phrases of five or more words that appear in at most five volumes in the Google Books index and are completely absent from a 4.7-trillion-token snapshot of the internet. These aren't common idioms; they represent distinctive language from specific existing books.
Among top-selling books, those with substantial AI text were saturated with these rare expressions. At the Top-50 tier, mean coverage hit 45.0% for substantial-AI books versus 37.7% for no-AI-text books — a gap of 7.2 percentage points. The gap held at Top-100 (5.5 points) and Top-200 (4.3 points), both statistically significant ($p < 10^{-8}$).
Here is the key detail: for substantial-AI books, this overlap rises with revenue — 7.6 percentage points per tenfold increase in revenue. For no-AI-text books, the slope is essentially flat (1.1 points, $p = 0.67$). The more money an AI-heavy book makes, the more it draws on the distinctive language of existing books.
To put this in perspective, the researchers compared both groups against 200 literary award-winning books — Booker Prize, Pulitzer, National Book Award, PEN/Faulkner, and Nobel nominees. Award-winning fiction had a mean rare-expression coverage of just 19.1%, roughly half the rate of the AI books.
The cited examples are telling: "cheekbones more pronounced, the shadows beneath her eyes deeper", "The question circled his mind like a vulture", and "my stomach flip-flopping like a fish out of water". The researchers note this aligns with recent findings that AI tends to overwrite physical sensations to convey emotion rather than naming feelings directly.
This measure captures aggregate textual overlap, not direct copying from any single specific book. It aligns with a derivative-production model, though it does not prove a given passage was directly lifted. Still, the pattern is clear: the AI books earning the most money rely heavily on language that resembles existing published works.
Why This Matters for Copyright Law
The paper addresses these legal implications directly.
In Kadrey v. Meta, a major lawsuit concerning the training of large language models on copyrighted books, Judge Chhabria specifically questioned whether AI-generated outputs cause market dilution by flooding the market with cheap substitutes. He noted that plaintiffs had not provided empirical evidence for this effect.
This paper offers empirical data addressing that exact market dynamic.
The researchers note that liability for copyright infringement, if established, stems primarily from upstream copying — the mass ingestion of copyrighted books for training — rather than the market competition generated by AI outputs. Competition itself is not infringement. However, the market-dilution theory demonstrates how that copying facilitates scalable competition with measurable economic impacts.
The paper also highlights a transparency issue: none of the 14,419 books disclosed whether they contained AI-generated content. Readers cannot easily distinguish AI text from human text based on preview materials prior to purchase. The absence of disclosure may remove potential market disadvantages for these titles, though the authors note further data is needed to evaluate this directly.
What This Means Beyond Self-Published Romance Novels
The conditions present in this segment of the publishing market exist across several creative fields:
- Low entry barriers — minimal restrictions on publishing volume
- Pattern-driven demand — audiences seeking established tropes
- Algorithmic discovery — human and AI content competing in the same recommendation feeds
- Unlabeled production methods — consumers cannot readily distinguish origins before purchasing
Music, stock photography, digital art, and short-form media share these characteristics. As the paper points out, an expanding supply of automated output competing for a fixed reserve of human attention represents a fundamental economic shift across digital creative markets.
AI-generated books do not need to outperform human writing on quality to alter market dynamics. By occupying rank positions, recommendation spots, and subscription reads through sheer volume, they capture market share that previously went to human creators.
Final Thoughts
The paper summarizes this phenomenon clearly:
"Generative AI can thus reshape a creative market through scale rather than quality."
Public discussion frequently centers on whether AI can write high-quality fiction or pass as human. Meanwhile, market structures are shifting primarily due to volume.
While individual AI-heavy titles earn less on average, their high output volume reduces relative returns across the market. Authors spending extended periods on individual works face increased competition from high-volume releases occupying visibility slots.
Furthermore, the higher-earning AI titles show significant overlap with phrases found in existing published works, highlighting ongoing questions around training sources and market impacts.
While regulatory and platform responses continue to evolve, treating automated publishing as commercially negligible is no longer supported by market data. The shift in volume is already influencing creative distribution.
Citation
Chakrabarty, Tuhin, et al. "Generative AI Floods and Dilutes the Market for Books." arXiv, 22 July 2026, arxiv.org/abs/2607.20349.