AI-generated books are flooding Amazon and tanking sales for human authors

In this articleBig catalog presence, below-average salesRevenue per book is falling even for titles with no detected AI textAI books are breaking…

By Vane August 15, 2026 6 min read
AI-generated books are flooding Amazon and tanking sales for human authors


Analysis of 14,419 Amazon titles reveals AI-generated books are displacing human authors through sheer volume

Revenue per book is falling across the platform, even for titles where no AI text was detected. A study of over 14,000 self-published e-books confirms that AI-generated titles are crowding out human writers, not by quality, but by numbers.

Researchers examined 14,419 randomly selected self-published e-books released between January 2023 and March 2026. For each title, they pulled daily sales figures from an internal dataset maintained by one of the five major US publishers. That dataset tracks about 500,000 Amazon titles and covers roughly 95 percent of all e-books sold daily on the platform, according to the researchers.

Unlike earlier studies that tried to detect AI text from short book previews alone, the researchers classified each book based on its full text. They used the Pangram v3.3 detector, whose developers report a false-positive rate of 0.04 percent. Pangram 4 has since been released. Books were sorted into three bands based on the share of text flagged as AI-generated: none, light (up to 25 percent), and substantial (over 25 percent).

Big catalog presence, below-average sales

Books with substantial AI content make up 20 percent of the catalog studied but account for only 12.1 percent of sales and 11.3 percent of revenue. Books with no detected AI text represent 62.9 percent of the catalog and generate 72.5 percent of revenue. At first glance, that seems to confirm the view that AI books remain low-quality “slop” stuck at the bottom of the market.

The study shows that this view misses the real market dynamics, though. Between Q1 2023 and Q1 2026, the cumulative catalog grew 38.3x while the number of titles selling per quarter grew 19.2x. Quarterly revenue only grew 8.9x. Far more books are now competing for a revenue pool that’s growing much more slowly.

Revenue per book is falling even for titles with no detected AI text

In six of eight genres, revenue per book dropped when comparing titles released in 2023 and 2025 over the same post-release window. Looking only at books with no detected AI text, revenue fell in seven of eight genres.

That rules out the explanation that the average is dropping simply because poorly selling AI books are padding the catalog. The authors call this effect “dilution” but stress that their comparisons are observational and associational, not experimental proof of causation.

The only exception is Fantasy/Supernatural/Horror, where AI text arrived latest and gained the least traction. There, revenue per book for titles with no detected AI text rose 35 percent. The researchers say this reversal argues against a general market trend as the cause.

The patterns are stronger in genres with high Kindle Unlimited availability, where readers draw from a shared subscription pool. In those genres, the revenue-share lead of books with no detected AI text is 8.4 percentage points smaller than in genres with low Kindle Unlimited availability. The researchers attribute the gap to genre-specific traits and don’t draw a causal link to Kindle Unlimited itself.

AI books are breaking into the top ranks

The share of new Top 25 entries with substantial AI content rose from near zero to 31 percent over the study period. The top ranks also turned over faster. The share of books with no detected AI text that stayed in the Top 25 from one quarter to the next dropped as low as about 28 percent at one point before settling around 62 percent by the end of the study.

Output is concentrated among a handful of prolific producers. Of 385 author identities that published more titles with substantial AI content after their first AI book, 287 increased their monthly output afterward. The highest-grossing pseudonym earned $1.7 million in gross revenue before platform fees across eight titles. The single highest-grossing book with substantial AI content brought in $643,000 on 80,431 copies sold.

The study echoes the New York Times report on “Coral Hart,” who reportedly published over 200 romance titles under 21 pen names in a single year and sold about 50,000 copies. AI spam extends beyond books, too. One man scammed millions of dollars through streaming platforms using AI-generated songs.

Top-selling AI books overlap heavily with rare language from existing works

To measure how much successful AI books overlap with rare language from existing works, the researchers used the Allen Institute for AI’s infini-gram tool together with the Google Books index.

They identified rare expressions that appear in five or fewer Google Books volumes and are completely absent from a 4.7-trillion-token web snapshot. Phrases like these point to language patterns closely tied to published books.

Among the 50 highest-grossing titles with substantial AI content, these rare expressions covered 45 percent of the text, compared to 37.7 percent for the top 50 books with no detected AI text. For award-winning or award-nominated fiction, the figure was just 19.1 percent.

Within AI books, overlap rose 7.6 percentage points for every tenfold increase in revenue. No such correlation existed for books with no detected AI text. The method doesn’t trace the origin of individual passages or prove that any specific book was copied. It measures aggregate language overlap.

Chakrabarty told Ai2 that an AI detector only returns “an estimate—a score for how likely a passage is to be synthetic,” without pointing to where the language came from. If a suspect text also contains rare expressions that don’t appear on the web and show up in only a handful of books, one “can say with some confidence that it was taken from books.” That kind of evidence “acts as circumstantial evidence that supports an AI detector score” and at the same time “helps debunk some hackneyed arguments that liken human reading of books to AI being trained on books” – a standard line of defense from AI companies in copyright disputes.

The results have immediate implications for copyright lawsuits against AI companies. In Kadrey v. Meta, Judge Vince Chhabria ruled in Meta’s favor in June 2025 but sent a strong warning. He said it was hard to imagine that using copyrighted books to build a product generating billions in revenue while producing a potentially endless flood of competing works would qualify as fair use.

The plaintiffs had presented no empirical evidence of this market dilution, though. The current study now delivers exactly the kind of evidence that was missing, showing that books with no detected AI text earn less as AI titles enter the market in large numbers.

Making the legal case harder is the fact that AI content can’t be identified on the platform itself. Authors must disclose AI involvement when publishing through Kindle Direct Publishing, but Amazon doesn’t pass that information along to customers. The platform has struggled for years with AI titles that hijack the names and styles of well-known authors and has mainly responded by capping publications at three per day.

The language overlap findings align with a November 2025 study showing that language models can reproduce passages from copyrighted books nearly word for word. Also in fall 2025, another paper showed that just two books are enough to fine-tune a model on an author’s style.

What it means

For writers relying on the platform, the margin for error has shrunk. The market is saturated with new entries, and the revenue available for each individual title is dropping. Even if a book is not AI-generated, it now competes against a flood of synthetic content. The data suggests that the most successful AI books do not write their own words but instead mimic the rare phrasing found in specific existing books, a capability that complicates legal arguments about copyright infringement.


Scroll to Top