GuideUpdated 2026-08-16

Are AI Companies Buying Secondhand Books for Training Data? What Is Actually Known

Booksellers report unusual bulk orders for unrelated titles, but the buyers and intended use remain unconfirmed. The episode exposes demand for pre-AI text.

By DiscoverAI Editorial Team4 min readHow we evaluate

Bottom line

UK and Irish booksellers report opaque bulk orders that they suspect may supply AI training datasets. Here is the confirmed evidence and why pre-AI books are valuable.

Editorial accountability

Who checked this guide

Meet the editorial team →
Evaluation type
Research-based verification
Last materially checked
Evidence
5 listed sources

Hands-on testing is identified explicitly. Research-based coverage uses cited product documentation and other named sources; it does not imply every paid plan was used. Read the full methodology.

Editorial basis

What this guidance is based on

Editorial basis
Source-led analysis
Primary references
5
Products covered
2
Last checked
2026-08-16

Important limits

  • Features, availability, and pricing can change after publication; confirm consequential details with the provider.
In this guide
  1. The short answer
  2. What booksellers reported
  3. Why old books are attractive as AI data
  4. What is known about prior book acquisition
  5. What transparency should look like
  6. Why this trend matters for AI quality
  7. The larger trend

*This is a research-based news analysis of reporting published August 15, 2026, plus public court records and provider documentation. The identities and purposes of the buyers behind the reported UK and Irish orders have not been established. We do not claim that any named AI company placed those orders.*

The short answer

Secondhand booksellers in the UK and Ireland have reported unusually large, apparently patternless orders for books across unrelated subjects, languages, and editions. Some sellers suspect the purchases are destined for scanning into AI-training datasets, but the cited reporting did not identify the ultimate buyers or prove how the books will be used.

The suspicion is plausible because books provide edited, long-form text—especially titles published before generative AI filled the web with synthetic material—and because documented AI-data programs have previously acquired and scanned physical books. Plausible is not confirmed. The immediate story is an opaque supply chain, not proof that a particular lab ordered these titles.

What booksellers reported

The Guardian reported that sellers received bulk orders from buyers in several countries, sometimes for hundreds or thousands of unrelated books. Sellers described unusual combinations, a willingness to pay full prices, repeated delivery patterns, and disruption to ordinary inventory and fulfillment.

Examples ranged across languages, obscure magazines, literature, biographies, and specialized historical subjects. That lack of a normal collector or subject pattern contributed to speculation that a machine-generated acquisition list was seeking broad textual coverage.

Marketplaces, intermediaries, recyclers, research organizations, resellers, digitization businesses, or AI companies could all participate in such a chain. Without contracts, shipping records, buyer confirmation, or a documented destination, the final use remains unknown.

Why old books are attractive as AI data

Books offer sustained arguments, edited prose, domain knowledge, translation pairs, uncommon vocabulary, and material that may not exist in clean digital form. Older editions also predate the recent flood of AI-generated web text, which can make them attractive to teams trying to reduce synthetic-data contamination.

Physical ownership does not automatically settle every right involved in scanning, copying, training, redistributing a dataset, or generating outputs. Those questions depend on jurisdiction, the works, licenses, exceptions, contracts, and the use. Buying a copy and acquiring rights to reproduce its contents are not necessarily the same transaction.

What is known about prior book acquisition

Public litigation and reporting have described large-scale book acquisition and scanning connected with model development. In the Anthropic copyright litigation, court records distinguished questions involving lawfully purchased books from allegations involving unauthorized digital libraries. Those findings are specific to that case; they do not identify the buyers in the current reports.

Anthropic told The Guardian that books are one source among publicly available web data, commercially acquired datasets, and company-generated data, and that its acquisition programs do not buy and destroy rare or antiquarian books. That statement does not confirm involvement in the reported orders.

What transparency should look like

AI developers should document major training-data categories, acquisition methods, rights controls, vendor requirements, retention, and whether physical materials are preserved or destroyed. Dataset intermediaries should maintain provenance records that survive resale and aggregation.

Booksellers and marketplaces can offer a disclosure path for unusually large digitization orders, publish preservation standards, flag rare-material risk, and distinguish ordinary resale from bulk data acquisition. That need not expose legitimate buyers' private information; it can provide aggregate transparency and contractual accountability.

Authors and publishers should avoid assuming a suspicious order proves infringement. They can preserve sales and rights records, monitor licensing activity, review current rights-reservation mechanisms where available, and seek qualified advice when evidence connects a work to an unauthorized dataset.

Why this trend matters for AI quality

The race for pre-AI books signals a data-quality problem. More internet text is now generated, summarized, translated, or rewritten by models. Training future systems indiscriminately on those outputs can amplify errors, flatten style, and obscure provenance.

High-quality human work is becoming more economically valuable at the same time that its creators demand clearer consent and compensation. Sustainable AI development needs both: strong source material and a credible chain of rights, attribution, preservation, and payment.

The larger trend

Model competition is shifting from raw data volume toward trustworthy data supply. The winning dataset is not merely large; it is useful, documented, legally supportable, and maintainable when a source is corrected or withdrawn.

The unexplained book orders are a warning precisely because the evidence stops short of the final buyer. An accountable AI-data market should make provenance easier to establish, not require the public to reverse-engineer it from shipping patterns.

Sources and verification

Product details and claims were checked against the following primary sources.

Frequently asked questions

Have AI companies been confirmed as the buyers?

No. Booksellers suspect AI-related acquisition, but the August 15 reporting did not establish the ultimate buyers or intended use.

Why would an AI company want books published before ChatGPT?

Older books can provide edited, long-form human writing that predates widespread generative-AI output and may not be available in clean digital form.

Does buying a physical book include AI training rights?

Physical ownership and rights to scan, reproduce, license, or train on a work are distinct questions that depend on contracts, jurisdiction, and legal basis.

Did Anthropic place the reported orders?

The cited reporting did not establish that. Anthropic described its general practices, but that statement neither confirms nor proves involvement in these orders.

Continue exploring

A useful next step

Editorial illustration of a large source dossier splitting into two AI analysis paths and rejoining at a human citation-verification desk
ComparisonContent & Search

ChatGPT vs Claude for Long Documents in 2026: A Practical Test

A practical, evidence-led guide for people searching for ChatGPT vs Claude long documents.

Claude is often a strong starting point for sustained document analysis, while ChatGPT offers a broader surrounding toolset. The reliable choice is the one that preserves citations, constraints, and nuance on your own representative document. Includes a repeatable framework, measurement plan, limitations, and primary sources.

Read guide

WorkflowWork & Operations

How Nonprofits Can Use AI for Grant Writing and Fundraising in 2026

A practical workflow for using AI assistants to draft, refine, and track grant proposals without losing the human voice funders expect.

A practical workflow for using AI assistants to draft, refine, and track grant proposals without losing the human voice funders expect. Written for nonprofit development directors, grant writers, and executive directors, with a decision framework, step-by-step workflow, measurable outcomes, and clear limitations.

Read guide

WorkflowWork & Operations

How to Write Small Business Proposals and RFPs With AI in 2026

A repeatable process for using AI to draft, tailor, and polish business proposals that win contracts without spending weekends on paperwork.

A repeatable process for using AI to draft, tailor, and polish business proposals that win contracts without spending weekends on paperwork. Written for small business owners responding to RFPs, bids, and client proposals, with a decision framework, step-by-step workflow, measurable outcomes, and clear limitations.

Read guide

WorkflowWork & Operations

Nonprofit Impact Reporting: Using AI to Measure and Communicate Results in 2026

How to turn program data into compelling impact reports, dashboards, and stakeholder updates using AI—without needing a data analyst on staff.

How to turn program data into compelling impact reports, dashboards, and stakeholder updates using AI—without needing a data analyst on staff. Written for nonprofit program managers and executive directors reporting to funders and boards, with a decision framework, step-by-step workflow, measurable outcomes, and clear limitations.

Read guide

Keep the useful part coming

Practical AI guidance for lean teams.

Get one weekly email with important tool changes, carefully selected resources, and workflows you can actually use. No hype; unsubscribe any time.

Tools mentioned in this article

Claude

Anthropic's thoughtful, safety-focused AI with exceptional long-form reasoning

4.5

Claude excels at deep analysis, long-form writing, and nuanced reasoning. Built by Anthropic with a focus on safety and helpfulness.

FreemiumChatbotsWriting

ChatGPT

The general-purpose AI assistant that started it all

4.6

OpenAI's flagship conversational AI model, powering everything from casual chat to complex reasoning, coding, and creative work.

FreemiumChatbotsWriting