GuideUpdated 2026-07-31

AI's New Battlefield: Why 'Model Distillation' Just Became the Hottest Flashpoint in US-China Tech Relations

American AI companies Anthropic and OpenAI have accused Chinese firms — including DeepSeek, Moonshot AI, and MiniMax — of systematically using a technique called 'model distillation' to extract capabilities from proprietary US models and train competing Chinese alternatives. Chinese companies and researchers argue distillation is a standard, legitimate research tool used globally. The dispute has opened a new front in the US-China AI rivalry — one that could reshape who has access to the most capable AI. Here's what distillation actually is, why it's become a flashpoint, and what the fight means for the future of AI development.

By DiscoverAI Editorial Team10 min readContent & SearchHow we evaluate

Bottom line

Model distillation — a technique where a smaller 'student' model learns from a larger 'teacher' model's outputs — has become the newest battleground in the US-China AI competition. US companies Anthropic and OpenAI allege Chinese firms have systematically used their APIs to extract frontier model capabilities and train competing Chinese models. Chinese researchers argue distillation is a standard practice used by both countries, and that restricting it would entrench the dominance of well-funded US AI companies. This article explains the technique, the arguments on both sides, and what the dispute means for businesses caught between the world's two AI superpowers.

In this guide
  1. The Short Answer
  2. What Is Model Distillation? A Plain-Language Explainer
  3. The Accusation: What Anthropic and OpenAI Are Alleging
  4. The Counterargument: China's Defense of Distillation
  5. The Legal Landscape: What Law Actually Applies?
  6. What the Dispute Means for the AI Industry

The Short Answer

The model distillation dispute is both technically fascinating and strategically consequential. Here's the practical bottom line:

Distillation is a real, widely used technique — not a euphemism for theft. Knowledge distillation has been a standard machine learning practice since at least 2015, used by researchers at Google, OpenAI, Meta, and universities worldwide. It's how you make AI models smaller, faster, and cheaper while preserving most of their capability. Every major AI company uses it.

The dispute is about scale, intent, and terms of service. The US accusation isn't that distillation itself is wrong — it's that Chinese companies have used it at industrial scale, against explicit API terms of service, to build directly competing products. The Chinese counterargument is that API terms of service restricting distillation are anticompetitive — an attempt to use legal fine print to prevent competition that technology alone can't stop.

There's no clear legal resolution in sight. US law doesn't clearly address whether training a model on another model's outputs violates anything — copyright law, computer fraud law, and trade secret law are all imperfect fits for the problem. API terms of service are contracts, not statutes. The dispute may ultimately be resolved through regulation (new rules about what you can and can't do with AI model outputs), industry norms (voluntary agreements not to distill competitors' models), or continued escalation (technical measures to prevent distillation).

Your practical takeaway: The distillation fight won't directly affect most business AI users in the near term. But it could shape which AI models are available, at what cost, and under what restrictions over the medium term. If the US restricts distillation, Chinese models may become less competitive with US models — reducing choice. If distillation continues unrestricted, Chinese models will keep improving rapidly — increasing competition and lowering costs. Either way, the smart strategy is the same: use multiple AI providers, don't bet exclusively on any single company or country's models, and stay informed as the regulatory landscape evolves.

What Is Model Distillation? A Plain-Language Explainer

The concept is simpler than the name suggests. Here's how it works:

The basic idea: You have a very large, very capable AI model — the 'teacher.' It's expensive to run and requires massive computing infrastructure. You want a smaller, cheaper, faster model — the 'student' — that's nearly as good. So you feed the teacher millions of prompts, collect its responses, and use those prompt-response pairs to train the student. The student learns to mimic the teacher's behavior without needing to understand why the teacher answered the way it did.

An analogy: Think of a master chef (the teacher) and an apprentice (the student). The apprentice can't replicate the chef's decades of experience, intuition, and creativity. But if the apprentice watches the chef make 10,000 dishes and carefully notes every ingredient, technique, and timing, the apprentice can produce dishes that are nearly indistinguishable from the chef's — without ever understanding the culinary principles behind them. That's distillation.

Why it's valuable: Training a frontier AI model from scratch costs hundreds of millions of dollars in computing alone, plus years of research. Distilling an existing model's capabilities into a smaller model costs a tiny fraction of that — thousands or tens of thousands of dollars in API fees. It's the difference between building a car factory from scratch and buying a finished car to study its design.

Why it's controversial: The teacher model's capabilities represent billions of dollars of investment by the company that built it. Distillation allows a competitor to capture much of that capability at minimal cost — without doing the underlying research, without building the infrastructure, and without taking the risks the original developer took. The question at the heart of the dispute: is this fair competition, or is it taking a free ride on someone else's massive investment?

The Accusation: What Anthropic and OpenAI Are Alleging

The US companies' complaints, which surfaced prominently in July 2026, make several specific claims:

Systematic, not incidental: Anthropic and OpenAI aren't alleging that a few Chinese researchers occasionally used their APIs for distillation experiments. They're alleging systematic, industrial-scale extraction — millions or billions of API calls designed specifically to capture frontier model capabilities for use in training competing Chinese models.

Terms of service violations: Both Anthropic and OpenAI's API terms of service prohibit using their outputs to train competing models. The allegation is that Chinese companies knowingly violated these terms, using intermediary accounts, shell companies, or third-party services to obscure the scale and purpose of their API usage.

The specific companies named: DeepSeek (known for highly competitive open-weight models), Moonshot AI (creator of the Kimi K3 model), and MiniMax (a major Chinese AI company) have been identified in reporting as targets of the accusations. All three have either denied the allegations or defended distillation as legitimate research practice.

The competitive impact: US companies argue that distillation at this scale allows Chinese firms to free-ride on American AI investment — capturing the benefits of models that cost billions to develop for the price of API access. This, they argue, undermines the economic incentive to invest in frontier AI research and creates an uneven playing field where Chinese companies can compete without bearing comparable R&D costs.

The security dimension: Some US officials have added a security concern: if Chinese companies can extract frontier AI capabilities through distillation, those capabilities could be transferred to Chinese military or intelligence applications without the safeguards and oversight that US companies build into their own deployments.

The Counterargument: China's Defense of Distillation

Chinese AI companies and researchers have pushed back forcefully against the accusations, making several arguments:

Distillation is standard global research practice. Google's 2015 paper introducing knowledge distillation is one of the most cited papers in machine learning. Every major AI lab — including OpenAI and Anthropic — has published research using distillation techniques. US companies use distillation to train smaller, more efficient models. The technique itself is not controversial in the research community — only its application to commercial competitors' outputs.

API terms of service shouldn't restrict research. Chinese researchers argue that using publicly available API outputs for research purposes — including training better models — is a legitimate activity that API terms of service shouldn't be able to restrict. They draw an analogy: if a company sells a book, it can't prohibit you from learning from that book and writing a better one. If a company sells API access to an AI model, it shouldn't be able to prohibit you from learning from that model's outputs.

The accusations are protectionism. Chinese officials and researchers argue that the real motivation behind the distillation accusations is protecting US AI companies from competition. Chinese open-weight models (DeepSeek, Qwen, Kimi) have become genuinely competitive with US models — and often significantly cheaper. Restricting distillation would remove one of the key techniques that enabled that catch-up, entrenching the advantage of well-funded US incumbents.

US companies also distill — they just don't call it that. Critics note that US AI companies train their models on vast quantities of internet data that includes outputs from other AI models — effectively a form of distillation, whether intentional or not. The distinction between 'training on internet data that happens to include AI outputs' and 'training on AI outputs deliberately' is blurry in practice.

The security argument cuts both ways. If distillation is restricted, the global AI ecosystem becomes more dependent on a smaller number of US-controlled proprietary models. That concentration creates its own security risks — including the risk that US AI companies could restrict access for geopolitical reasons. Chinese officials argue that a diverse, globally distributed AI model ecosystem — enabled by distillation and other techniques — is more resilient and secure than one dominated by a few US companies.

One reason the distillation dispute is so hard to resolve is that existing law doesn't clearly address it:

Copyright law: Can you copyright the output of an AI model? US law is unsettled. The Copyright Office has generally taken the position that works created by non-humans aren't copyrightable — which would mean AI outputs aren't protected by copyright and can't be infringed. But this isn't definitively resolved, and some argue that AI outputs should be protected as derivative works of the copyrighted training data. If AI outputs aren't copyrightable, distillation can't be copyright infringement.

Computer fraud law: The Computer Fraud and Abuse Act (CFAA) prohibits unauthorized access to computer systems. But using a publicly available API — even in violation of terms of service — may not constitute 'unauthorized access' under current legal interpretations. The Supreme Court narrowed the CFAA's scope in 2021, ruling that exceeding authorized access doesn't necessarily violate the statute. Whether violating API terms of service to distill a model is a CFAA violation is genuinely unclear.

Trade secret law: If a company can show that its model's outputs reveal trade secrets — proprietary information that derives economic value from not being generally known — distillation could potentially be trade secret misappropriation. But AI model outputs are generally not trade secrets: they're publicly available (to anyone with API access) and the model itself is the valuable asset, not individual outputs. This is a difficult legal theory to sustain.

Contract law: API terms of service are contracts. Using the API in violation of those terms is a breach of contract. But the remedy for breach of contract is typically damages — and proving damages from distillation (as opposed to lost API revenue) is difficult. An injunction prohibiting future violations might be available, but shutting down an already-trained model based on a contract claim would be extraordinary.

The bottom line: Existing US law doesn't provide a clear, effective remedy for distillation-based competition. This is part of why the dispute is escalating — and why it may ultimately require new legislation or international agreements rather than litigation under existing law.

What the Dispute Means for the AI Industry

1. The API model has a vulnerability. The distillation dispute exposes a structural vulnerability in the 'sell API access to your AI model' business model. If API access allows competitors to extract your model's capabilities and build competing products, the API business model is self-undermining — every customer is potentially a future competitor. AI companies may respond by: restricting API access more tightly, raising API prices to capture more of the distillation value, implementing technical measures to detect and prevent distillation, or shifting toward on-device and enterprise deployment models where the model is less accessible.

2. Technical anti-distillation measures are coming. AI companies are already working on technical measures to prevent or detect distillation: output watermarking (embedding detectable patterns in model outputs), query pattern analysis (detecting the large-scale, systematic querying characteristic of distillation), and output degradation for suspected distillation queries. These measures will create an ongoing arms race between distillation techniques and anti-distillation defenses.

3. Open-weight models change the calculus. Open-weight models (Llama, DeepSeek, Qwen, Mistral) don't have API terms of service — anyone can download and use them for any purpose, including distillation. The growth of capable open-weight models means distillation is becoming democratized: you don't need API access to a proprietary frontier model to practice distillation; you can use the best available open-weight models. This undermines the API-based anti-distillation approach and may accelerate the shift toward open-weight models as the distillation foundation.

4. The global AI ecosystem may fragment. If the US restricts distillation (through law or technical measures), and China permits it, the global AI landscape fragments: Chinese AI models improve rapidly through distillation while US models are protected but develop more slowly without the competitive pressure. If both countries restrict it, AI development slows globally but incumbents' positions are protected. If neither restricts it, competition intensifies and costs fall, but the economic incentive to invest in frontier research may weaken. These are genuinely difficult tradeoffs with no obvious right answer.

5. Businesses benefit from the competition — for now. The distillation-enabled competition between US and Chinese AI models has been a major driver of falling AI costs and improving capabilities. Chinese open-weight models, many of which benefited from distillation techniques, have given businesses access to capable AI at dramatically lower costs than US proprietary alternatives. If distillation is restricted, some of that competitive pressure diminishes — and AI costs may stop falling as fast. Businesses should enjoy the current competitive environment while recognizing it may not last.

Sources and verification

Product details and claims were checked against the following primary sources.

Frequently asked questions

Is model distillation the same as stealing AI technology?

No — but the line between legitimate learning and unfair extraction is genuinely blurry. Distillation is a standard machine learning technique, analogous to a junior developer learning from a senior developer's code. The junior developer isn't 'stealing' — they're learning. But if the junior developer's entire training consisted of copying the senior developer's output line-by-line and reselling it at half the price, that crosses into territory most people would consider unfair. The distillation dispute lives in the gray area between these extremes. The scale matters (a few experiments vs. industrial-scale extraction), the intent matters (academic research vs. building a direct commercial competitor), and the terms of access matter (did the distiller agree to API terms prohibiting this use?). There's no simple answer, which is why the dispute is so hard to resolve.

Will distillation restrictions make Chinese AI models worse?

They would likely slow the rate of improvement, but not stop it entirely. Chinese AI companies have multiple paths to improvement beyond distillation: original research (Chinese AI research output is world-class), training on non-distilled data, using open-weight models (which don't have API restrictions) as teachers, and developing novel architectures and training techniques. Distillation has been an accelerator, not the sole driver, of Chinese AI progress. Restrictions would remove one tool from the toolkit — a significant but not catastrophic loss. And restrictions might actually accelerate Chinese investment in original AI research by forcing companies to develop capabilities they previously accessed through distillation.

How can I tell if an AI model I'm using was trained through distillation?

You generally can't — and this is part of the problem. There's no reliable technical method to determine whether a model was trained using outputs from another model. Some researchers are working on detection techniques (looking for statistical signatures of distillation in model behavior), but these are research-stage and not yet reliable. From a practical standpoint, most competitive AI models — US and Chinese, proprietary and open-weight — have likely been influenced by distillation to some degree, whether intentional or incidental (training data inevitably includes AI-generated content from other models). The question is usually about scale and intent rather than whether any distillation occurred at all. For businesses, this means evaluating models based on their demonstrated performance, reliability, and licensing terms rather than trying to determine their training provenance.

Should I avoid using Chinese AI models because of the distillation controversy?

Not based on the distillation controversy alone. The controversy is between AI companies and governments — it doesn't create legal risk for end users of Chinese AI models (unless you're using those models to do something that's independently illegal or restricted). Chinese open-weight models remain some of the most capable and cost-effective AI tools available, and many Western businesses use them successfully. The relevant considerations for whether to use Chinese AI models are: your organization's risk tolerance regarding geopolitical uncertainty, any regulatory requirements that apply to your industry (government contractors may face restrictions), data privacy considerations (where does your data go when you use the model?), and the availability of suitable Western alternatives. The distillation controversy adds to the geopolitical uncertainty but doesn't, on its own, create new legal exposure for end users.

Continue exploring

A useful next step

View topic →
WorkflowWork & Operations

How Nonprofits Can Use AI for Grant Writing and Fundraising in 2026

A practical workflow for using AI assistants to draft, refine, and track grant proposals without losing the human voice funders expect.

A practical workflow for using AI assistants to draft, refine, and track grant proposals without losing the human voice funders expect. Written for nonprofit development directors, grant writers, and executive directors, with a decision framework, step-by-step workflow, measurable outcomes, and clear limitations.

Read guide

ComparisonWork & Operations

ChatGPT vs Claude vs Gemini: Real Small Business Task Showdown 2026

We tested all three AI assistants on six specific small business tasks — proposals, customer emails, financial analysis, policy drafting, content creation, and meeting summarization — to help you pick the right one for your actual work.

Most AI assistant comparisons focus on benchmarks and abstract capabilities. We tested ChatGPT, Claude, and Gemini on the tasks small business owners and nonprofit leaders actually do every week. Here's which one performed best on each task — and which to choose for your specific work.

Read guide

ReviewWork & Operations

ChatGPT Review 2026: The AI Assistant That Defined a Category, Thoroughly Tested

We tested ChatGPT across 75 real-world business tasks — writing, analysis, coding, research, and creative work — to give you an honest assessment of what the world's most popular AI assistant actually delivers for small businesses and nonprofits in 2026.

ChatGPT is the most widely used AI tool on the planet, but popularity isn't the same thing as suitability for your specific needs. We spent three weeks testing ChatGPT against real small business and nonprofit tasks to answer the question that matters: is it the right AI assistant for your organization, or are you using it because everyone else does?

Read guide

ReviewContent & Search

Google Gemini Review 2026: Google's AI Assistant for the Workspace Era, Tested

We tested Gemini Advanced across business writing, research, data analysis, and Google Workspace integration to determine whether Google's AI is the smart choice for organizations that live in Gmail, Docs, and Sheets.

Google Gemini is deeply integrated into the Google ecosystem that millions of businesses already use daily. We tested Gemini Advanced across 60 real business tasks — and directly compared it to ChatGPT, Claude, and Perplexity — to help you decide whether Gemini's Google integration makes it the right AI assistant for your organization.

Read guide

Keep the useful part coming

Practical AI guidance for lean teams.

Get one weekly email with important tool changes, carefully selected resources, and workflows you can actually use. No hype; unsubscribe any time.

Tools mentioned in this article