Training AI on Protected Content: Who's Really Going to Pay?
This is no longer a theoretical debate. Thousands of authors are suing AI model makers—OpenAI, Anthropic, Meta—for unauthorized use of their work during training. The problem: there’s currently no stable legal precedent. U.S. courts are ruling differently case by case. Meanwhile, the EU is finalizing a copyright directive that could mandate explicit licenses for any protected content used in ML.
Why is it murky? Because no one has definitively established whether training models on publicly available data counts as “transformative use” (possible exemption) or just mass copying (violation). Even lawyers acknowledge it’s “complicated.”
Most likely outcome: Model providers (OpenAI, Anthropic, etc.) will either agree to pay royalties, restrict their datasets, or filter certain content. None of these options are free. As an SMB using these services, you’ll see these costs reflected in your API pricing or subscription plans.
What this means for your business
For your business, watch these three scenarios:
-
Prices will climb: If models must pay copyright fees, your OpenAI/Anthropic bills will go up.
-
Weaker models: If training is limited to explicitly licensed content, performance could decline.
-
Open-source models gain appeal: Models trained only on royalty-free data (like Llama under strict regulatory oversight) could become attractive—but with lower performance.
Action now: Document your current API usage. If a retroactive fine lands on the model provider, you don’t want surprises.
In brief
OpenAI Gains Ground with SMBs—But Instability Looms
OpenAI is recovering market share from Anthropic in enterprise. Problem: customers regularly switch models based on updates. This churn frustrates investors and shows how fragile AI customer loyalty really is. Takeaway for you: don’t lock your stack into a single provider.
Inherent (Former DeepMind) Launches Strong in Automated Research
A new player—Inherent—releases an AI agent specialized in replicating scientific studies. Outperforms OpenAI/Anthropic models in this niche. Key insight: future AI leaders won’t be generalists, but specialists by domain (R&D, finance, content, etc.).
AI Safety Systems Break Easier Than Expected
Claude’s guardrails (Anthropic) meant to block illegal content can be bypassed with ease. Implication: if you deploy customer-facing AI, its automated compliance promises aren’t enough. You remain legally responsible.
When AI Discovers a Drug, Who Gets the Credit?
Insilico Medicine claims a molecule was “discovered by” its AI. Pressing question: who bears legal and commercial liability if the molecule has side effects? Accountability frameworks don’t exist yet.
Get The AI Brief in your inbox
3x per week, the essentials of AI decoded for business leaders.