A $1.5 billion settlement. Millions of pirated books. And a legal distinction so fine that one federal judge used it to let an AI company off the hook for training its models — while still holding it liable for how it got the training data in the first place.
If you use an AI writing assistant, a coding copilot, or an AI-powered research tool in your business, two court cases decided over the last year quietly rewrote the rules for the tools you rely on every day. Neither case involved you. But together, they answer a question every small business owner using AI should be asking: when I use an AI tool built on someone else's copyrighted content, am I exposed?
The $1.5 Billion Question: Bartz v. Anthropic
The case started with three authors — nonfiction writers Charles Graeber and Kirk Wallace Johnson, and thriller novelist Andrea Bartz — who discovered that Anthropic had used their books, without permission or payment, to train its Claude AI models.
What Judge William Alsup decided split the case into two very different questions, with two very different answers.
Question one: Is training an AI model on copyrighted books fair use? Yes — but only for books Anthropic legitimately owned or licensed. The court described the training process as "transformative — spectacularly so," comparing it to how a person reads widely to learn to write. Since the training data itself stays hidden inside the model and doesn't substitute for buying the original book, the judge found the use didn't harm the market for the authors' work.
Question two: Does it matter how the company got the books? Absolutely. Anthropic hadn't just used books it purchased — it had downloaded more than 7 million copyrighted works from "shadow libraries" like Library Genesis (LibGen) and Pirate Library Mirror, sites that host pirated copies of copyrighted books for free. The court was skeptical that any amount of downstream fair use could erase the fact that the books were obtained illegally in the first place. Training might be legal. Piracy is piracy regardless of what you do with the pirated file afterward.
That distinction turned into the largest copyright settlement in U.S. history: $1.5 billion, working out to roughly $3,000 per pirated work before legal fees, with around 120,000 authors and rightsholders filing claims. As part of the deal, Anthropic also agreed to destroy the pirated datasets and certify they're not embedded in any commercially deployed model. The final approval hearing is scheduled for May 2026, with payments beginning shortly after.
The Case That Cuts the Other Way: Thomson Reuters v. Ross Intelligence
If Bartz shows where the fair use line bends in an AI company's favor, Thomson Reuters v. Ross Intelligence shows where it snaps back.
Ross Intelligence built an AI-powered legal research tool to compete directly with Thomson Reuters' Westlaw. To train it, Ross used "Bulk Memos" built from Westlaw's headnotes — the editorially written summaries that distill what a court actually decided in a case — and its Key Number System, a legal classification scheme Thomson Reuters spent decades developing.
The court found the headnotes and classification system were original, copyrightable work, and that Ross's training memos matched the source material so closely that "no reasonable jury could find otherwise" than that copying occurred. Crucially, the court also rejected Ross's fair use defense, for a reason that matters a lot if you're evaluating AI tools for your own business: Ross's tool wasn't generative AI creating new text — it was a search engine competing head-to-head with the product it trained on. That direct market competition, combined with using a competitor's proprietary analytical work rather than published creative content, tipped the fair use analysis against Ross. The case is now on appeal to the Third Circuit, with oral arguments held in June 2026; the ruling will set the first binding appellate precedent on fair use in AI training.
Reading the Two Cases Together
Put side by side, the cases sketch an emerging (if still unsettled) framework:
- Training on lawfully acquired, published creative content — the kind of thing you can buy at a bookstore — has a real shot at being fair use, even without the author's permission.
- Training on pirated copies of that same content is a separate legal problem, regardless of how the training itself is characterized.
- Training on a competitor's proprietary analytical work — internal databases, classification systems, or other non-public assets — faces much steeper fair use odds, especially when the resulting product competes directly with the source.
None of this is fully settled law yet. But it's the clearest signal so far of how courts are thinking about the AI tools you're increasingly building into your workflow.
What This Actually Means If You Use AI Tools
You're very unlikely to be sued for using ChatGPT, Claude, GitHub Copilot, or a legal-research AI tool. The lawsuits target the companies that built the models, not the businesses that use them. But "unlikely to be sued" isn't the same as "no exposure," and a few practical takeaways are worth building into how you evaluate and contract for AI tools:
1. Check what your vendor's contract actually promises. Many AI vendor agreements are still written with traditional software-licensing language that doesn't clearly cover AI-specific risk. Look for indemnification language that explicitly covers claims arising from the training data and the model's outputs — not just the software itself. A vendor that will indemnify you for a bug in their code but stays silent on training-data provenance hasn't actually protected you from the risk these lawsuits highlight.
2. Ask where the training data came from. You don't need a forensic audit, but a vendor that can point to licensing agreements, publisher partnerships, or documented data-acquisition practices is in a fundamentally different risk position than one that can't answer the question. Post-Bartz, "we don't disclose our training data sources" is a less reassuring answer than it used to be.
3. Understand what happens to your own data. Separate from the copyright question, most AI vendors default to using your inputs to further train their models unless your contract says otherwise. If you're feeding an AI tool client financial records, draft contracts, or proprietary business data, you want a "no training, no retention" clause in writing — not just a settings toggle you have to remember to flip.
4. Watch the Ross Intelligence appeal. If the Third Circuit's ruling (expected sometime after the June 2026 arguments) narrows fair use for AI tools that compete directly with the data they trained on, it could reshape pricing, licensing, and availability across an entire category of AI-powered research and analysis tools — the kind that increasingly touch bookkeeping, tax research, and financial analysis software.
A Practical Checklist Before You Sign (or Renew) an AI Tool Contract
You don't need a law degree to do basic due diligence on the AI tools already running in your business. Before your next renewal, walk through this:
- Read the indemnification clause, not just its title. Does it cover claims that the outputs infringe someone's copyright, or only claims about the software malfunctioning? Those are legally different promises, and vendors that haven't updated their templates since before the generative-AI boom often only cover the latter.
- Ask directly: "Where does your training data come from?" A vendor with a clear, documented answer — licensed content, first-party data, public-domain material — is a lower-risk vendor than one that deflects the question. You're not asking to audit their entire dataset; you're asking whether they've thought about it at all.
- Find the data-retention default. Search the contract (or the product's admin settings) for words like "training," "improve our services," or "retain." If you don't see an opt-out, assume your inputs are being used to train the vendor's next model version, including anything sensitive you've pasted in.
- Check if the tool competes with a specific proprietary dataset. This is the Ross Intelligence pattern: a tool that was trained on a competitor's internal analytical work and now competes directly with that competitor sits in much shakier legal territory than a general-purpose writing or coding assistant.
- Keep a simple log. Note which AI tools you use, for what, and whether they touch client or financial data. If a vendor's legal exposure ever becomes your problem — because of a service disruption, a feature getting pulled, or a licensing dispute — you want to know your own footprint without having to reconstruct it from memory.
None of this takes more than an hour, and most of it only needs to happen once per vendor, not every time you log in.
Why This Belongs in Your Bookkeeping Conversation
It's tempting to file AI copyright litigation under "interesting news, not my problem." But if your business uses an AI-powered tool anywhere in your financial workflow — categorizing transactions, drafting client communications, summarizing tax guidance — you're a downstream user of exactly the kind of technology these cases are about. Vendor risk in your tech stack is a real business risk, and reviewing what your AI tools' contracts actually say costs nothing and takes an afternoon.
It's also a good reminder of a broader principle: the more transparent and auditable your own systems are, the less exposed you are when the ground shifts under a vendor you depend on. That's true of AI tools, and it's just as true of your financial records.
Keep Your Own Records Transparent and Auditable
Whatever legal questions get sorted out about the AI tools you use, your own financial records shouldn't be a black box. Beancount.io offers plain-text accounting that's fully transparent, version-controlled, and easy to audit — no proprietary format locking up your data, no vendor you have to trust blindly. Get started for free and see why developers and finance professionals are switching to plain-text accounting.