Using Copyrighted Material for AI Training not Fair Use, Third Circuit Court of Appeals Finds

'bout time. The audacity of claiming fair use is "we'll just copy it!" While AI (or really LLMs) have some impressive capabilities I cannot shake that it is just automated plagiarism.

I fear it is already too late. How does an AI company prove they have removed all the copyrighted material from their model? I don't see them just throwing away all their training expense.
 
Well, the actual case was about a company outputting verbatim summaries from the legal search service they crawled, which wasn't judged to be transformative and pretty clear cut plagiarism. While the article quotes an AI company rep saying that if a model is highly transformative then it may well be deemed fair use in future cases.
 
Isn't all user posted content technically copyrighted automatically by that user in the USA?
That's why it's never going to stand.

Software, by default, is copyrighted. This would pretty much make ANY vibe-coded project, anything by Claude Code, etc., illegal (unless you actually owned the rights to all the training material, which is practically none of them.). But goodbye anything off of Stack Overflow, GitHub, etc.

As a Software Engineer, yeah, I like this current decision, but there's far too much money involved to allow the decision to stand.
 
Yep. The horse is long out of the barn.
I would enjoy watching the writhing if a legal decision could instantly pull the plug on all of the frontier models. 😅

Sure, it would be one hell of a ride, but it would be totally worth it.

I'd also pay to see every crypto-bro's face after a cryptocurrency ban, and the response to the end of all prediction markets, but I guess I'm just weird that way😅
 
I would enjoy watching the writhing if a legal decision could instantly pull the plug on all of the frontier models. 😅

Sure, it would be one hell of a ride, but it would be totally worth it.

I'd also pay to see every crypto-bro's face after a cryptocurrency ban, and the response to the end of all prediction markets, but I guess I'm just weird that way😅
For some reason, this is how I picture zarathrusta: sitting in a lawnchair, sipping on a cold one, and watching the world burn..

1791338074923.jpg
 
Wayyyyyy too late for this.
Yep. The horse is long out of the barn.

I mean, meh. A lot of people would panic, and some institutional investors would lose a lot of money, but businesses would scramble to rehire some of the people they laid off due tot he anticipated labor savings from AI that never materialized, and pretty soon we'd be back to normal.

I'd just love to see the reaction on Sam Altman's smug, punchable face.
 
"Ross shut down in 2021" - I'll join the way too late thing.

If this had been litigated to conclusion in 2021, maybe. Not a chance the US federal government is going to knee-cap this now.

At most, there's some fines levied. Similar in size/impact to the fines financial institutions get for fraudulent behavior; far less than the money gained.
 
That really not what the ruling say, training (or human reading-watching and learning via that action) of copyrighted material could still be fair use, the model output need to be transformative too. Here for this case, the AI was detecting what was the answer and made literal copy paste of the matching material for its answer as a direct competitive subsitute. That just sensational headlines being used by publication:

HTfBXz_bUAA4yiZ?format=jpg&name=large.jpg
 
It is unfortunate, but the people creating the content (outside of big operations like Paramount, Disney, and big publishers), rarely have the resources to make these people hurt for it. The people being hurt the most don't have a chance in hell of making AI companies pay for the unfair use. Yes, it is by the strict definition "transformative" but its end goal is to wipe out the original content creator's ability to do his job. The end goal of me studying someone's art and painting something similar is not to wipe out that other person's ability to get paid for their work. The end goal of AI is to wipe out all the jobs of the people who's data it is using to train its models. I feel in these cases the law is not sufficient to solve the problem at hand.
 
I'd argue that what should happen is this:

Any model trained on public data becomes the property of the public. The AI companies that develop it should be legally blocked from monetizing it. They should only be able to monetize models where they outright own all of the material said model is trained on. And copyrighted material should never be used for any reason.

And if that kills the AI industry so be it. We'd probably be better off if it did.

Let the Chinese win this race. They can have it. This is a race you do not want to win. It is a one way ticket to dystopia.

Absolutely nothing would make me happier than to see every last AI company bankrupt and all of their investors 100% wiped out.

What we should be focusing on is narrow windows of acceptable AI for special purpose tasks (flying a jet, filtering noise, etc.), specifically where the developer of the model outright owns 100% of the data said model is trained on. And I'm not talking any of that bullshit where social media companies claim they own the rights to the content of their users. Any user data would be outright prohibited unless they have the individual permission from every individual user, separate from any agreement they might need to agree to in order to use a platform.

We should be trying to steer the AI industry away from these "everything models" trained to communicate like people, and into only special purpose models that look and feel like the tools they are.

None of us got a say on if we wanted our entire society changed by these ass clown tech bros. We should have been asked first. They should have needed permission before rolling anything new out. Like an FDA pre-market clearance for a drug. Letting these unethical shitbags just roll out whatever they want, do whatever damage to society they want, in the hopes of becoming too big to fail is completely and utterly unacceptable in every way shape and form.
 
Last edited:
I'd argue that what should happen is this:

Any model trained on public data becomes the property of the public. The AI companies that develop it should be legally blocked from monetizing it. They should only be able to monetize models where they outright own all of the material said model is trained on. And copyrighted material should never be used for any reason.

And if that kills the AI industry so be it. We'd probably be better off if it did.

Let the Chinese win this race. They can have it. This is a race you do not want to win. It is a one way ticket to dystopia.

Absolutely nothing would make me happier than to see every last AI company bankrupt and all of their investors 100% wiped out.
You can't currently copyright works that are AI generated. That's why companies like Oracle are still seeking software developers who don't use AI.

There's a difference between an internal tool, or a product you don't charge money for (e.g. Google and Facebook make money through ads, not software), and a professional program in which you can now pirate Adobe Photoshop to your hearts content, as you're no longer breaking copyright laws. This is now going to lead to more DMCA as well as companies using trademarks and patents. The era of copyright is over now.
 
What we should be focusing on is narrow windows of acceptable AI for special purpose tasks (flying a jet, filtering noise, etc.), specifically where the developer of the model outright owns 100% of the data said model is trained on. And I'm not talking any of that bullshit where social media companies claim they own the rights to the content of their users. Any user data would be outright prohibited unless they have the individual permission from every individual user, separate from any agreement they might need to agree to in order to use a platform.

I love the use of AI for things like cancer screening for instance. That literally helps save lives. Hard to argue against that sort of approach. I also don't have issue with running AI models against air traffic control data to highlight possible problems. I see it as a sort of extra set of eyes sitting behind an air traffic controller that could flag things they might have missed. There are dozens and dozens of industries where targeted AI can do great work and is already doing great work. But, they can't get the massive investment dollars for these targeted things and that's why they've pushed hard for LLM that they claim can "do anything" and replace hundreds and thousands of workers.
 
It is unfortunate, but the people creating the content (outside of big operations like Paramount, Disney, and big publishers), rarely have the resources to make these people hurt for it. The people being hurt the most don't have a chance in hell of making AI companies pay for the unfair use. Yes, it is by the strict definition "transformative" but its end goal is to wipe out the original content creator's ability to do his job. The end goal of me studying someone's art and painting something similar is not to wipe out that other person's ability to get paid for their work. The end goal of AI is to wipe out all the jobs of the people who's data it is using to train its models. I feel in these cases the law is not sufficient to solve the problem at hand.

The legal system moves slowly. Sometimes I think that's an intentional feature that just happens to benefit those that make multimillion dollar "campaign contributions". What an incredible coincidence!
 
And I'm not talking in some stupid Terminator way. I'm talking about it destroying society as we know it.

If you are even loosely paying attention, that is the goal. A few mega billionaires *want* to recreate society. They've point blank said it. I bet you can't guess who gets to be on top in this scenario!

If power can be used, it WILL be abused.
 
I mean, meh. A lot of people would panic, and some institutional investors would lose a lot of money, but businesses would scramble to rehire some of the people they laid off due tot he anticipated labor savings from AI that never materialized, and pretty soon we'd be back to normal.
My company just got the projected Anthropic bill for next year - 10X increase.
I'd like to think that means things go back to normal, but we'll see. The options are more likely either a slowing down or migration to some Chinese models.
Things are going to get interesting.
 
My company just got the projected Anthropic bill for next year - 10X increase.
I'd like to think that means things go back to normal, but we'll see. The options are more likely either a slowing down or migration to some Chinese models.
Things are going to get interesting.
They have to increase the prices. All of the frontier models are so incredibly deeply in debt with no hopes of ever paying it off at current prices.

And that goes back to why these AI's went "rogue". Because as Mark Baum stated, these AI companies don't have a moat. AI's not going away, but more people and companies are switching to local models rather than the frontier models. They do good enough of a job, and you don't have to worry about some large corporation stealing your data.

If you watched the NYC meeting today with the OpenAI "whistleblower", the proposed regulation was that all NEW AI companies would need regulation and approval. I.e. the conspiracy theory that the large AI companies are trying to shut the door to any future competition.
 
They have to increase the prices. All of the frontier models are so incredibly deeply in debt with no hopes of ever paying it off at current prices.
Those 2 company do not have specially big debts, they were able to get investor money insteed (for equity at really high valuation) all along and did not had to take much loans (at least that we know of). They do have very big future commitments too.

They will continue to cut price will be my guess. and it will continue to be aggressive, they offer pretty much the cheapest product already;
https://artificialanalysis.ai/models#intelligence-comparison-tabs

GPT-6 Luna max is pretty much the cheapest option that exist, chaeper than GLM or deepseek flash, Haiku 5.5 max is better and cheaper than those, 6.1 sol seem to be the best per dollar model out there.

Anthropic IPO documentation showed that their revenues were growing 12x while their compute cost only by less than 3x annually, if they keep that pace even for a little bit they will become the most profitable company of all time fast.

Competitive pressure will make them continue to aggressively cut their price like we saw the couple of last months (and token efficacy work will also continue, not just raw token pricing), will probably look like this, the giant decline we have seen in compute cost per dollar of revenues will stop, with competition and clients being more cautious with their spent (and routers tooling to use cheaper options when it make sense automatically working well):
1791477378164.png
 
Last edited:
Those 2 company do not have specially big debts, they were able to get investor money insteed (for equity at really high valuation) all along and did not had to take much loans (at least that we know of). They do have very big future commitments too.

They will continue to cut price will be my guess. and it will continue to be aggressive, they offer pretty much the cheapest product already;
https://artificialanalysis.ai/models#intelligence-comparison-tabs

GPT-6 Luna max is pretty much the cheapest option that exist, chaeper than GLM or deepseek flash, Haiku 5.5 max is better and cheaper than those, 6.1 sol seem to be the best per dollar model out there.

Anthropic IPO documentation showed that their revenues were growing 12x while their compute cost only by less than 3x annually, if they keep that pace even for a little bit they will become the most profitable company of all time fast.

Competitive pressure will make them continue to aggressively cut their price like we saw the couple of last months (and token efficacy work will also continue, not just raw token pricing), will probably look like this, the giant decline we have seen in compute cost per dollar of revenues will stop, with competition and clients being more cautious with their spent (and routers tooling to use cheaper options when it make sense automatically working well):
View attachment 830119

I guess it depends. I think most of the industry revenue thus far comes from enterprise contracts, and the Enterprise users are starting to really balk at the cost from what I can tell.

They are going to have to really show significant verifiable value to continue the way they are.

I mean, I personally will never use a model hosted by anyone else, paid or free for any reason, and nothing could ever convince me otherwise, but I understand I am likely unusual in taking that stance.

One thing I found interesting in a news report I heard the other day (can't remember the source) is that given how much the AI talk has been dominating everything lately, the adoption is actually lower than I realized. A recent survey suggested that only ~50% of Americans have ever used AI and according to surveys on menlovc.com only some 25% of Americans use AI daily, with only about 55% of regular users actually paying for a service.

Since it is self reported, I assume this figure is low, as it doesn't include behind the scenes stuff that people don't realize they are using, but even sop, that number is much lower than I would have thought. It suggests there is at least room to grow.

Even so, it feels like a bit of a pipe dream that they can attract enough paying users to both make up for the investment in training to date and keep the lights on for ongoing compute needs.

I just don't think there is enough room for revenue growth here to keep these companies afloat without essentially one or two surviving and the rest folding. For survival they need massive market share, and that's not going to happen as long as they split the market between themselves.

Since about 2023 the AI folks have had this "its inevitable, you have to get on board" attitude, but I'm just not convinced.

Especially what with that same menlo report suggesting that approximately 70% of those who are not regular AI users are now entrenched on the "no" side, not trusting it and not wanting it. Many have been radicalized by heavy handed approaches to forcing adoption, by forcing the technology into everything against users wishes.

This suggests that without some serious convincing, the market can't really grow by more than ~30% of non users.

So we have 25% who are regular users.

55% of those regular users, or ~13.75% of Americans currently pay for AI services.

70% of the remaining 75%, or about 52.5% of Americans are currently in the "over my dead body" camp.

This means that unless the "over my dead body" folks can be convinced, at most consumer revenues can grow by about 3.5x, and with Enterprise revenues starting to hit the "OMG we have to control our AI spending, tokens are costing us way more than we can save by reducing headcount" backlash, those might actually shrink.

It really does feel like the entire industry is a house of cards just waiting to collapse on itself.
 
Last edited:
I just don't think there is enough room for revenue growth here to keep these companies afloat without essentially one or two surviving and the rest folding. For survival they need massive market share, and that's not going to happen as long as they split the market between themselves.
Microsoft and Google. They let the first-movers show them how it's done and now that they've blown their wad, they're busy stealing their customers with their deep pockets. MS like a black widow with the contracts and Google with the hardware.
Anthropic has like $4B in net revenue for $512B in debt lol. They're charging dollars/1M tokens, while China is offering tokens for fractions of pennies. They're screwed!
 
My cynical ass thinks this is a response to Pewdiepie using distillation training to use for-profit AI models to train free, locally runnable, open source LLMs.

Essentially, this is JUST IN TIME to use this law to block distillation and smaller actors from using 'copyrighted' LLM content.
 
Back
Top