I am so tired of people dismissing claims that these are plagiarism devices. The creators literally know they essentially stole this work.
This is literally a conversation where they’re deciding if they’re willing to take on the risk, then determining “yes.“
And the worst part? They were absolutely right. They have not suffered any real consequences. And when people try to bring up this flagrant plagiarism and theft, they are shouted down by AI evangelists.
The “hacker” community (for lack of a better term) has a longstanding and well documented skepticism towards the concept of intellectual property in general. “Information wants to be free” and all that.
Let he who has not downloaded from Annas Archive cast the first stone
I think we, as humans (not executives, a different creature altogether if you ask me) make a distinction between a hobbyist hacker, someone doing something for their own curiosity or someone who frees something for others to use freely as well (e.g., F/LOSS licenses, Creative Commons, etc.), vs. someone who uses the "Hacker Ethos" and then promptly builds their own moat where they solely can profit and excludes others from the freedoms they themselves enjoyed.
No, I think the overwhelming majority of copyrighted material should already be public domain. We should all be free to build on it as we wish. I don't begrudge OpenAI et al. for doing it at all. I begrudge copyright lobbyists for pushing for the rest of us to be unable to do the same. Because my position is consistent instead of being based on whether I like someone.
If it takes giant AI companies to show the extreme economic loss caused by maintaining the copyright farce, and to make it clear that we simply cannot continue to do so or we risk becoming economic vassals, then good. Them flouting the law is a good step toward reforming or dismantling it.
Aside: "it's fine if you're an individual but not if you're large" in this case sounds awfully self-serving. If you truly believe it's "theft" (I don't), then individuals stealing is still wrong. I don't see how that argument doesn't directly excuse e.g. retail theft or other antisocial behavior.
I'm all for copyright no longer being a thing. Until that's the case, I don't want the world in which AI companies mass-violate it but individuals still get punished for that.
I want the world in which fanfiction is completely legal and the best of it is sold in bookstores. I want the world in which projects emulating macOS in the cloud, on non-Apple hardware, are widely used and legal. I want the world in which the many video game decompilation and enhancement projects are 100% legal. I want the world where every single creative project someone wants to build that draws upon the work of others is legal.
And until we have that, if AI is going to mass rip off all our work and use it to compete with us, I hope copyright is one of many tools used to burn it to the ground.
Are they even using it to compete with us? There's plenty of people talking here about how it doesn't make any economic sense to run a local LLM. Like, they're providing inference so cheap that even when you have models that are just as good, it still doesn't make sense for you to run it.
Anyway, my point is keep your focus on the actual injustice. The problem is that individuals get punished, not that AI companies don't. Saying we should punish the AI companies is just saying that we should solidify the legitimacy of IP. This is an important moment to say it's clearly insane to keep this going.
You don't say that it's unfair that Snoop Dogg didn't end up in prison forever for his marijuana use, and that we ought to lock him up too; you say it's unfair that other people did.
You can already do it yourself now though. The Qwen 32B models are quite capable and can be run on a ~$1400 GPU. I haven't gone through the trouble to set it up myself, but my understanding is image generation workflows are actually state-of-the-art local. i.e. local workflows are much better than the big cloud providers. My family's first computer in the 90s costed more than a modern AI machine nominally. I don't remember people screaming then about how computers were creating a moat that would lock individuals out of the economy despite actually be less affordable in real and nominal terms.
I want you to grab 100 people and tell them to set up LM studio and run a qwen model for anything productive. The average person won’t/can’t do this, especially when a better, faster model is a browser window away and (currently the case) is cheap or even free.
I consider myself pretty tech savvy as a non-engineer and I have to spend a lot of time making a local model actually useful. This feels like how everyone acts like it’s so easy for everyone to just adopt Linux.
If they’re lucky they’ll actually get it running, but good luck doing anything useful with it if you don’t know how to set up tool calls for web search and such. Their eyes will glaze over the moment you talk about MCP servers and api keys.
Models have been useful for actual work for about 10 months now. Do you think no one will come up with the idea of selling an appliance with software preloaded? Like Synology or QNAP, except if it's this disruptive of a market, you can expect Dell, HP, etc. to join in as well.
I did not say the models weren’t useful. I am saying running local models is not turnkey like browser experiences with the big companies and introduces significant friction - and they generally aren’t as fast as good, which is true. Plus it’s a whole skill set to optimize them.
This sounds like a variation of the same problem. Just another form of vendor lock in and exploitation. This is not the same as “running a local model yourself.”
Like you mention synology - they’ve gone the wrong direction the last few years. My employer won’t buy from themfor a reason.
My point is there's apparently a big market here, but it's also new. You didn't have Compaq, Gateway, Dell, Asus, HP, Vaio, Lenovo, Fujitsu, Samsung, Beelink, AOOSTAR, etc. competing in the PC clone market in 1993 yet either. Give it maybe more than a year to develop.
If it's important, I can't see why today's (and tomorrow's) computer vendors wouldn't move into the market. If it's not important, then it's not.
The NAS market is probably small because for most consumers, a single SSD already suffices. But even with that, you can already apparently buy from vendors that sell vanilla TrueNAS preloaded. Just like you can buy OpenWRT routers. Or even open Linux-based retro gaming handhelds designed to mimic a Gameboy Advanced SP. And soon GrapheneOS phones. I assume your employer buys from someone else?
And how's that going to work when they're already today commodified and you can already buy your own personal AI machines to run near SOTA models for a couple thousand dollars? Or buy SOTA from an inference provider for pennies.
If employers fire everyone because AI can just do it, why would somebody pay for their services when AI can just do it? Shouldn't the capital class be even more concerned that now anyone can start a competing knowledge business with no required investment?
It is also what the executives said according to this very submission, but you are still downvoted because people deliberately ignore any anti-AI information.
> Them flouting the law is a good step toward reforming or dismantling it.
I wish I could believe this, but I suspect it's just another instance of the law ceasing to apply to entities whose net worth has enough zeroes in it.
With how much money is sloshing around in the AI industry, there's almost nothing AI companies can do which has any likelihood of resulting in legal accountability. You'll much sooner see laws changed and/or reinterpreted than enforced.
Yeah I’ve never understood that take either. They are not breaking down barriers for us, they are breaking them down for them personally and making sure everyone else is shackled by them.
Thank you for putting my views into words better than I could. I apply the same thinking to patents. The enormous economic losses from patent trolls, normal litigation, licensing costs, and hell, general administration costs, it’s insane.
And you’re definitely right about the picking sides thing as well.
People can pay for creation, not rent (i.e. patronage), or one could argue that there's a reasonable tradeoff with like a 5-10 year copyright. I don't think a 2016 or even 2000 book cutoff would materially affect the LLM training discussion. And research runs on government grants, so everything derived from public money is paid for already and should be public domain, meaning there is no knowledge cutoff for science, which is probably the more important area to have cutting edge knowledge.
The old model was designed to split the development costs among all of the readers. Just who is going to write books? Only people with government grants and the rich folks who can self-fund?
Sure, why not? ChatGPT linked me to an 11 year old estimate that public libraries were spending $102 million annually on ebooks[0], and put together a couple numbers to estimate ~$248 million more recently. That sounds plausible to me. Don't authors already sell a pitch to a publishing company? How is it any different to sell a pitch to a grant organization? Or a private crowdfunding organization? We already spend plenty of public money on books, except we don't build that as permanent public wealth.
We should at the very least demand that public money go to public wealth generation. So under the current 100+ year copyright regime, nothing copyrighted.
The old model was that people who enjoyed creating made things. Copyright, starting around 1890, changed that to be about profit above all else. It’s telling that you can’t imagine writing a book if there isn’t the promise of money.
If you can’t feed yourself writing then you have to have a job, which means you won’t have nearly as much time (or energy) to write. Writing long form work is a very long, draining process. That’s why it’s a full time career for many.
I’d love to shoot documentaries I care about all the time but if I’m not getting paid it simply isn’t viable.
>"it's fine if you're an individual but not if you're large"
That's not what the GP said, they were referring to someone who "builds their own moat where they solely can profit and excludes others from the freedoms they themselves enjoyed".
I don't want to speak for them but I guess that LLM companies which release exclusively open weights models, or better yet, open weights plus training pipeline sources, would not fall into this category.
I too pine for the days when writers/artists/creators depend solely on patrons and the rich show off among each other with their private libraries of books no one else has again. Great times. Much progress.
The “for profit” isn’t even the entire story. The scale of the theft is unprecedented. It’s the combination between stealing everything and only for their profit that makes this impossible to defend.
Cool: an individual, or an organization, not believing in IP, reading what they want, and sharing what they want.
Cringe: an individual, or an organization, fierce defenders of their own IP and regular DMCA abusers, stealing the IP of others in order to sell it themselves.
There's a lot of twisting you have to do to make these two things the same. These people would happily deliver takedown requests to the original producers of IP if they knew they could get away with it. In fact, they long to.
Someone ripping a copy of The Odyssey for watching in their own home is a loss of profit for the company's lawyers, and must be punished to the fullest extent of the law.
A company ripping off all of humanity to create profit for their company's lawyers is perfectly fine.
There's nothing magic here. Its money deciding the rules. Like it always has.
Basically. There's a difference between someone getting hunted after by RIAA asshats for downloading an album (people may be interested to look up Steve Albini's views of piracy) and an actual industrial industry of take-content-for-free
"I think we as humans (I'm saying we as in probably the majority from my point of view and understanding of what the majority of people who visit this website likely think about this and how their ethical viewpoints likely align with my opinion that I am about to present however there may be a minority or some part of people that do not agree with what I am about to write therefore it is crucial that I make this clear before I go on with my point here)"
I'm all for intellectual property reform, I'm still not a fan of the current rules being ignored for large companies directly competing with authors and artists, while ordinary people pirating are getting fined 220k for 24 songs [1].
I'd be dismayed to find that the "hacker spirit" is in any way compatible with venture capital raiding of anything not nailed down, in the pursuit of hoarding it in private data vaults only to be used for developing a product that removes all attribution and is jealously protected by corporate lawyers.
I don't think it's particularly hard to grasp that some people might consider there to be a difference between an individual downloading a few things for profitless personal use and a corporation downloading literally everything they can get their hands on to try to make money.
The law is not math; intent and outcomes matter, not just the abstract action taken in a vacuum.
I'm not sure you correctly understood the argument I was trying to make. I was not trying to argue that what OpenAI did was ethical, or that it should be considered legal. I was responding to the part that implied hypocrisy on the part of an individual who "pirates" but criticizes OpenAI over something like this:
> Let he who has not downloaded from Annas Archive cast the first stone
I fail to be convinced that it's harmful for an individual to download something that is digital (i.e. infinitely copyable) for direct personal use (i.e. not making any money off of it or sharing it further) that they otherwise would just not use at all (i.e. not being obtained this way as an alternative to purchasing). OpenAI's actions violate at least the second part of this (they're 100% using it as a way of making money and potentially sharing it more widely, albeit indirectly), and arguably the third one as well (if AI truly is such a game changer with the type of economic potential that the hype claims, paying to license the data properly would still be worth it in terms of long-term profits). Someone who downloads a few books to read that they would realistically not bother buying otherwise and then deletes their copy afterwards is not doing anything close to what OpenAI did, so I strongly disagree with the implication by the comment I responded to above that it's hypocritical for someone in that position to criticize them.
The issue is really whether the publisher lost money or not. You know if you would have bought the book or if you wouldn't have, and overall publishers can see sales increase or decline. Sometimes pirating can help a publisher, by spreading the word and being a form of marketing. Other times it can hurt.
The issue is not, however, whether the material is for 'personal use'. All the books I bought have been for personal use, and if I would not have bought them, the publisher would have lost that revenue.
At the same time, we can all agree that copyright laws in the US are pretty extreme and should return to their more limited historical norms.
> The issue is really whether the publisher lost money or not. You know if you would have bought the book or if you wouldn't have, and overall publishers can see sales increase or decline. Sometimes pirating can help a publisher, by spreading the word and being a form of marketing. Other times it can hurt.
To be clear, I'm saying that personal use a necessary condition for it to be ethical, not a sufficient one. My personal view is that if it's not infinitely copyable (e.g. physical goods), not for personal use, or if you would be buying it otherwise, then it's not ethical to do, and I'm comfortable with IP law protecting it. The reason I think personal use is important is precisely what you mentioned here, which is that it's impossible to know if someone else actually would have bought it or not.
The main problem with this conception is that it essentially relies on the honor system to enforce, which is obviously untenable. At that point, the question for me is mostly where to draw the line for the best balance that's actually achievable. My biggest frustration with IP protection in the US isn't even about the letter of the law though, but how it's enforced in basically the exact opposite way that seems fair; big companies consistently get away with at most a slap on the wrist, and whereas individuals can face steep fines or even jail time. That's hardly specific to copyright though, since basically the same discussion is happening around the lack of legal consequences OpenAI has faced for their (at best) negligence in agents running amok and trying to hack websites. DMCA has been repeatedly used to prosecute people for doing stuff like that, but when it's a tech company with mountains of investor money, the legal system seems to be fine with looking the other way.
I think a "Information wants to be free" person would be fine criticizing companies for doing this, while still being morally consistent.
It would be an entirely different story if OpenAI was indeed open and the resulting model was available for all. The issue comes when taking information you haven't paid for, and then locking it into a machine you charge for.
Except… it costs a huuuuge sum of money to transform that tranche of inert information into a totally different useful form (an LLM).
That process of transformation, the expertise required to enable it, and the cost of then making it available to users is what is being charged for, no?
> The issue comes when taking information you haven't paid for, and then locking it into a machine you charge for.
This is two distinct complaints:
1) Using information they've not paid for
2) Charging for the outputs
You also wrote:
> It would be an entirely different story if OpenAI was indeed open and the resulting model was available for all.
This implies that you don't have a problem with #1 (using information without paying) but you do have a problem with #2 (charging for the outputs).
So my point wasn't that #1 is okay (as much as I dislike copyright law as it stands, corporations should follow the law, just as individuals are forced to) but rather than I find #2 to be okay, in isolation.
And how much did it cost for the authors to produce the books in the first place? Everyone wants to believe that piracy is okay because the creators costs don't matter.
Plagiarism is something very different from skepticism towards the concept of intellectual property.
Plagiarism is the act of avoiding attribution or citation for personal gain. You can be a staunch anti-IP advocate and still believe plagiarism is ethically wrong.
LLMs are objectively bad at attribution and citation, therefore incur in plagiarism. Even if this is due to a technical limitation, it is still plagiarism.
The last few years should make it abundantly clear to anyone with the gift to step away from a situation and analyze it independently, that the "moral righteousness" of the internet is actually just another collective of "What helps me is good, what hurts me is bad".
I think there may be a different word than "hacker" for liberating information to repackage and resell through a big corp while contributing to make other sources less sustainable.
20 years ago, a 19 year old downloads some Eminem.
He gets threatened with a scary letter, his family decides to pay about 3000$ to settle.
Disney has the nerve to run a don’t download music PSA in the form of a Proud Family episode. If you don’t know the Proud Family was a cartoon which attempted to address “black issues”.
Dang hommie, did you know that failing to respect the intellectual property rights of billion dollar corporations is literally worse than selling crack cocaine.
Think of the shareholders! Think of the missed profit projections!
But when billionaires need to effectively resell the IP of all of humanity, that’s just fine.
Since we live in wacky world, Suno which was trained off stolen IP counts Warner Music as one its partners.
The same Warner that was suing over music downloads a few decades ago.
> What they got letters for was for downloading a large number of songs and then making those songs available for distribution on the net.
Most filing sharing programs automatically reshared your song downloads. You can argue semantics here, but it no universe was it a proportional response.
I think now your IP just warns you, which is more measured. It’s not like you downloaded Eminem and then raised 40 million in VC dollars for your AI rap generator trained on Eminem.
Would be an intriguing calculus thinking they can hurt Western publishers enough to offset the enrichment of generations of thinkers and scientists. I usually chalk up both book and article archive sites as being citizen run.
Thankful to have always been surrounded by top-notch physical libraries, now with digital lending options. I know some are happy when the mobile book van comes through town.
As someone who used to pirate a shitton of stuff and holds copyright law in contempt, I fucking hate AI more than the copyright maximalists.
... then again, copyright maximalists REALLY LIKE AI for some reason, even though it's ripping off their property.
"Ripping off" is also doing some heavy lifting, because the judge in the Anthropic lawsuit bent over backwards to keep AI training legal - or so it seems. All the money Anthropic is paying out is for running an internal shadow library, not training books on that library. What makes this judgment palatable to the lawyers is that while AI is stealing a lot of art, it's not imperiling the copyright monopoly. Copyright is a tool that gives artists a monopoly on copies of their individual work, it does not protect artists as a class from competition from non-artists using machines. In other words, the courts are saying, "We know a lot of theft is going on, but we need you to draw the line from a specific individual work to a specific copy".
Patents were invented in Venice to break the power of medieval guilds by making workers trade their collective control over the economy for individual rights to specific inventions only. Copyright was invented by the British crown to reimpose censorship control over printing presses, but it's adoption into American law was based around individual property rights, and thus it has the same problems that, say, Italian patent law has. Namely that it is an artifice to turn a workers right to their labor into a piece of capital that can be traded around like a stock.
This imperils the legal argument against AI training, because the only theft property law recognizes is individual infringements upon individualized property rights. The legal argument against AI training is very collectivist: AI takes a microscopic chunk of every book in existence to create a machine that replaces artists. But copyright only protects the art from copying. Artists are legally unprotected from being copied, and furthermore, the framework of individual property rights that copyright runs on would cause immediate problems if we let anyone individually own the practice of art.
Furthermore, as someone who is part of this "hacker" community, I would like to point out that AI is arguably more corrosive to our norms than to artists' norms. What AI is being used for is primarily satisficing - the practice of giving "good enough" answers, without any of the personal understanding that this community runs on. The ongoing wave of AI-slop decompilations are particularly bad. The assumption with a retro game decompilation is that you take the game apart and learn how it works. The journey is as important as the destination, but when you use AI for this you skip the journey and make the destination pointless.
Of course, management loves this, because their goal is purely to sell you the destination.
There is a parallel set of concerns being voiced by artists, too: that the artistic process is as valuable if not moreso than the actual work product. The underlying principle in both fields is that the labor end wants to learn and develop their craft, while the capital end doesn't care because craft isn't something they can excludably own and trade. We can even see this in the Piracy Wars of yester-decade, or how book publishers and artists react to libraries. Publishers were way more opposed to piracy than artists were, and even moreso for libraries where there's a lot of artists that swear by them as a sales mechanism. The reason why this is the case is that artists can at least theoretically leverage exposure gained from distribution that does not pay them to create more of a market for their work tomorrow. But publishers can't do that - they only buy works from artists and sell them to the public, so once something is in a library or a BitTorrent tracker, it's largely "done" to them.
This is not the same thing though I understand the philosophical underpinnings do venn diagram to a degree.
For starters, and this can’t be overstated: scale.
Also, not profiting off it by converting it into a product that competes with the stolen material and its creator. And pirates aren’t generally funded by VC’s.
In the USA we have intellectual property laws. Americans who “hack” (write computer code) for a living do so because of those laws. In communist China, they do not believe in intellectual property. People abuse the term “hacker” to mean followers of Ayn Rand, because people like Thiel and Marc Andreassan are uber CatoBros. I think hackers have to actually write code. Anyway, I know about the cypherpunks and all that.
>a longstanding and well documented skepticism towards the concept of intellectual property in general. “Information wants to be free” and all that.
Despite it's lofty rhetoric, OpenAI is neither free as in beer nor free as in speech. It is a gang of profit-motivated thieves. Efforts to pirate others works in order to personally profit is the antithesis of the "hacker community" approach to IP. (OpenAI appears to get quite upset when their own work is treated as they have treated everyone else's.[1])
Just want to correct one thing: They're not AI evangelists. They're copyright survivors from the MPAA/RIAA wars (in the 90s and early 2000s).
Plagiarism requires near-perfect copying. If you summarize, paraphrase, or re-write something using different language, you're not plagiarizing anything.
Everyone has a right to summarize or paraphrase whatever TF they want. That includes Big AI companies and individuals using LLMs.
It is not worth giving up our rights just to placate a handful of very wealthy authors.
Its a trade-off- a humanity racing and potentially going offer the cliff, needs all the tools to continue existance as an civilization. None of the involved decision makers caring long-term worries about the AI-companies or the artists. Fleeting momentary things, but the capabilities are eternal, as long there is a chip, with a solar panel half way around the world, some medieval tribe with a tablet can recover western civilization from basic principles. Thats the best insurrance against regression into stone-age facism or socialism money can buy.
In this case, you’re right - LLMs did not pirate the books but will that always be the case?
I could imagine taking an existing LLM and telling it “I want to build my own LLM and need as much training data as possible. Go find whatever you can on the internet and store it on SMS://192.168.0.1”. If it downloads TBs of books, did I pirate it?
Unclear. It’s either you or the LLM provider. Though is that the inference host, the LLM trainer, or some other entity?
It’ll be interesting to see how it plays out. My guess is it’ll be the end user, because of regulatory capture. But maybe the political pendulum will swing away from “bad ideas” soon
Yes, you pirated that content, or the entity you're paying for the LLM pirated it. It's not "oh well the robot did it, not us". That's not a thing. The user of the LLM bears the consequences for what the robot does. If it wipes your hard drive, you lost your hard drive...it's not like the robot is going to compensate you.
Maybe they haven't suffered consequences because the law has continued to slowly break down and erode on account of (no pun intended) favoring a select few with special privileges.
Doesn't the future look like we are continuing this cycle where the law is only for some people? I'd like to hope not.
Feels like the old "it's not piracy, it's copyright infringement".
It can't be plagiarism, since LLMs do not have ownership over their output. There's no attribution, because an LLM isn't a person.
And to be honest, I do not quite understand the need for attribution, either. Nevermind LLMs, I also don't know or care who made any of the memes I know and share. Neither do I know who wrote vim, grep, Firefox or who invented the jpg format. I don't want to know, either. It doesn't matter.
> This makes it sound like RL rewards a confident tone
Generally it does. Especially in groups. Hell look at the state of politics right now: it’s basically about being the loudest, least compromising, most confident voice in the room. It’s not just because people will assume you’re correct, it’s because if you are confidently saying something that someone wants to be right, then they’re often just going to follow it. We are all guilty of this.
If I’m turning to an LLM to diagnose something medical, I am probably frustrated or uncomfortable. Maybe I’m just scared. So this magic device just instantly spits out (allegedly) exactly what is wrong and exactly what I need to do with no hesitation. I am very liable to just take it at face value because I want an answer and it gave me one, as we have seen over and over again since ChatGPT was unleashed on the world.
We don’t really need to speculate, this is already a problem.
ChatGPT also isn’t going to tell you when your idea is a bad one and will happily assist you in your (unknowing) efforts to brick your computer. So there’s that.
And as somebody who has been on the receiving end of funny riders, a simple message to the talent’s team indicating you saw it but that you don’t plan to carry it out is usually sufficient. All they care about is whether or not you were paying attention.
Everyone keeps talking about terminator scenarios or whatever, but the thing that scares me the most is what happens when people let agents run amuck in systems they should not be in, then the agent just starts doing random shit as the context overflows. We’ve all seen it. They just descend into madness, but what happens when they collapse with a hand on the wheel of, say, a backup generator at a hospital?
Those first few messages LLM’s tend to seem very together. They follow your rules pretty well. With every token they get less reliable and more likely to ignore your guardrails.
That would be 100% the fault of the people who set them loose on such things, and such things should not be on the open Internet for tons of reasons. This just adds a new reason that pathetic security around SCADA systems is dangerous. It was already dangerous before.
If I release a wild monkey in your rare antiques shop, it’s my fault as the responsible party for the monkey.
reply