This isn't black and white, it's a sliding scale. Advertising, menu boards, box art, etc. do show the product in the best light possible and take artistic liberties, but are substantially closer to the real product than an AI that's never seen the product spitting out an image of a burger.
It's reasonable to be annoyed when the menu board _isn't based on the product at all_ and thus is nowhere close to the real thing.
If I got served food at a restaurant and realized that the pictures were from a different restaurant, I would be pissed off in exactly the same way and amount as if a menu used AI slop. The only difference is that with AI slop you can tell ahead of time and decide not to spend money there.
Agreed - not only does the obvious LLM writing/editing make it hard to read, but the weird stylistic choice to have long sections where every single short sentence is its own line made the read even more grating.
I usually like blog posts about homelabs but came out of this one disappointed.
Comparing Norton selling their company (even to a scummy owner like Symantec) to Pétain collaborating with Nazi Germany feels a bit ... extreme ... to me.
Yeah, I don't doubt that he did what he thought was best, I think most if not all humans do that most of the time. But I do think the general sentiment about him did shift a bit after he made that choice, which was the point.
I fully agree with you, and I think this trend of glorifying disabilities is cringe - ADHD, autism, etc. are life-altering medical conditions and not desirable.
However, I don't think this specific project is intending to glorify ADHD or help people claim they have it - it's just piggybacking on the idea that telling current-gen LLM models that you have ADHD (allegedly) produces better results for everyone.
"Google has confirmed that an exploit exists in the wild but has not disclosed information about the threat actors, targeted organizations, or attack campaigns while the update is still rolling out."
Technically yes, in practice the odds your local resolver is validating DNSSEC is slim (and if you're intentionally configured it to do so, switch to a provider that isn't Quad9).
Honestly, even as a Linux guy who doesn't use Windows (almost) at all anymore, I trust Microsoft to be generic and corporate in their spying.
They'll spy on users usage habits, and likely backdoor their OS for the feds, but they won't individually exfil your browser cookies and steal money from your bank account.
A random closed-source black-box "privacy" tool though? You have zero guarantees.
> They'll spy on users usage habits, and likely backdoor their OS for the feds, but they won't individually exfil your browser cookies and steal money from your bank account.
What’s the difference between backdooring for the US gov’t and directly robbing/bugging you (assuming for the sake of argument that they don’t already do at least the latter)?
Used to be that you’d only ask that question if you were brown and/or write your name from right to left. Doesn’t matter much whether you were from Svalbard or Sanaa these days…
I think they are making the assumption that if such a govt backdoor existed, it wouldn't be used for generic mass spying, but instead targeted at specific individuals they are trying to find/arrest.
That may or may not be true, but I feel like the former would be a whole lot harder to hide, and the risk of having to reveal it in court would likely mean they aren't going to blow their cover over someone who wasn't worth it.
Bias disclaimer: Amazon is my current employer, but I don't work on AI or anything else mentioned in the article.
Yes, this is a result of copyright laws. The other commenters are wrong/uninformed.
If it was up to the companies training LLMs, they wouldn't destroy the books: It's a waste of company resources, it's needlessly destructive/evil, it generates bad PR, etc etc. There are essentially zero advantages, other than it is what is required under US copyright law (or at least, it is what their highly paid lawyers believe is required under US copyright law).
Perhaps partially? I assume that cutting the pages out of the spin makes them easier to scan at least partially. That said I have no insider knowledge of this type of operation so I don't know how much easier that actually makes it.
But ultimately, it's a moot point, because the legal requirement means the books must end up destroyed. Even if the people at Amazon wanted to scan the books in a way that required no destruction at all, it's not currently (legally) possible for them to do so, so they might as well take the easy way out today.
Cutting the pages out makes them machinable. Non-destructive scans involves gently turning pages, and paying a lot of attention to the state of the spine. Destructive scans involve guillotine cutting the spine off, scanning the covers by hand, putting the pages into a hopper, clamping them in and hitting a button. While that book is scanning, you're already cutting the spine off the next book. If the machine jams, try to work the jam out gently, scan the pieces, and let the computer stitch it together.
A judge a while ago decided that as long as the physical copy is destroyed, and "transformed" into an electronic copy, you can do the upload. But if you preserve the physical copy after scanning it, you are in violation of copyright because you "copied" the book.
That's literally the only reason they are trashing them. It's a legal requirement.
Can you find a citation for this? I have heard this claimed rule recently from other people, and I haven't seen this decision (nor do I know what level of court or jurisdiction it might be). This is not a rule that I heard many years ago when working on and adjacent to copyright issues (including book scanning!), although of course the issue has been newly litigated again recently, so there may be new interpretations coming out.
Edit: Someone else linked to an order in Bartz v. Anthropic which appears to emphasize that destroying the original copies improved the defendant's position with respect to the fair use analysis. Is that the decision you're thinking of?
Ah, so I can buy and scan a dvd, destroy it and then legally distribute the legal copy via torrents. Good to know, because that is what the LLM thieves are doing.
Distributing an exact copy of the text would be illegal; if you could get an LLM trained on a book to output the exact copy of the text from the book then that would be illegal as well.
They still made a bunch of copies and reused it to train multiple models after making the first copy. Unless they copy and destroy a book each time they use it as as training sample
It's reasonable to be annoyed when the menu board _isn't based on the product at all_ and thus is nowhere close to the real thing.
reply