Hacker Newsnew | past | comments | ask | show | jobs | submit | thenewnewguy's commentslogin

This isn't black and white, it's a sliding scale. Advertising, menu boards, box art, etc. do show the product in the best light possible and take artistic liberties, but are substantially closer to the real product than an AI that's never seen the product spitting out an image of a burger.

It's reasonable to be annoyed when the menu board _isn't based on the product at all_ and thus is nowhere close to the real thing.


If I got served food at a restaurant and realized that the pictures were from a different restaurant, I would be pissed off in exactly the same way and amount as if a menu used AI slop. The only difference is that with AI slop you can tell ahead of time and decide not to spend money there.

Agreed - not only does the obvious LLM writing/editing make it hard to read, but the weird stylistic choice to have long sections where every single short sentence is its own line made the read even more grating.

I usually like blog posts about homelabs but came out of this one disappointed.


Comparing Norton selling their company (even to a scummy owner like Symantec) to Pétain collaborating with Nazi Germany feels a bit ... extreme ... to me.

Yup, that's exactly what I said. Quoting myself verbatim:

> Norton selling his company is exactly like Pétain collaborating with Nazi Germany

I wish I didn't say that, but what can you do.

Sometime y'all need to probably chug a glass a water and re-read the comment you're about to reply to.


Pétain was a soldier with what, 50 years of experience when he made that choice?

He picked the wrong side obviously, but is it also possible he did what he though was best for France? Both things can be true.


Yeah, I don't doubt that he did what he thought was best, I think most if not all humans do that most of the time. But I do think the general sentiment about him did shift a bit after he made that choice, which was the point.

I fully agree with you, and I think this trend of glorifying disabilities is cringe - ADHD, autism, etc. are life-altering medical conditions and not desirable.

However, I don't think this specific project is intending to glorify ADHD or help people claim they have it - it's just piggybacking on the idea that telling current-gen LLM models that you have ADHD (allegedly) produces better results for everyone.


It's open source software - it can be baked into pretty much anything given they follow the license terms.

The question is whether the 10 million active users of Codex / ChatGPT Work count towards the download count.

That would contradict the cause implied in the title here.


Does anybody have a source for the "actively exploited" part of the HN title?

by nature of being in the "known exploited vulnerabilities catalog" (https://www.cisa.gov/known-exploited-vulnerabilities-catalog...)

"CISA maintains the authoritative source of vulnerabilities that have been exploited in the wild."


This line, I think? >This CVE is in CISA's Known Exploited Vulnerabilities Catalog

"Google has confirmed that an exploit exists in the wild but has not disclosed information about the threat actors, targeted organizations, or attack campaigns while the update is still rolling out."

Technically yes, in practice the odds your local resolver is validating DNSSEC is slim (and if you're intentionally configured it to do so, switch to a provider that isn't Quad9).

If you care about software being closed source or proprietary you probably aren't running Windows.

Honestly, even as a Linux guy who doesn't use Windows (almost) at all anymore, I trust Microsoft to be generic and corporate in their spying.

They'll spy on users usage habits, and likely backdoor their OS for the feds, but they won't individually exfil your browser cookies and steal money from your bank account.

A random closed-source black-box "privacy" tool though? You have zero guarantees.


> They'll spy on users usage habits, and likely backdoor their OS for the feds, but they won't individually exfil your browser cookies and steal money from your bank account.

What’s the difference between backdooring for the US gov’t and directly robbing/bugging you (assuming for the sake of argument that they don’t already do at least the latter)?

Used to be that you’d only ask that question if you were brown and/or write your name from right to left. Doesn’t matter much whether you were from Svalbard or Sanaa these days…


I think they are making the assumption that if such a govt backdoor existed, it wouldn't be used for generic mass spying, but instead targeted at specific individuals they are trying to find/arrest.

That may or may not be true, but I feel like the former would be a whole lot harder to hide, and the risk of having to reveal it in court would likely mean they aren't going to blow their cover over someone who wasn't worth it.


>Generic mass spying.

Like PRISM, Boundless Informant, Stellar Wind, etc?


It's a typo that carried over to the second code example of trying to read the file too? Is the sample code completely untested?


the example is fixed, thanks for spotting that


Bias disclaimer: Amazon is my current employer, but I don't work on AI or anything else mentioned in the article.

Yes, this is a result of copyright laws. The other commenters are wrong/uninformed.

If it was up to the companies training LLMs, they wouldn't destroy the books: It's a waste of company resources, it's needlessly destructive/evil, it generates bad PR, etc etc. There are essentially zero advantages, other than it is what is required under US copyright law (or at least, it is what their highly paid lawyers believe is required under US copyright law).


Isn’t it being destroyed because it makes the scanning process easier?


Perhaps partially? I assume that cutting the pages out of the spin makes them easier to scan at least partially. That said I have no insider knowledge of this type of operation so I don't know how much easier that actually makes it.

But ultimately, it's a moot point, because the legal requirement means the books must end up destroyed. Even if the people at Amazon wanted to scan the books in a way that required no destruction at all, it's not currently (legally) possible for them to do so, so they might as well take the easy way out today.


Non-destructive scans are at least 10x as much, and tend to lower quality.

https://software.annas-archive.gl/AnnaArchivist/annas-archiv...

Cutting the pages out makes them machinable. Non-destructive scans involves gently turning pages, and paying a lot of attention to the state of the spine. Destructive scans involve guillotine cutting the spine off, scanning the covers by hand, putting the pages into a hopper, clamping them in and hitting a button. While that book is scanning, you're already cutting the spine off the next book. If the machine jams, try to work the jam out gently, scan the pieces, and let the computer stitch it together.


Perhaps this is a use case that's worth investigating, to invent better, less-destructive scanning processes! Or some way to rebind them afterwards.


No.

A judge a while ago decided that as long as the physical copy is destroyed, and "transformed" into an electronic copy, you can do the upload. But if you preserve the physical copy after scanning it, you are in violation of copyright because you "copied" the book.

That's literally the only reason they are trashing them. It's a legal requirement.


Can you find a citation for this? I have heard this claimed rule recently from other people, and I haven't seen this decision (nor do I know what level of court or jurisdiction it might be). This is not a rule that I heard many years ago when working on and adjacent to copyright issues (including book scanning!), although of course the issue has been newly litigated again recently, so there may be new interpretations coming out.

Edit: Someone else linked to an order in Bartz v. Anthropic which appears to emphasize that destroying the original copies improved the defendant's position with respect to the fair use analysis. Is that the decision you're thinking of?


Ah, so I can buy and scan a dvd, destroy it and then legally distribute the legal copy via torrents. Good to know, because that is what the LLM thieves are doing.


Distributing an exact copy of the text would be illegal; if you could get an LLM trained on a book to output the exact copy of the text from the book then that would be illegal as well.


They still made a bunch of copies and reused it to train multiple models after making the first copy. Unless they copy and destroy a book each time they use it as as training sample


They could just not do the evil thing.

(This is why I will never be a billionaire)


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: