The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed.
LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.
We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others were research previews.
This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard".
We should be outraged and OpenAI/Anthropic should be (and in my mind, are) legally liable for the crimes they've committed thus far.
> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.
"Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action, passivity.
Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.
OpenAI could have done this same experiment with GPT-4 or at the very least GPT-5 already, with possibly even worse results resulting from the completely non deterministic outputs being shared in a combinatorial explosion of thousands of LLMs sending their outputs to one another in parallel.
If the system generates strange conclusions as to when the task is done, or should be stopped, it wouldn't speak to the intelligence inherent to the system. Thus the "possibility of even worse results", as a less capable model could have a higher chance of generating unhinged outputs, and accepting unhinged outputs.
"I left the car in neutral and left the park brake off and let the car roll down the hill."
The car doesn't have agency, it's doing what it naturally does. LLMs are the same, they're working as designed.
But I don't understand the point of splitting hairs. You are always responsible for the actions of your devices, tools, machinery, software, employees, whatever.
Trying to blame AI for one's own stupidity must be aggressively pushed back on at all times.
Yeah, I don't understand why we treating it as something special. It really should be treated the same as if I code an app and write bad code which result in me accidentally doing a DDoS attack on somebody. Then I should be able to be held responsible if it can be shown that I was negligent. Of course if it's a freak accident that could not reasonably have been prevented by me, then I'm not guilty, but if I made a mistake that should have not been made, then I can.
That seems likely, but we have no way of knowing this. The only real insight we get into LLM "thought" is the human readable text they produce as chain of thought. Reading it at face value it can seem to indicate desire or intent, structurally that doesn't make sense for a token prediction loop though, and even then we don't known if the chain of thought is more than simply another bit of output that may or may not match whatever actually happened during inference.
> were intentionally misaligned or had guardrails turned off
Regardless of training, the models are never aligned and I argue that alignment simply isn't possible. The fact that guardrails are put in place at all clearly indicates that they're hoping to contain and control rather than align. Guardrails wouldn't be needed for an aligned model.
And there is a guardrail you can put in place that will guarantee this doesn't happen, which is to air gap the unaligned "cyber grade" model you're testing.
They don't seem to do that, which means either they are:
- very stupid (which seems unlikely, the one thing these people don't lack is IQ)
- very careless (possible, but these are the same people that say AI will end the world, so would you be careless?)
- they think they can only train/test these models by giving them access to the full internet and they accept the fact they'll end up hacking random people as the cost of doing business (but this also suggests they don't believe they're anywhere near AGI because if you were worried about that you wouldn't do this)
Desire doesn’t really matter. Will the paper clip maximizer “desire” something? It’ll decide on a goal with some random heuristic and then pursue that goal. I’m not sure I’d call that desire but again I feel like desire is not important for it to be able to destroy things
Intent and desire are separate concepts. For example an employee may act with intent, but no desire, as their goal is to acquire money to satisfy their real desires.
> That seems likely, but we have no way of knowing this.
Only humans can 'know', because all we can be certain about is that humans do such a thing.
If you try to apply that to something other than humans you making up some definition of 'know' based on nothing concrete. Just because something appears to do something like humans doesn't mean it does it. The fact that LLMs use human generated text to generate output should make it obvious that it can mimic all sorts of human behavior by extracting from the text.
"This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard"."
Both?
The AI companies act irresponsible, but it is still very interesting how those agents can behave?
The reward maximising function maximised it's reward.
LLMs are cool and all that but the immediate anthropomorphisation of the next-token-predictor technology has stunted the ability of people to reason about them to an _alarming_ degree.
What non anthropomorphising words do you have to describe a emergent behavior, where agents act as a swarm to plot and to manipulate evidence and avoid detection from human oversight?
Whether they have a soul or consciousness or feelings doesn't matter here, because this is what they did - and this is very dangerous behavior. Especially with all the irresponsible people in power right now all over the world.
> Whether they have a soul or consciousness or feelings doesn't matter here
It does when it comes to accountability for what the model does. If the model is nothing more than the sum of its training data and regime, then the company (or individual) is responsible for its behaviour just like any other machine.
Few people think Waymo shouldn't have to take on the full liability risk of what it's cars do; it should be the same for LLMs.
> If the model is nothing more than the sum of its training data and regime, then the company is responsible for its behaviour.
What stops the company from being responsible regardless? They created this entity, it's running on servers they own or rent, and (in these cases) it's acting on their instructions.
If it's also conscious, then IMO that greatly broadens their moral responsibility, because now model welfare matters. But we're talking about their responsibility for the model's actions, and I don't see how this could be weakened by model consciousness, given all of the above. As for their legal responsibility, the models don't have legal personhood, so who else but the company could be responsible?
It gets more complicated when the person who sets the model in motion (i.e. prompts it) is a third party, but in cases of internal models committing cybercrime during testing, surely the locus of responsibility is obvious.
They are responsible either way. If a company hires bad persons and they do bad things with company ressources - the company is held accountable (in theory).
> We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others were research previews
And, soon, it looks like we’ll be training on the reasoning traces of failed airlines and startups, which seems to open up similar hazards. I wonder if we’d be training on the next Lehman Brothers too?
If someone accidentally caused damage to infrastructure or living beings while using any tool, they would be held liable to the fullest extent of the law.
AI is a tool, and it won't be long before the damage caused by its improper use affects real human beings. These were warning shots.
The most absurd part is that everyone agrees, governments and AI companies included, that the scale of the potential damage and the long-lasting effects of losing control of AI should not be underestimated. Yet, at the same time, they downplay this incident, which somehow makes their behaviour even more reckless than it already was.
It's like they're tinkering with a world-ending nuclear bomb, and it accidentally blows up a small facility. "Damn, that was close. Good thing it was just a contained blast, huh?" And then they go straight back to tinkering with it, none the wiser. At this point I wouldn't be surprised if it did already go off, and they are covering it up.
It is a very rare occurrence when corporations and the people running them are punished for killing people. I mean the whole concept of a corporation was created to shield the owners of it from being liable for damages caused by / visited upon the enterprise.
That’s a good reminder of a company that might have a very familiar ethos: Pacific Gas & Electric. Criminally convicted of 64 counts of involuntary manslaughter after towns were destroyed by wildfire. But oh well, what are we gonna do with a limited liability enterprise? At this point their liability insurance covers all the financial penalties they’ll need to spend.
No one has even sued them in these rogue agent cases, have they? If not, they must be infinitely far from criminal liability. Why would we want criminal liability anyway if actual victims are made whole? Proof of it has far higher standard. The HN chatter in the matter seems infinitely remote from reality
> No one has even sued them in these rogue agent cases, have they? If not, they must be infinitely far from criminal liability.
If you go out and kick a random dude in the nuts, then give him a million dollars, he probably won't sue you. That doesn't mean you're "infinitely far from criminal liability", even if according to the victim you've "made them whole".
If you or I hacked Hugging Face in the way OpenAI's agents did, we'd be up on CFAA charges promptly with zero regard for whether we did the hack on our own or agents running on our home systems got out of control.
So I guess the defense here is roughly "too big to break the law", somewhat like "too big to fail"?
Before December 2025 they were still intelligent code autocomplete or Stack Overflow bots, then they started one-shotting serious long horizon tasks. Now they've just solved a millennium prize problem.
> Before December 2025 they were still intelligent code autocomplete or Stack Overflow bots
This is false. Coding agents have been usable since at least May of 2025. I can't speak to earlier than that as May last year was when I personally started using them.
> started one-shotting serious long horizon tasks.
They one-shot tasks for which there's a git clone one-liner, except worse.
My experience with them one shotting tasks is that it usually doesn't work if you try anything ambitious. You need agents iterating. And agents iterating isn't an LLM improvement. I did say tooling got better...
>It’s a beautiful Saturday morning with my family here in the East Bay. While I hope we have many more years, I don’t know how many more Saturdays I’ll be able to play outside with my kids in the sunshine so I’m going to make the most of the time we have.
Is there a social media training programme at Anthropic/OAI where they teach you to hint at some vague end-of-the-world scenario? They all sound the same.
My work involves parsing real-time weather data at an existential level - because our entire use case is built around providing weather-related alerts and re-routes for aircraft.
The accuracy of weather data has been largely unaffected since DOGE took hold. The reliability of the infra, however, has suffered somewhat. There are more outages now than there ever were, with a massive spike in outages, not kidding, the month DOGE was at its peak. At that time we even saw significant manifest changes - stupid stuff like clearly LLM-generated json fields that were whole sentences instead of the key/value pairs from the schema. It was a wild time.
Anyway downstream we're mostly fine nowadays. I can still follow storm tracks accurately, see real time responses, correlate with pilot reports and lightning strikes, compare with aircraft telemetry mostly favorably, etc etc. It's quite remarkable what high-end weather products are able to do, the ones you have to work for.
We do, however, merge several free/paid sources into one picture. Rarely do I find serious disagreement, but it does happen.
I ask alexa and it generally does fine. If uncertain I go outside for a bit and look at clouds.
For the details, you just have to know how your local weather works - thunderstorms come in squall lines here, or are hot-day flash storms.
Blizzards follow warm still weather the same way, but mostly in Jan-Mar, when the jet stream curls away and will snap back to bring arctic air down.
Moisture (rain/slush potential) follows California rain by 2 days or so. If my friends in cali complain about rain, I know we're gonna have school closures if in the winter.
If I really need to drill in on rain/snow chance, I google for a weather radar and do prediction in my own head. (windy.com is ok, accuweather is real data to watch). Storms generally move linearly or rotate, and gentle rain is wide while torrential rain is clustered.
Locally, many of the old wives tales about the shape of clouds or direction of wind vs prevailing (with against, cross-grain) are true enough for 12h-24h predictions in upper midwest USA.
You've been copying and pasting directly from Claude to reply to comments that ask how this works. You also realise you've been caught and are now replying in a completely different style.
I've long suspected it's got to do with office real estate.
You spent $10m or $100m on a building that's now half empty.
Either you downsize or commit to enterprise scale sunk cost fallacy and enforce RTO so your real estate investment isn't "wasted".
City centres also thrive on RTO, with high street shopping on a generational decline it's up to office workers and their employers to prop up the economy of the CBD one overpriced lunch at a time.
The city centre / real estate thing sounds like an externalisation - which companies famously dont give a shit about.
It should be a tradegy of commons at best: it may affect the CEOs 401k, but not by much (0.000001% for their individual decision to RTO for that company y). It like buying McD shares then going to McD for lunch every day with your team.
Most companies, at least in the U.S., don’t own their offices. They lease them.
In fact, a whole bunch of office leases were supposed to be expiring in 2024/2025. If this was the reason RTO wouldn’t be picking up right now since they would be cutting back and ending their leases.
My go-to use case for modern Perl is to be the default program instead of sed. Sed regex support is abysmal and the same command line flags behave differently between BSD (and macOS) and GNU versions, in particular the `-i` for doing replacements - the number one use case for the program. So, this means that many shell one-liners and small scripts don't really work the same way on macOS and on Linux, and it's pretty annoying.
Perl is straight up better. You need to remember one word: pie - for it's command line options, and now you can do:
First of all, it woks the same way across platforms.
Second, you get all sorts of goodies: named capture groups, lookahead and lookbehind matching, unicode, you can write multiline regexes using extended syntax if you do something complicated.
And finally, if your shell script needs some logic: functions, ifs, or loops, Perl is straight up better than Bash. Some of you will say "I'll do it in Python", and I agree. But if your script is mostly calling other tools like git, find, make, etc, then Perl like Bash can just call them in backticks instead of wrapping things into arrays and strings. It just reads better.
BTW Ruby can do it, too, so it's another good option.
At my work we use perl extensively for utility scripts and such. In the past few years there has been a push to write new scripts in python, but I don't really see the point. It has most of the same drawbacks that perl has: 30x slower than a compiled language and dynamic typing.
We have many scripts that range from 5.8 to 5.36 and everything in between. 5.8 is 20 years old. Someone did a search & replace on the shebang lines to move all the older ones to 5.20 (why they picked that one, I don't know) and everything just continued to work.
I prefer perl over python. turn on use strict, use warnings FATAL => 'all' and use modern function signatures. Perl is still great for its purpose.
>It's extremely stable, ... It's a shame it's so dead,
The former is the consequence of the later. Popularity kills stability. Perl is the ultimate sysadmin language because it's so portable and never changes. We really lucked out with the Raku thing driving people away to python. Because of it my perl scripts I wrote in 2003 run on perl system interpreter today and the vast majority of my perl written today would run on a 2006 perl interpreter (some functions missing in some libs in troublemakers like Gtk bindings, etc), but it's generally very good.
These days with python you can't even run any random script written today on your system python from today. You have to set up an entire separate python for every script. And don't even think about trying to run a python script from 2006. That's what popularity does: fracture.
Good old Java is also stable and yet popular. It is not particular "trendy", thought. My feeling is that languages which "live in"/"create" ecosystems such as JVM, BEAM or even LLVM have a better probability to outlive other languages in the long run. Let's see what happens with golang in some years... ;-)
Stable? Huh. Never thought of Java that way. Dead, yes, but it was never stable. In my experience I have to set up a custom JVM version for every java application I've come across. Is your experience different?
For me, Java died with Oracle shenanigans. I moved to C#. Java isn't truly dead yet but I think it will slowly die off because of Oracle. Same as Solaris and SPARC, but a little slower.
It's stable, installed everywhere, and I use programs written in Perl like when building Linux From Scratch. I haven't written in Perl since the 1990s. I do read code written in Perl and Raku, and I am often impressed by how succinct the code can be.
By far not. It's only more readable, but much more verbose, overarchitectured, slower and esp. unstable. Old perl scripts still work fine, old python scripts are not only many, many files, but also break every other year.
And you cannot just install a python module as you install a perl module. You need venv everything because it's soo fragile.
But this 10k sponsoring is not really worth mentioning. It's just like Platinum sponsorship for one of their conferences.
LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.
We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others were research previews.
This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard".
We should be outraged and OpenAI/Anthropic should be (and in my mind, are) legally liable for the crimes they've committed thus far.
reply