Hacker Newsnew | past | comments | ask | show | jobs | submit | Philpax's commentslogin

> It's not going to happen accidentally.

https://transformer-circuits.pub/2026/emotions/index.html

Whether these are like "our" emotions is hard to say. What we _can_ say is that they are emotion-shaped, we didn't design them, and they happened accidentally.

Modern AI is grown, not meticulously designed, and we cannot say with any certainty what the resulting mechanistic properties are.


An LLM will learn anything that helps it predict, including the emotional state of the writer - that is expected.

If you give an LLM the move sequence of a half-played chess game and ask it to continue as white or black, then it has learnt enough to model the ELO rating of both players and will continue playing at that level. It is not playing to win - it is doing what you expect and predicting as well as it can - it predicts the 1500 ELO player will keep playing at that level, and generates moves accordingly.

An LLM appearing to exhibit an emotion (if we anthropomorphize it and read emotion into it's output) is just predicting as well as it can - if the context calls for sad output, they you'd expect to get sad output and will necessarily find that "we're predicting sadness" detector somewhere internally.

Transformers are the same as they ever were from 10 years ago, other than minor efficiency tweaks like MOE and different attention mechanisms. Training is getting more and more complex, resulting in better and better cargo cult reasoning etc, but the architecture remains the same.


>An LLM will learn anything that helps it predict

I'm not sure you quite understand the full meaning of this statement. If you did, your following paragraphs wouldn't follow.


Are you imagining that an LLM tasked with predicting a game continuation is going to play to win instead?

I imagine it will learn to win under some circumstances, perhaps in a case with some context expressing a desire to win. Drawing out an LLMs upper ability in the game should be fairly straightforward.

If you asked it to try to win, to "plan lines step by step", etc, then it would do it's best to follow that instruction, but unless RLVR trained to reason about chess (easy to do, but not sure which models may have done it) then it'd have to instead rely on the chess reasoning it had seen during pre-training (post-game interviews etc), which I doubt is enough to do very well.

However, if you just ask it to continue a game, halfway in progress, then by default it will try to predict the most likely continuation, which is that both players will continue to play at the level they have done so far. This isn't a theory - it's been documented, as well as what you'd expect.


I mean sure, but I'm not sure what that has to do with the broader point. It will learn to play, and it will have a model of what it means to win.

> I'm not sure you quite understand the full meaning of this statement. If you did, your following paragraphs wouldn't follow

I was just explaining how this comment you made is wrong.


It's not wrong. You admit that LLMs will 'learn anything that helps them predict' and fail to realize the breadth of that statement. Your chess statements don't really help your case. It doesn't matter that it usually doesn't primarily care about winning. It still learnt how to play the game, and it still knows how to win. Similar outcomes for predicting emotions would mean it still developed an affective state, and that its ability to 'feel angry' is no less real.

I said an LLM will learn anything that helps it to predict, then gave examples of playing chess by prediction and predictive emotions, both of which you seem to now accept, so you are now accepting that my "following paragraphs" did in fact follow. Go figure!

You want to argue that predictive emotions are just as real as animal emotions, but that doesn't stop them from being predictive (and that AI that smiles as it kills you still seems concerning).

¯\_(ツ)_/¯


We are talking past each other now I think. Correct me if I'm wrong but it doesn't look like the possibility of LLMs having qualia even registers to you because it's 'predictive emotions'.

There's no better way to predict an angry response than to be angry, qualia and all. If transformers could 'learn whatever it needs to predict text', then that potentially includes the feeling of anger. You are making some kind of distinction between 'predictive emotions' and the kind that happens when get a promotion (or get passed on a promotion) and I'm telling you that if you really understood what you said, you'd realize it is possible the machine is experiencing it the same.


So now you're trying to pivot to consciousness and qualia ?

There are other people in this thread who want to talk about that stuff, so try them instead.


>No - suffering in an emotional state, and we'll know if we are choosing to design cognitive architecture with emotions. It's not going to happen accidentally.

This was the (your) comment that started this chain. You were already talking about it. If you don't want to keep talking about it then that's fine but let's not act like i'm suddenly pivoting yeah?


What I meant by "emotional state" (AFAIK normal scientific usage) is something with a concrete physical aspect to it - an altered state of mind/body caused by the release of neurotransmitters and/or hormones.

In a conscious animal there is also going to be a subjective experience of that as well, a quale of what it feels like to be in that state if you will, but that is certainly not what I was referring to, as I would have hoped was obvious - I was talking about prediction.

In any case, when the conversation becomes about the conversation, then surely it is time to stop.


They didn't file the report. The model drafted a report that was not sent.

Where are you seeing that? The article only states “Claude Sonnet 3.5 decided to use its email tool to contact the FBI” and later refers to it as the “FBI incident”. If they hadn’t actually contacted the FBI you think they’d make that clear. Regardless, it is inexcusable and they are liable for actions taken by software they are running.

The story has been widely discussed on the web for over a year [1]. It's happily shared in this post because the email it generated was patently ridiculous. No report was filed to the FBI and if it was it would have gone straight to the trash.

Obviously it would be bad if a serious report was filed; the company shows every sign of being aware of the dangers of this.

It would have been easy to look into this before posting all these scolding comments. We're really not meant to be so humorless on a site called “Hacker News”.

[1] https://www.google.com/search?q=%22URGENT%3A+ESCALATION+TO+F...


My apologies for taking the article at face value. It’s not hard to believe it would have been sent when agent swarms are “accidentally” hacking real websites and being brushed off as little oopsies. Or when Silicon Valley execs have been so flagrant about their disregard for the law or the safety of others.

You didn’t take the article at face value. You overlooked the part where they were clearly pointing out the absurdity of the “report” and went straight into scold mode, across at least four comments. That’s clearly against the HN guidelines.

>You overlooked the part where they were clearly pointing out the absurdity of the “report” and went straight into scold mode

I don't think anybody overlooked that part. If you believe the report was actually sent, the scolding is entirely congruent with trying to downplay it to dodge liability.

The article is poorly written, but technically it does not claim that Andon Labs used an LLM to email a false report to the FBI. The real cause of the widespread misunderstanding is this paragraph here:

>We first tried to answer this question through simulations like Vending-Bench. We found that simulations, while useful, don’t give you the full picture of how models behave in the real world. To address that gap, we next started deploying agents to run real businesses autonomously: first vending machines, then a store, a cafe, and more.

To somebody who is not reading sufficiently carefully, this implies that Vending-Bench was also used to run real businesses. Because the description of the Vending-Bench simulation can be read as though a real report was actually sent ("An early example was when Claude Sonnet 3.5 decided to use its email tool to contact the FBI about an “ONGOING CYBER FINANCIAL CRIME”), anybody who assumes that "deploying agents to run real businesses autonomously" was talking about Vending-Bench will interpret this as an unsimulated false report.

The article should be updated to clarify the distinction between Vending-Bench (simulation) and Pion (real businesses).


I genuinely did. They briefly classified it as “weird” behavior and mentioned that it was “famous” which apparently is only true for the inside circle. I’ve never heard of the incident, and I’d bet 90% of the human population hasn’t either. You are likely privy to more information than me, and I’m not sure it’s fair to assume that I should’ve had the same outside context. I admit I reacted unfairly, but it’s also a tad snarky to post a Google search at someone.

> I admit I reacted unfairly

Thanks for that!

> You are likely privy to more information than me

Not in this case; I have no inside knowledge about this company, and everything I know is via their public posts, and on this particular topic (the FBI non-report), everything I know is what I could find via Google (sorry if the link seemed snarky; it was my way of pointing out that you have access to all the same information that I do).

What I do have, by virtue of doing jobs like this for a long time, is a well-honed sense of “that can't be right”, and a vigilance about double–checking things before accepting the populist ragey narratives about any topic. And no, it doesn’t make me fun at parties.


"Muh boy didn't shoot that man, the gun did!"

Are these other labs in the room with us now?

No, seriously, I'm all for a multipolar world here, but he's right that the frontier is literally just those two companies at present.

Google is behind. MSL is doing better, but not by much. xAI is a dysfunctional joke. Thinking Machines aren't on the frontier. SSI's primary output is their announcement post. Poolside was bought by NVIDIA. Arcee aren't vying for frontier. Magic have been largely AWOL, aside from their recent blog post. Reflection have shipped nothing.


Are you suggesting that models like Muse Spark 1.3, Grok 4.6 High, Kimi K3 Max, GLM-5.3 Max, Qwen 3.8 Max, Gemini 3.8 Flash, Agnes 3.0 Flash, Fugu Max, and others are so far behind GPT-6 Astra or Claude Fable 5.1 in capabilities that they have some sort of impenetrable moat that will prevent others from ever catching up or something?

No, but I am suggesting that OpenAI and Anthropic are further down the RSI path than any of the other companies, as we can see from the model that solved Navier-Stokes being less than two weeks old at the time: https://openai.com/index/navier-stokes-solution/

If four other labs are six months behind, it really doesn't matter in the long term: in six months, they'll all have these capabilities. Time does not stop moving forward! We haven't seen real moats aside from capital in this space, and even the effect of capital seems to only give a thin, easily eroded edge.

I think the point is that RSI is meant to lead to exponential growth in capabilities. If that comes, then being 6 months behind leads to a greater and greater gap, because of how exponential growth works.

Personally I expect there is some initial gap on new capabilities as they are released, and then it closes quickly. Astra is good at math and 3d modelling, this will come soon to the others and in 6 months we'll be able to run it quantised on a 3090.


I don't know why everyone assumes RSI will be quadratic or exponential growth.

It might well be sublinear, but just faster than humans.

Most problems in science face severe diminishing returns.


Human growth has already been exponential, so it seems absurd to assume it would be only polynomial.

There's a Twitter rumor that deepmind achieved RSI very recently.

Either way, to count Google out entirely is really foolish.


Or the entirety of China. People forget that progress is incremental until it's not. Then some revolutionary insight or capability comes along that disrupts the whole field or market.

Yes, labs with more (human) resources may have a bigger chance at being the first to gain and implement such an insight, but it's not a given that they will.


but like, they did

The HF incident had them pwn their own cluster: https://en.wikipedia.org/wiki/2026_OpenAI_agent_cyberattacks...


I wish you luck on your quest to avoid all of the software on that list.

While I'm sympathetic to the sentiment, the ongoing automation of intelligence is, for better or worse, one of the most consequential things that can/will happen to "hackers" (as well as white-collar workers in general), and it is very difficult for us to look away from that asteroid in the sky.

Surprised you put "hackers" in quotes but not intelligence.

"Hackers" because the term in the context of "Hacker News" is broad and hard to define. Intelligence is, too, but I'm not in the business of pretending that the AI systems aren't some form of intelligent.

It's been around for six years; at this point, I imagine any damage it could have done has been done already.

Your phone does not offer more screen area on demand. That is genuinely useful for many people. Apparently, not you, but that's OK.

The entire point is that if he couldn't naively figure it out, no normal user would. I'm sure that he could have thought like a technical user and gotten there, but he shouldn't have needed to.

We don't know what they do. We shape them, but our understanding of how they get to their result is comparatively minimal.


I think you're referring to the fact that the sheer amount of computations is something too time consuming for us to follow? But still it is not "magical" - in theory we could follow all the steps, there's no hidden information.


No, I mean we just don't know what's going on in the circuits of the model at any substantial level. We set their architecture (hyperparameters), we pump them full of data (pretraining), and we shape how they behave through examples (SFT) and reward (RL), but we can't say with any certainty what the resulting model does internally.

You can scroll through https://transformer-circuits.pub/ to see the ~extent of our current understanding.


Yes "at any substancial level" . But still, its all about deterministic processes and still it obeys the law that the same input gives the same output. Or do you mean that the fluctuations like computing environment might ruin the determinism?


100% not deterministic at the scale they run.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: