Hacker Newsnew | past | comments | ask | show | jobs | submit | danieltk76's commentslogin

Why was this not a thing BEFORE continuing to develop AI? Makes me think that if they actually believed in AI causing extinction, they would have already had a kill switch.

The article mentions this is about a legal requirement, not Anthropic considering adding one. They state they and many others already have one.

There’s even a benchmark for kill switch efficacy!

https://arxiv.org/abs/2511.13725


Found the omission of Claude odd, turns out Claude considers that approach prompt injection and ignores it.

Not if the extinction happens after they're dead. Then they wouldn't feel obligated to do so because it won't affect them. Instead, speaking hypothetically, if they truly believed that AI would cause extinction, then they would only implement the kill switch sufficiently many others believed it and they could claim plausible deniability for not truly understanding what AI would become.

** Note that I'm not claiming that AI will cause extinction, just continuing your hypothetical reasoning.


Related, on pacing:

> Slowing down in order to address their alignment risks felt like trying to study the psychology of humans by performing experiments on bacteria.

Note author’s small financial ties to the subject (Anthropic CEO) https://darioamodei.com/post/we-must-pace-the-frontier


I find it hard to believe that this wasn't a serious consideration until recently.

It was a serious consideration, and almost everyone around here laughed at it.

There valid reasons people laugh at it though. Kinda the same reason serious people laugh when you tell them the gun has digital failsafe.

It became a very serious consideration for Anthropic this year, with the advance of Chinese AI.

That’s a bit unfair, Dario Amodei has written in this topic a lot since around mid 2010s IIRC, Anthropic too published a good amount of stuff on similar topics. I don’t think the lack of consideration is really the issue here. It’s more a question of incentives

We're pretty crap in a capitalist society to think about those things ahead of time. Firstly, the idea that AI could "runaway" was simply a concept or a thought it wasn't baked into a real product that could do that. We're now getting close or perhaps we are at that point of where you can't race at speed for investors without now considering a real kill switch.

You could say this in hindsight for many times in which disasters or engineering issues have occurred.


I was working on some soc2 audit materials for my startup, and claude refused to edit a document because it would "tamper with the record". I suppose claude noticed I started the doc a few months previously and was worried I was tampering with evidence. No matter what I did it wouldn't help me. I suspect there will be more situations like this in the future.

We tested whether prompt canaries and honeypots could detect AI attackers.

Prompt canaries were surprisingly unreliable. I suspect the defenses model providers have added against prompt injection also make models less susceptible to prompt canaries.

Honeypots worked much better, having roastable accounts that are otherwise useless seems to be a good catcher.


I could see the value of using burp as a way to have a GUI look at what your agent is doing, but I know at Vulnetic and other ai security vendors we use our own proxies or custom scripting rather than burp.

I've heard of Vulnetic, and it looks quite cool, though I feel it's a bit expensive for individual users (probably fine for companies)

reachout at danielk@vulnetic.ai and Ill give you some free credits to try us out!

yea but part of this is the consolidation of funding too. Standards to raise seed capital are soooo lofty now compared to 3 years ago. If you are in your in, if not good luck.


great, but nobody can use it for another 100 days right?


"GPT‐6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business"


Then make a pretty blog post "over the coming days" instead of now and send an internal memo to that "limited set of organizations"?

hahaha i still have access


Daybreak blue is definitely a good model (I think a further post trained GPT 5.6 sol). Alot of the capabilities they talk about Astra having though have been available with good harness engineering for a year now.


Where would you recommend to look into regarding Harness Engineering for Cyber-security as well as for other use-cases.


This is moreso about the (human-intended) tools, data, and environments you have available to you. Wanna do defense? Get more telemetry. Wanna do red? Get solid test-bed environments. Mature infosec programs are benefiting the most, good-guy-side wise, at the moment; because they've got these things in order already.

As far as harness engineering goes, it boils down to your ability to clearly define goals or success criteria, and safely facilitate the necessary access via the harness. There is no easy single piece of advice here, sadly. Though it would be helpful if you said what 'for Cyber-security ... other user-cases' means in your case.


DARPA’s AI Cyber Challenge (AIxCC): Competition Design, Architectures, and Lessons Learned: https://arxiv.org/html/2602.07666v2


this might be dated. the tech moves very fast in this space and architecture from 6 months ago is dated.


We have posted some articles on it (blog.vulnetic.ai), Im also happy to chat!


The guardrails are horrendous for cybersecurity. you will get booted quickly down to Opus 4.8


is there just a GC where people go sign these things?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: