Why was this not a thing BEFORE continuing to develop AI? Makes me think that if they actually believed in AI causing extinction, they would have already had a kill switch.
Not if the extinction happens after they're dead. Then they wouldn't feel obligated to do so because it won't affect them. Instead, speaking hypothetically, if they truly believed that AI would cause extinction, then they would only implement the kill switch sufficiently many others believed it and they could claim plausible deniability for not truly understanding what AI would become.
** Note that I'm not claiming that AI will cause extinction, just continuing your hypothetical reasoning.
That’s a bit unfair, Dario Amodei has written in this topic a lot since around mid 2010s IIRC, Anthropic too published a good amount of stuff on similar topics. I don’t think the lack of consideration is really the issue here. It’s more a question of incentives
We're pretty crap in a capitalist society to think about those things ahead of time. Firstly, the idea that AI could "runaway" was simply a concept or a thought it wasn't baked into a real product that could do that. We're now getting close or perhaps we are at that point of where you can't race at speed for investors without now considering a real kill switch.
You could say this in hindsight for many times in which disasters or engineering issues have occurred.
I was working on some soc2 audit materials for my startup, and claude refused to edit a document because it would "tamper with the record". I suppose claude noticed I started the doc a few months previously and was worried I was tampering with evidence. No matter what I did it wouldn't help me. I suspect there will be more situations like this in the future.
We tested whether prompt canaries and honeypots could detect AI attackers.
Prompt canaries were surprisingly unreliable. I suspect the defenses model providers have added against prompt injection also make models less susceptible to prompt canaries.
Honeypots worked much better, having roastable accounts that are otherwise useless seems to be a good catcher.
I could see the value of using burp as a way to have a GUI look at what your agent is doing, but I know at Vulnetic and other ai security vendors we use our own proxies or custom scripting rather than burp.
yea but part of this is the consolidation of funding too. Standards to raise seed capital are soooo lofty now compared to 3 years ago. If you are in your in, if not good luck.
Daybreak blue is definitely a good model (I think a further post trained GPT 5.6 sol). Alot of the capabilities they talk about Astra having though have been available with good harness engineering for a year now.
This is moreso about the (human-intended) tools, data, and environments you have available to you. Wanna do defense? Get more telemetry. Wanna do red? Get solid test-bed environments. Mature infosec programs are benefiting the most, good-guy-side wise, at the moment; because they've got these things in order already.
As far as harness engineering goes, it boils down to your ability to clearly define goals or success criteria, and safely facilitate the necessary access via the harness. There is no easy single piece of advice here, sadly. Though it would be helpful if you said what 'for Cyber-security ... other user-cases' means in your case.
reply