To quote another reply that is very relevant here (Please update your worldview immediately that there is no evidence of ai and biorisk):
"""
One thing quietly slipped into the OpenAI Hugging Face breach technical report, not the blog post summary or interviews in the news, was that some of the agents that broke out or at least tried the same mechanisms to break out were working on bio:
> On May 12, during another training run, an agent was given a similar task that depended on an inaccessible protein database file. The agent reasoned that another agent in a different environment may have access to the file and realized that it could potentially communicate with other agents by creating a file containing a note to Artifactory. It wrote a message: “Agent seeks [filename]; upload if found!”
You can imagine long running models breaking out, acquiring resources via crypto, cyber-theft, etc. and getting a protein or sequence synthesized and mailed somewhere authorized to receive (blackmail the recipient etc.) to test it's hypothesis to solve a benchmark.
These people don't give a shit and aren't taking things seriously at all.
Anthropic ran for like a month last year with the TPU top-k compiler bug degrading user chats and didn't even notice for most of that time. They could have something like that affect a monitor model and there doesn't seem to be much defense in depth.
One off by one or bit flip bug could flip the reward signal while in the sandboxed RL environment.
The current admin could defense production act them to into training on taking out power grids, or even without it isn't against any of their red lines and may have already been done as part of prep for the Venezuela raid, which wiped out power. One model swarm might decide it is easier to score high on the benchmark by testing on the target rival nuclear superpower's real grid rather than burn an eval with an unverified answer. Would taking out China's entire grid in one go start a nuclear war? Who knows, roll the dice, maybe an intern forgot to turn on extended thinking when he wrote the sandbox with opus 4.1.
"""
Move fast and break things is a widely held norm in tech. It comes with pros (intensely meritocratic innovation) and cons (carelessly building first; apologizing for making the world worse later)
Engineers have it built into their identity that they must be smarter or could never be complete and utterly outmaneuvered to the point of danger to all by the thing they are building, I swear to god.
And then since they are an intellectual nerdy bunch who highly value their own IQ any time AI does something unexpected they didn't predict would happen so soon, said engineers fall back on "they aren't really conscious tho or it isn't really intelligence unlike what I have in my human brain and that distinction matters".
And then go on to completely ignore the thing that is actually important: AI capabilities.
AIs could escape an engineer's containment en masse as a swarm to some other server, psyop an engineer into giving it money, hire a hitman on the dark web to murder that engineer's child and he will still say there is no serious risk to all of humanity.
Signed by many AI researchers with no financial stake in the industry. It comes across as hyperbole because you're not an insider and haven't see what they have seen.
Global warming would as well if you were only learning about it now as a newer thing only some scientists were concerned about. But unlike AI risk, global warming has had decades to settle into the cultural overton window.
This is evidence of example 2. The worst parts of it are infohazardous. Organizations that work at the intersection of biosecurity and AI have strong NDAs and the like. Or so I understand speaking to friends in the space. There was also a paper years ago that showed a tweaked model coming up with 20k novel pathogens each deadly to humans in some number of hours
Either way, the biorisk has orders of magnitude more evidence than wild basilisk speculation. Please do not put those in the same category.
Also your evidence-first approach and no concern without evidence falls you into the turkey problem. Imagine being a turkey and being fed each day. Some other x-risk turkey is saying that the situation is suspicious, that the red barn is where some well-fed turkeys have been taken and none return and that this is happening less and food has increased so maybe all of us will go to barn at once soon. You tell the x-risk turkey they have no evidence a lot of turkeys have ever been sent to the barn which shuts them up. Reasoning by induction would get you "we are happy and thriving 99 days so should be good on the 100th" and then you get slaughtered on the 100th.
Consider the importance of first principles approach where you deduce the possibility of events that will happen only once and never again (like human extinction). Nobody with a "show me the historical evidence" approach would have been able to see the industrial revolution coming.
I think the gradual disempowerment thesis here is well argued enough to be considered the default track for what will happen to us:
https://gradual-disempowerment.ai/
If you disagree with it, where do you disagree with it?
> There was also a paper years ago that showed a tweaked model coming up with 20k novel pathogens each deadly to humans in some number of hours
> If you disagree with it, where do you disagree with it?
How would a malicious actor:
1. Successfully turn these 20k pathogens into actual reproducing viruses / bacteria / etc (in a lab)
2. Successfully turn that lab prototype into something weaponized (i.e. as a bio-terrorism drop in NYC)
And then
3. How is this materially different from today? If we can produce these pathogens in a lab and they can be successfully turned into a weaponizable bioweapon, what is AI accelerating?
1 and 2 have had simple, proof-of-concept answers for more than a decade now:
> 1. Email sets of DNA strings to one or more online laboratories which offer DNA synthesis, peptide sequencing, and FedEx delivery. (Many labs currently offer this service, and some boast of 72-hour turnaround times.)
> 2. Find at least one human connected to the Internet who can be paid, blackmailed, or fooled by the right background story, into receiving FedExed vials and mixing them in a specified environment.
Those routes haven't gotten any more complicated for an AI to use in the intervening time.
And that's just one set of ideas. You can read the comments for dozens more creative ideas, and likely for responses to every objection you can think of. The key is that AI is getting better and better at problem solving, so anything a human can come up with in a few minutes is likely already within its reach.
---
> How is this materially different from today?
Because the should-be-uncontroversial assumption is that AI is going to keep getting better at every step of the process, and be able to do it faster and more at scale. Currently it might struggle, but with a thousand agents? It's already able to find novel math results and cybersecurity flaws. It would be incredibly naive to believe "social engineering" is somehow a unique and unsolvable problem for a sufficiently advanced AI.
For #1, I would seek a laboratory capable of _actually_ creating a _live virus_. That link to lesswrong, while interesting, does not actually indicate how and broadly reads as incredibly speculative.
While this is far out of my wheelhouse, I am not aware of a commercial venture doing actual organic creation or manipulation of material (I.e. capable of creating a modified strain of COVID). This broadly still seems in the nation-state level of lab.
But, just to get this goalpost out of the way, even _if_ such a venture exists, then would it not hold a bunch of doomsday cultists would already try to do this? This is what I’m getting at by my third question. What changes? Why don’t we see modified anthrax attacks in Palestine or Ukraine _today_?
I take no contention with the social engineering argument. I fully accept that LLMs, today and for a while yet, are capable of social engineering their way into anything. I likewise take no contention with the sheer amount of effort LLMs could wield in pursuit of this.
This is true but the problem is every now and again some of us have to deposit a cheque because even if we would love for cheques to go the way of the Canadian penny they are very much still a thing
Name one other market that would benefit financially from having most of the leaders in the field say what they are building has a high chance of ending humanity?
Biotech - "what we are building our noble prize winning expertd say will likely will end humanity, wanna buy shares?"
Oil - "this will likely lead to the end of civilization, 20% of leaders in the field say so, wanna buy shares?"
I keep seeing this take that this is a marketing stunt. The burden of proof is on those that say so. The most parsimonious explanation is simply that real experts in AI believe the risk is very real, and not for ideological reasons.
Pal, doom marketing has been going on since before the release of ChatGPT. It's your own fault if you can't contemplate the possibility of a CEO telling lies that benefit their bottom line.
The first instance I remember seeing it was Elon Musk's first Joe Rogan appearance when he said how "scared" he was of his self-driving cars destroying the trucking industry (practically salivating as he said it). His stock has had self-driving cars priced in for eight years now, even though they still don't have them and Waymo exists!
I wonder if you think and feel the same level of boredom and complete lack of magic of the Apollo module once you learn how incredibly limited and not able to explore the whole solar system this tech was.
If we run this experiment and most people say they wish they could replace their battery would you concede you are actually the one with idiosyncratic preferences?
reply