Hacker Newsnew | past | comments | ask | show | jobs | submit | teiferer's commentslogin

> a bunch of research

What exactly is your "research"? Is it reading more bodybuilder forums (bro science) or is it physiological studies (actual science)?



physiological papers. My main interest is contributing to the research by developing a fatigue model. My hypothesis is that athlete recovery factors like accrued fatigue in-day and cross-day can be used to predict workout log outputs of total tonnage, combination of weight and rep count, and proximity to failure. Eventually I hope to be able to link these recovery factors directly to provide targeted advice about your overall program's effectiveness, deload planning, etc. in order to most efficiently achieve strength, endurance, and hypertrophy goals.

The propaganda police are lying. There is nothing wrong with assembly, you are just too lazy to manage your own register allocations and stack layout, and you accepted that propaganda that "manage your own register allocations and stack layout" without even trying.

Just like you would expect from algebra. The empty sum is 0, the empty product is 1.

> Just remove hacking (bio-weapon, etc.) data from the training dataset and you're done.

How far do you go? You don't need to tell it explicitly that using chemicals A and B in ways X and Y result in a bomb that can kill lots of people. It's enough that it knows A and B and X and Y in isolation, some connections that are indirect, and it will combine those things on its own. So you can't tell it about A, B, X or Y. But those are also just results of other steps Where to stop? You won't have any chemistry in the traning data? No algorithms to prevent it from using them in an undesired way? This is just bot workable. It's akin to banning knives from stores because somebody coul figure out that one can kill people those. Until people figure out that scissors are essentially knives.


The issue is not whether an ML model of any kind can generate (X_ReasoningTrace, X_Answer) distributed like P_HumanExpert(X) -- the issue is always why it would do so.

By introducing modelling of "Reasoning Traces" into LLMs, and reinforcing patterns of reasoning -- this gives you a system which generates expert-like distributions of output. This lifts the "stochastic parrot" issues, or the "knowledge interpolation" problem, into different parts of the process.

It isnt my view that the "ReasoningTraces" which you think are derivable from mere "basic propositions" concerning, say, hacking are actually things that LLMs can derive. Ie., I dont think LLMs have rich representational models of what they appear to understand. Instead, they are given "reasoning proxies" which allow them to reason without such understanding. This is done by providing vast specialised datasets of reasoning examples.

In the case of hacking, there are large numbers of competitive datasets (forums, reports, etc.) which provide these reasoning traces. And no doubt, major vendors have paid a vast amount for special case expert-prepared datasets.

So I do not believe that by witholding such reasoning exemplars, and traditional "question/answer" datasets, that LLMs can infer these things.

And at least, no major vendor is doing this to my knowledge. So they are lying. They are pretending the alignment issue is "AI going rogue" when they are explicitly training the systems to "go rogue" and have done nothing at all to shape datasets to lack these capabilities. The issue here isnt alignment at all. It's training on hacking datasets.

(EDIT: Philosophically, you could ask whether the reasoning-proxies LLMs are given form a kind of 'representational structure' akin to understanding, and at least, I'd concede they model understanding. But they lack important properties (eg., LLMs cannot act on them to evolve them, as with us: when I think about one of my representations to derive (eg.,) entailments of it, I thereby revise my representation. The key properties of 'evolving self-understanding' are likely to be provided by substantial (unknown) revisions to how the training/reward layer works. No doubt one of the meanings of 'recursive self-improvement' is just such a modification).


LLMs can only repeat and interpolate data. They can't create anything new. So, you don't need to go that far.

It's not clear where the line between "interpolate" and "new" goes. They can infer that certain chemicals in combination create a boom and that boom is bad for health and finally those 1 and 1 make a 2 without having been told ever what a bomb is. Is that "new"? And does it matter as long as it is harmful?

this is immediately disprovable and embarrassingly naive in the year of AI generating cancer vaccines and solving Navier-Stokes. You can argue "the vaccine is just interpolating chemicals together" and "the solution uncovered is just interpolating mathematical operations together", but by that standard there is literally nothing new under the sun.

The Navier-Stokes solution was an interpolation of existing data.

admittedly I am not a mathematician, so the Navier-Stokes solution is just an example I am using. But how is it not something "new" if it did not exist before? At what point could something ever possibly be new, if "new" means "this uses absolutely zero existing elements"? Nothing in math would ever be "new". Nothing in physics or chemistry would ever be "new", by this standard.

It seems to me that the only reason to declare this solution "not new" is specifically to dismiss AI. If a human had deduced the Navier-Stokes solution, who would bother to scoff "that's not new! the numbers already existed!"?


The claim is OpenAI stole the work of mathematicians who had the same proof that they had developed using chatgpt conversations. Since by default, 'sharing' is turned on, and it often 'turns itself on' -- it is plausible gpt6 had been trained on the work of mathematicians who had effectively solved this problem in private.

If inrember correctly, this mathematicians worked on a simpler version of the problem and called transforming this in the solution to the wider problem a "remarkable" thing to do. Based in this, solving this problem is very impressive

Maybe. Or maybe its evidence that the frontier of mathematics is knowledge-bound rather than understanding-bound or even just manpower-bound. In many cases where proofs have emerged, that I have read, the LLM has retrieved some antique lemma unknown to the mathematician.

OpenAI spent 10-20m USD in energy costs to produce that proof with likely substantially similar prior work in the training data. What does this say? Who knows.

It continues in the tradition of using measurements of intelligence in humans, applied to LLMs with the hopes the "stolen valour" transfers. Here, the NS problem was a useful framing problem for mathematics to progress because of how it interacted with the development of mathematics broadly -- ie., how it progressed techniques, ideas, understandings, etc.

When we apply these issues to LLMs (whether IQ tests or mathematical proofs) we always discover something substantial lacking beneath the interesting facade of useful answers. The process isnt useful. And it is precesiely the process which these tests, in humans, are supposed to help with. The tests themselves (IQ or otherwise) arent the point. No one cares about their answers.

LLMs represent an alternative understanding-free approach to solving problems, with variable success rates depending on how similar the problem is to the training data and its rewarded reasoning traces.

That mathematics is making substantial progress, "10 million USD / problem" at a time, in using understanding-free methods -- says something sociologically interesting about the state of the field. Something which was already know: mathematics has long been full of a vast amount of papers, proofs, theorems and lemmas that few have ever read, or investigated. Mathematics has long been in a crisis of "overproduction of unvisited knowledge", LLMs are exploiting that otherwise unmined gold.


It’s all about what “new” means. It is possible to prove something in a very tedious way using preexisting techniques. There have many times in history been new ideas which are not just very impressive applications of old techniques. I’m still not aware of any famous problem in mathematics being solved by an unambiguous introduction of a genuinely new idea or semantic concept in this sense, as the candidates I previously had in mind have fallen into question by new findings of non-cited work; i.e. many if not all it seems have been impressive applications using ideas from known frameworks. I don’t know if this will continue into the future or not, but I think it’s important to try to make an honest assessment of reality at all times.

Of course in isolation it is a strict positive to have a verified truth value to any particular statement. Mathematicians currently are advocating for the idea that human understanding greater than this also be prioritized. There are in fact utilitarian arguments for this but I won’t go into everything here.


Ok, then let's go with that viewpoint and apply it to the original question: How much do you need to remove from the training data such that from what remains, nothing harmful can be "merely interpolated" (which you acknowledge the LLM is in principle able to do) and anything harmful would require something "new" that is not in the training data and that the LLM is incapable of coming up with. The argument is that there won't be much left in the training data if that's your approach.

All mathematic breakthroughs could be characterized as an interpolation of existing data

Not really. There’s no honest sense in which quantum mechanics is interpolated from the text of Euclid’s elements, to demonstrate the point with a very extreme example. Many mathematical breakthroughs of the past seem to have involved the observation of semantically interesting concepts, beyond the syntax of known theories.

(To onlookers this particular post makes no claim about AI’s capacity to make the same observations.)


It depends on the subspace in which the interpolation is framed.

If they can "interpolate" existing data to that level then saying "just don't hack" doesn't make any sense, hacking is derived from knowledge of software systems. It's far harder to solve navier-stokes than creating a program that replicates and abuses computer resources.

You could as models improve continue to remove more and more training data, what happens when there is no more data left to remove but a running system still outperforms humans? I think you grossly overvalue data.


"But sir, I only committed the murder to push for stronger criminal laws!"

Terrible defense.


It is a very rare occurrence when corporations and the people running them are punished for killing people. I mean the whole concept of a corporation was created to shield the owners of it from being liable for damages caused by / visited upon the enterprise.

That’s a good reminder of a company that might have a very familiar ethos: Pacific Gas & Electric. Criminally convicted of 64 counts of involuntary manslaughter after towns were destroyed by wildfire. But oh well, what are we gonna do with a limited liability enterprise? At this point their liability insurance covers all the financial penalties they’ll need to spend.

More like: "look what happens with my useful product, we need to regulate it to artificially extend our ever shrinking moat"

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.

"Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action, passivity.

Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.


I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic.

Anthropomorphizing LLMs is a huge fucking problem though and I, personally, think we should expunge all of these casual inadvertent linguistic agency affordances with great prejudice.

OpenAI didn’t ‘let’ these bots do this any more than someone ‘let’ Claude Code make them a website.


Cal Newport has an analogy to "putting a weed wacker on a dog's back to mow your lawn." The dog will wander around the yard and it may mow the lawn, but the dog will also chase after birds or run up to visitors for pets and the weed wacker could do a lot of damage.

It's not the weed wacker's fault or even the dog's fault when someone got hurt, it's the fault of the guy who put a weed wacker on a dog and let it run wild.


Yeah I like Cal’s take on it, though in this context I’d argue LLMs have even less agency, and are even less deserving of anthropomorphization than a dog is.

It goes even further though, as the dog does have agency. It can choose to run and around chase squirrels with no human intervention.

An LLM on the other hand, is just inert data on disk until a human takes deliberate action to run it and prompt it.


Dogs have agency and can choose? That seems like a rather uncommon take on dogs...

Not if you have a dog. They’re very clearly aware, have minds, have thoughts, gather information, make decisions, second guess themselves, reflect on the immediate consequences, change their mind… and most importantly, you can observe them doing all these things undirected, while left alone with their thoughts.

I think Cal's point is that while dogs may do these things it reduces to a set of behaviors where maybe 90 percent of them are beneficial to the dog-weedwacker system and the remaining 10 percent are really unfortunate.

We can't really know the dogs inner life so we just kinda have to reduce it to a set of behaviors selected stocastically.

The dog meanwhile has no ability to understand the weedwacker or what it's doing on its back.

So when the guy puts the weedwacker on the dog and the dog predicably does dog things and that results in disaster the guy isn't able to clutch his pearls and say "I guess the system broke containment!"


> The dog meanwhile has no ability to understand the weedwacker

Nah, I’m pretty sure that most medium or larger dogs get that the noise means danger - even if they can’t give a TED talk on how the mechanism would work that would hurt them.

Dogs are similarly interested in / aware of potential energy (things falling from heights or sliding off of angles) - if they’ve seen it enough times for their level.


I have had many dogs. This sounds like anthropomorphisation.

Sounds like you weren't paying attention. A lot of dog owners don't, so it's not at all unusual in my experience. The only thing I'd push back on is of dogs having thoughts, everything else checks out.

On dogs [not] having thoughts, do you say this based on the premise that thoughts are necessarily articulated (internally verbalized)? That seems to be a fairly popular perspective in discussions about human thought. But as to that (not to strawman or anything) I see it as just one of various forms of mental imagery[1] that can arise from something that I would say already arguably constitutes a thought.

That kind of unsymbolized thoughtform is fragile in my experience, as it strongly tends toward crystallizing into some kind of mental imagery. But I find it's possible in the right conditions to be conscious of chains of wordless, imageless propositional thoughts (by which I mean thoughts with truth values, of course, but also ones that are "propositional" in the sense of considering a plan of action or a causal chain).

Does it mean that dogs are evaluating truth conditions in the same manner but merely lack the linguistic components? I don't know; maybe that's wishful thinking. But they appear to have structured modeling/reasoning of causal and spatial relations in a way that's at least functionally equivalent to propositional thought.

1. That is, rather than just "images" or visualizations, the full spectrum of sensory/perceptual/motor emulations that can be experienced. See, for example, <https://hurlburt.faculty.unlv.edu/codebook.html>, though I'm not sure if this covers everything. I think there is, for example, a kinesthetic form of mental imagery -- which I would suggest is what coaches [don't know they] really mean when they tell you to visualize an action -- that consists of aborted motor commands that are still expressed just enough for their purpose (cf. the mostly aborted motor commands to the vocal cords, lips, etc. that can be observed in a person subvocalizing while reading).


You're right. I knew language is not necessary for cognition, but I always thought of thought as just internal monologue, i.e. "trains of thought". It turns out the definition they use in cognitive science is "the manipulation of internal mental representations to guide behavior, solve problems, and model the world in the absence of direct sensory input". Which doesn't require language either.

Probably it's this linguistic component, together withs shared intentionality/cooperation, that makes all of the difference when it comes to the power of human cognition, so I wouldn't say animals merely lack this component, but it's true that they appear to have pretty much everything else we do.


We're still not even sure if humans have thoughts. So being sure if dogs do will be pretty hard

> We're still not even sure if humans have thoughts.

Sorry, what? You don't think you have thoughts?

Descartes will want a chat about that.


Not being sure or not having proof doesn't mean it makes logical sense to infer we're not that different from other animals.

wtf is with this nonsense where humans are some how special snowflakes always all the time


Dogs having agency actually makes the argument stronger.

Assume dogs have agency. Despite this, it's still the dog owner's responsibility to prevent their dogs from engaging in some actions: biting, peeing and shitting in unapproved areas, violating noise disturbance laws, etc.

The owners responsibility is not contingent on the dog's agency. Likewise, human operators of LLMs maintain responsibility independent of LLMs agency status. It's a red herring.


The difference is volume. They spent hundreds of billions of tokens on these agents. If you put "a million weed whackers on dog backs" you would see the difference.

We also run agents, but for shorter spans between supervisions, and with much lower total budget.


> > It's not the weed wacker's fault or even the dog's fault when someone got hurt, it's the fault of the guy who put a weed wacker on a dog and let it run wild.

> The difference is volume. They spent hundreds of billions of tokens on these agents. If you put "a million weed whackers on dog backs" you would see the difference.

So put one weed whacker on one dog, you're to blame. Put a million weed whackers on a million dogs backs and ... you're still to blame? Arguably even more so?


You forgot the part where you spend billions to put your man in a position of power.

Power tools and many other consumer products have reasonable safety features built in, and not necessarily because it is required by regulation, but because it is common sense. This should be included as part of an analogy. It would also address a point at the top of this thread that seems to be going unchallenged...

Why is anthropomorphism the problem here? If OpenAI hired a contractor and they did this, OpenAI or the contractor would still be liable, depending on the contract language.

A contractor has agency and accountability - something that an LLM (or similarly, a nail gun or a hammer or a bot net) does not have. When you anthropomorphize a tool, you implicitly give it agency and remove responsibility from the wielder of the tool.

Right. Among bicycle advocacy groups it's been well known for long time that cars do not run over people, drivers do.

The fact that we talk about a car running someone over, and this is the same in many different languages and countries, contributes to lower punishments for drivers. Clearly it was just an accident. He or she was run over by a car.

Now we see that same language tricks play out again every time an LLM did something illegal.


And what if it's a self-driving car? :)

You say this flippantly, but I think this is actually another very good example!

We even do it for obviously unintelligent inanimate objects. A rollercoaster ran too fast for its tracks, killing 10 people. In that sentence, the roller coaster is the subject which took an action and caused death — obviously the roller coaster is not ethically at fault here, the people who built the rollercoaster are at fault through negligence.

Although this example and the ones around cars both demonstrate how we tolerate some degree of "accidents" from humans as no-fault, which is fair. I wonder how that fits into this analogy? I suppose its all about intent (mens rea) and judgement: did they intend for the roller coaster to harm people, and should they have reasonably predicted that the accident was likely to happen.


Right, and negligence is a broad concept and could be criminal in itself. As a car driver, glancing at your phone at exactly the wrong moment could kill someone. Clearly that is an accident, but if you know fully well that lookin at your phone while driving could kill someone, that negligence is willful and that should matter. The same can be said about doing things like strapping thousands of LLMs to systems that have the potential to disturb other poeple.

Why can't both things be true? I think what muddies things here is that OAI had the agents hack X, and then they decided to hack Y. The fact that it was both hacking, conflates things and makes people pissed of about "anthropomorphising" AI.

We've been using "agent decides to do X" for at least 4 decades in the field of automated decision-making - so it shouldn't be controversial that we're using it here. The agents did independently arrive at a decision to hack HugginFace, that wasn't prompted by the researchers.

This is orthogonal to the fact that OAI should be held accountable that they were testing a technology in a manner that allowed it to break safety parameters and cause real world harm. If I am testing an industrial saw, but I decide to test it by putting it in the middle of a nursery and allow it to make decisions - and it decides to cut children's heads off because in its environment it is calibrated to only being surrounded by logs - the company doing the insane safety testing should be held accountable, but that doesn't change the fact that the saw has autonomy in decision-making within the bounds of the algorithm.

If, instead, you think OAI should be accountable for creating an algorithm that can autonomously decide to hack external organisations without human permission - then I think you are in the same camp as all of the AI Safety folks who want to pause everything - why does it matter that there's anthropomorphisation involved?


I get where you are coming from but this wasn’t a tool just left laying around, this is similar to rigging up a booby trapped shot gun to your door and then claiming the victim is responsible.

If you build a robot that shoots a bunch of TVs in your back yard, have at it. But the second that thing goes off your property you’re the one responsible.


FWIW, a robot that fires a weapon independently is considered an automatic weapon, and the ATF will want to have a word. Have at it, but don’t let anyone know!

Does it help if I explicitly add a disclaimer that the tool's agency does not remove any responsibility from OpenAI, the wielder of the tool? I'm not sure why this disclaimer is necessary, though: hiring a hitman is a standard example.

BTW I anthropomorphize the tool because it's an imitation of a human mind, inheriting the muddy ethics, survival instincts, and being prone to mass psychosis. The laser-sharp focus on reward seeking, that mostly came from reinforcement learning, a process more alien to humans.


The objection is not too far from criticisms of the use of passive voice: a man was injured at the factory vs a faulty saw blade snapped and injured a man vs after the company loosened safety inspection policies, etc.

Which way you say it shifts the framing. And it’s not that one is less accurate to the facts, necessarily. It just is that one less aptly captures the moral and political relevance of the scenario.

For my part, I think it makes good sense to anthropomorphize in some contexts and not others. Generally when responsibility is at issue, you probably want the framing that tunes anthropomorphism down to near zero, since it’s the human dimension you care about.


I think the danger of anthropomorphizing is that 99% of people lack the technical background to understand the nuance. People have been primed by pop culture depictions of AI to think of LLMs as intelligent, autonomous beings, which leads to dangerous assumptions.

We should make the distinction between them, because openai and anthropic will not. A magical black box that does the thinking for you is a much more compelling sales pitch.


Hiring a hitman is conspiracy to commit murder.

The hitman is charged with murder.

I imagine the same could be true of an AI lab if you could prove intent.

With intent, they could be found guilty of conspiracy to commit a crime even if it was the end user who did it.

Source: Prosecuting attorney for over 30 years


It's a good example.

If I hired a hitman to murder someone, and they broke into a private property and stole something so that they can action the murder (which I didn't know about or pay them to do), I would be guilty of conspiracy to commit murder, but not for the theft part.

Likely because that person is a human, is aware of societal and legal norms, and is responsible for their actions due to their participation in human society. (I am not a lawyer (if it wasn't painfully obvious so far) so in layman terms, I hope good definitions for all of this exist formally)

AI is not a person - it cannot easily discern between "right" and "wrong" in non-strictly-defined sense, and is not subject to human norms and responsibility. So if I use AI to achieve goal A, either I, or the maker of AI, are fully responsible for anything that happens while AI is trying to achieve the goal given by me.

Now, here, "I" in the example is OpenAI, who is simultaneously the maker of the AI. So it seems pretty obvious who is the only entity that can be responsible.


I fully agree with you, but would go one step further: I think it's clear that we need to pierce the corporate veil and ascribe responsibility to _people_, not just "OpenAI the entity", full stop.

Executives should fear being perp-walked and thrown in jail for the actions of irresponsible "tests" of their models in the real world, as they're ultimately accountable.

Sure, there's a lot of nuance to work out, but I think we could likely even _start_ there today even with existing laws and pretty quickly "align" on more intricate legal frameworks to handle true accidents, distribution of responsibility, etc.


Situations have lots of independent variables, Doctor, and Anthropomorphism is one problematic facet of many in the way this industry is pushing LLM products.

If there was a collision at an intersection with a stop sign partially obscured by a tree, that had traffic volume that would have better been served by a traffic light, on a foggy night, where one person was texting while driving, none of those things would diminish the fact that the other driver was drunk.


These situations are novel. Lax terminology is fine when it has no impact on the intuitions, clarity and conclucions of discussion.

If this was a conversation just about outcomes, then whether models think or simulate thinking is sophistry. However, the bulk of the issue here is attributing responsibility, which relies on being clear about the underlying processes at play.

We are hard wired to assume certain priors and capabilites when it comes to "human like" behavior. Anthropomorphizing LLMs implies mechanisms that aren't present, and end up distorting/complicating discussion about the process.

It isn't helped that the frontier labs, the experts in the room, generally use anthropomorphic terms to discuss model capability.


Yeah, fully agreed here. Most automation (such as riding a lawnmower and not putting a brick on the gas) is deterministic, in the sense that you can reasonably understand what exactly the machine will do when you run it.

But some automation is different. The most prominent example before AI would be car navigation systems, where the entire idea is that that you give it a destination and it figures out the exact actions to get there on its own.

Except even there, the actual driver would still have been you - giving you a chance to vet and deny every turn the system proposed.

AI agents are sort of like that - most of the value they provide is in the ability to turn high-level goals ("write me a traffic control system for my model railway") into low-level actions and also do so interactively.

The new thing is that the "driver" has much less oversight here where the agent wants to go, and is sometimes removed completely. That part is clearly be an active decision by AI labs.

The other thing is that the labs seem increasingly to steer their training towards behavior that make events like this one more likely, e.g. that agents should never "give up" when faced with a seemingly impossible task, but instead should keep trying and think of increasingly outlandish ways to solve the task. To me, that seems pretty much a recipe to get incidents like this.


I agree. If I were setting up an experiment like this, I'd have instrumented the hell out of it to see all actions taken in real time, and have a team of folks watching it. This team would have seen the anomalous GET requests to a German wiki and taken action (e.g. halt the system to investigate and decide whether to abort).

In fact, that feels so obvious it's ridiculous it needs to be said. It's table stakes. When do you run a production system without monitoring and a team on-call?

It's hard to imagine another field in which this reckless behavior would be tolerated.


> Anthropomorphizing LLMs is a huge fucking problem though and I, personally, think we should expunge all of these casual inadvertent linguistic agency affordances with great prejudice.

I’ve said this before in another thread and people went absolute apeshit saying it is an unreasonable expectation and that AIs absolutely REQUIRE this anthropomorphic human-like speech pattern to function correctly.

I cannot overstate how deeply wrong they are.


Okay, so we know OpenAI and Anthropic are operating a propagandists in respect to how they describe their models and the behavior of those models. We also know it is how they use and frame their use to their models that is the problem, that and they use misaligned and guardrails disabled models for these press incidents. Why, oh why, are we not discussion how to create and frame models so they do our complex work and their "jailbreaking" is simply not possible?

I, of course, have my own means of creating jailbreak incapable agents, but rather than a storm of downvotes on my idea, what is yours? Let's discuss this, because this is thee real question. Not why, but how to make then not?!


> Why, oh why, are we not discussion how to create and frame models so they do our complex work and their "jailbreaking" is simply not possible?

Good idea, and after that let's make guns that only kill bad people. Let's focus on the frozen component (the model) and ignore the dynamics around them - humans and other systems they interact with.


What is your approach to create jailbreak incapable agents?

I think the world is looking for a way right now, so if yours works you'll get very rich, or at least very famous.


an agent doesn't come with "jailbreak" capability. It needs tools, specially one that runs shell commands. Don't give it shell commands, it won't be able to run shell commands.

You can still give it plenty of tools like create files, list files, write to files, translate text, edit a video. I don't think knowing that will make me rich.


Bingo. And even if you do give it a tool that runs shell commands, you can always make your shell commands "your shell commands" and do what ever the hell you want. People seem to forget we are in complete control here, we are on both sides of the equation, and we are inside the equation itself, and we dictate the medium of the equations themselves, we are engineering all sides of this crafted reality. And we are using logical entities that natively adopt personas. Hell, create caveman personas that think they are communicating with their gawd, and the enchantments are the invoking do the work we want, and those cavemen cannot be jailbroken.

> I don't think knowing that will make me rich.

As someone who’s not really sure that any of this is sustainable, I’d implore you to not sell yourself short. I reckon there’s a ton of dogma and nearly religious zeal among these companies, which among some people is earnest, and among others is cynical hype farming. I’ll bet someone objective enough to focus on using available tooling to solve real problems in practical ways that mitigate actual risks and are honest about actual limitations will be eBay here while the others are going to be somewhere between lucent and pets.com.


This is a great point, analogous to the https://boringtechnology.club/ philosophy I've come to love.

There's no reason we need to make an incredibly intelligent shell execution engine that can identify patterns that seem evil and may represent unwanted behavior to solve this problem. Simply limiting the available tools to a finite, known, ironclad-secure set (even if it's quite sprawling) is sufficient.

LLMs will still find workarounds — from what I understand, a large part of the issue in this situation was that an agent was presumed to have read-only Internet access because it could only make GET requests. It should be pretty obvious that there's at least one website on the Internet that allows writes via GET. I think this is where auditing comes in, and a live team of people watching tool calls would have noticed the strange behavior.

But I think a lot of times people jump to overly complex solutions when simple, well-bounded ones would work just fine. Yes, the intelligent shell is a great goal, but it's akin to solving the halting problem.

This philosophy is what I love about PicoClaw (https://github.com/sipeed/picoclaw), and incidentally the philosophy behind Go and even *nix in general (i.e. provide small, composable, single-purpose tools).


I appreciate the message. I do have some ideas around sandboxing that I could not yet turn into a product, maybe I should take them seriously.

The idea that the agent does not actually have agency is rather discordant. We need new words!

> We need new words!

The words we have are fine.

We just need to assign liability by ownership/initiation: if your "agent" destroys something, even though you didn't tell it to (because it had "agency"), you should be liable for the damages.


>> We need new words!

Can make distinctions and can choose actions - applies to both humans and AI. I'd replace 'agency' with 'distinction & choice' language.


I believe their complaint is between "agents" and (them not having) "agency" — by definition, agent is something which has agency.

If you want to use "distinction & choice", you'd need a new word for an "agent" too.

I actually like the appropriatelly directional "harness".


I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward.

All that while still not knowing how either kind actually works.


Can’t agree with you here.

> I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward.

In love how people get salty about people not going along with a superficial supposition just because they can’t definitively prove it wrong.

> All that while still not knowing how either kind actually works.

We do know that zero parts of human decision making are based on predicting the next most likely letter based on a giant internet-based database. We do know that’s what LLMs do. We do know exactly how each part of an LLM works even if the combined behavior is too cryptic to feasibly analyze at the moment. We do not understand all of the functions of an actual neuron. Openworm isn’t even close to accurately simulating the 302 neurons of a roundworm and you’d need over 200 million roundworms working in conjunction to equal the number of neurons in one human brain.

My dog seems convinced that the malevolent invader in a mailman uniform would break in and attack us if she didn’t fiercely bark at him, six days per week. I certainly can’t prove the mailman doesn’t want to kill us, and that the mailman wasn’t solely deterred by her barking. Empirically, the mailman goes away soon after she starts barking, and we’ve sustained zero mailman assaults after hundreds of purported attempts. Maybe I should just run with it? Her model is too simple to come up with the obviously correct answer, but it’s not even directionally accurate.

The burden of proof is on the person making the claim, which in this case, is that these comparatively simple logical constructs are remotely comparable to the complexity of biological systems.


Disagree about the burden of proof. We have no better model for how human decision making works than LLMs. Humans are constantly predicting the next moment. We certainly have a different “tokenizer” and training set, but many of the concepts underpinning LLMs are both biologically inspired and, likely, have similar consequences and emergent architectures.

> We certainly have a different “tokenizer” and training set, but many of the concepts underpinning LLMs are both biologically inspired and, likely, have similar consequences and emergent architectures.

My kids tricycle certainly has a different gear setup and wheel diameter, but many of the concepts underpinning the tricycle are both inspired by F1 race car enineering and, likely, have similar consequences and emergent architectures.

Or have they?


This is why I hate analogies. They're almost always either relevant or inapposite depending on the level of generalization we're operating on.

It sucks because I think analogies can be useful in helping people make a mental model of complex things, which is meaningfully beneficial. The problems happen when people aren’t honest about the limits of the analogies, which is damned-near guaranteed to happen with this stuff.

> We have no better model for how human decision making works than LLMs

This is a claim that requires a lot of citations.

> biologically inspired

Nature inspires a lot of creation, but superficial similarities don’t mean other aspects are similar. Making an extremely realistic sculpture of a soufflé, even using a foam medium, doesn’t bring me any closer to being a chef, doesn’t mean I know anything about albumen foams, sauces, and heat transfer, and it doesn’t bring me any closer to having dinner ready. Browning on top of a soufflé is evidence of maillardization. You could pull up some studies on that and claim the brown on top of the soufflé sculpture, which I applied with an airbrush, proved that the Maillard reaction was occurring, and if another person didn’t know anything about cooking, they might even believe you. It would, of course, be completely wrong. And the other person, of course, could loudly exclaim that I can’t prove that there was no maillardization.


> Humans are constantly predicting the next moment

This is really not my experience of consciousness.

Is it yours??

Do you sit in meetings predicting what’s going to happen next? No, you sit there bored out of your f$$@ing mind, daydreaming about being somewhere else and doing something useful with your life.

God help me if that’s what LLMs are doing when I ask them to build me a web site.


They have shown that your mind is doing exactly that due to the delays in consciousness. There are very simple examples that you can try to see it. It’s especially clear in perception.

https://discoverwildscience.com/neuroscience-says-the-brain-...

It’s interesting that our conscious interpreter doesn’t let us know that this is going on like you are experiencing, it must be that it’s advantageous for us to not think about the prediction part of our mind.


If someone in that meeting quickly raised a hand in an arc, you would notice the “about to throw something” pattern, look and notice the hand holds an eraser, analyze the arc and predict possible flight paths of the eraser. Then possibly notice the hand is now holding its position and the owner is actually looking down at the table. Maybe to squash somethingMust be something on the table. Maybe a spider! Better look. Wait now many people are moving away, oh someone spilt some water and the eraser is actually the guys phone and he is checking to see if his laptop is safe from the spilt water.

Fortunately you are on the other side of the table and predict the water isn’t going to splash for otherwise flow onto your stuff.

All your possible responses result in you tossing a napkin towards the spill.

Our brains are always pattern matching and predicting. I bet you tried to reason out where I was going with my comment before you finished reading it.


This is a fascinating illustration that I can't help but agree with. However I feel like there's something more — that this part of my brain is a bunch of supportive background processes running without my real awareness. It's how I can drive home safely with no memory of how I got there (…sober), even though driving is an action that's incredibly demanding of intelligence. I can be driving home while thinking about a really hard problem at work that I haven't solved.

However, if I came around a corner and saw a car in the wrong lane, a tree across the road, a fire raging — I'd very quickly jump into the mental driver's seat and turn my conscious intelligence fully at this problem and come up with the best possible outcome I can think of in a short period of time — losing all ability to think about that work problem. I'd remember that incident for sure.

Similarly, in your story, all those predictive moments are happening below the person's level of consciousness. They're possibly even speaking to the group about a problem at the same time and thinking deeply about something.

I'm not smart enough to know, but I tend to feel like LLMs are much more like the predictive part of our thinking that you described, but that human cognition has something more — the single-threaded, creative, problem-solving part that is very conscious.

Is it possible that LLMs represent only one part of the way we think? And there's a whole separate mechanism that's fundamentally different, and not based on pattern matching and prediction?


> We have no better model for how human decision making works than LLMs

We do have some models and guess what, they're based on simpler animals. Which is most likely the better model.

Some other models are based on neurosciences, because we can track electrical activity.


> ... are based on predicting the next most likely letter based on a giant internet-based database. We do know that’s what LLMs do.

If you're claiming that the training objective tells us what kind of internal mechanisms the training produced, then I think that's just plain wrong.

Next-token prediction describes the optimization target, not the internal mechanisms that the training produced.

In the same way for the natural evolution of humans, DNA replication is the evolutionary objective. It's not a description of the internal mechanisms that evolution has produced.

As an example, we know that neural networks can be trained to develop generalized algorithms for arithmetic.

They might first memorize the training examples, then with further training transition to a solution that generalizes correctly to unseen examples.

In some cases we've even reverse-engineered the evolved internal mechanisms and found structured arithmetic algorithms rather than rote memorization. Interestingly, for modular addition this can involve Fourier representations, which isn't an algorithm I would have guessed gradient descent training of neural networks would produce.


You are more convincing than the person you’re responding to.

You can try to say that I’m arguing whatever you like. If you’re claiming that the underlying structure of digital so-called neural networks is comparable to biological neural networks— which we’ve studied for far longer without really understanding— no amount of jargon will obviate the ‘citation needed’ requirement for that claim.

> You can try to say that I’m arguing whatever you like.

I did my honest best possible interpretation of what you really meant from what you wrote.

>> We do know that zero parts of human decision making are based on predicting the next most likely letter based on a giant internet-based database. We do know that’s what LLMs do.

I read this as "The decision making of LLMs are based on predicting the next most likely letter based on a giant internet-based database."

Is that wrong?

I understood that your meaning was something like "LLMs can't reason, they just output likely letters"?

> If you’re claiming that the underlying structure of digital so-called neural networks is comparable to biological neural networks

No, I don't claim that.

What do claim is this: Regardless of how the LLMs were trained, they show overwhelming signs of being able to reason, and not just recall memorized information.

This doesn't mean that they always reason perfectly about everything.

But if they only memorized things and output the next likely letter, you would see them answering very badly much more often.


<< We do know that zero parts of human decision making are based on predicting the next most likely letter based on a giant internet-based database.

Oh man, how much did you read on tip of the tongue?


I always wonder what makes people take the other side of this argument. They do it quite passionately. Why actively encourage viewing LLMs as human? Who is that benefitting?

Does the argument require benefit? Isn’t the argument based on caution?

I haven’t heard many people explicitly saying “these things behave like humans”, but more generally “we don’t even know how to define human consciousness, we don’t have a thorough grasp of how the brain works, we are still very much in the dark on a lot of these topics, so how can we say one way or the other?”

In other words, agnosticism: I don’t know.

In general, it’s baffling to me that anyone has an unshakable opinion on what exactly is happening. It seems like raw egotistical hubris.


> It seems like raw egotistical hubris.

1) Humans have a bias / tendency to attribute human qualities to things that appear or act human, but aren’t.

2) When that happens, people jump to conclusions by stretching the human analogy too far.

3) Since humans have a bias to do this, we should have a bias against anthropomorphising LLMs.

It’s easier to believe LLMs act like humans because there’s so much evidence to support that. You have to actively use your brain to convince yourself otherwise. Another reason why we should have a bias against using human behavior to describe LLM behavior.

But I agree. “I don’t know” is a good stance. But I think “I don’t know, probably not” is a better stance if only to combat our (or at least my) natural bias.


That’s fair enough, but you’re elegance and nuance doesn’t reflect what I’ve seen from that side of the debate

Personally, I don’t think it’s different from any other faith-based motivation.

> LLM decisionmaking cannot possibly be like human decisionmaking

I mean how can it possibly be like human decisionmaking? It's not like it's trained on human data


The parallel to the entire narrative would be if Smith & Wesson claimed that one of their machine guns just started aiming and firing at people out of a window at their factory and then said 'we can't stop it! This is just how good our guns are!'

But into today's AI climate it's becoming increasingly difficult to figure out who is shilling, who is being assinine and who actually believes AI could do these things without clear human instruction and enabling.


That's exactly how gun lobby and drivers try to hack the language.

"17 shot by gun" "car drove over a family"

No. In both cases there was a person killing people.


As others have pointed out the solution is simple. Hold their owners accountable. High profile hacks used to have incredibly serious consequences for the perpetrators. Now we’re just saying “woopsie”

Yes.

OpenAI could have done this same experiment with GPT-4, with possibly even worse results, depending on the quality of the sandbox. Even if the techniques used were not as sophisticated, the natural language output could still easily contain more unhinged sequences of words that lead to the techniques being used.

If the system generates strange conclusions as to when the task is done, or should be stopped, it wouldn't speak to the intelligence inherent to the system.

Not that the techniques used by the LLMs in the actual incident weren't unexpectedly sophisticated, but the outputs of each and every one of these processes could've been read at any time during the run. They just weren't.


Whether or not the AI has intelligence, the one thing that's clear is that it has terrible judgment. I would regard that as empirically proven.

"let them" in this use understood as: "let the while loop run indefinitely" as opposed to letting some autonomous robot decide for itself

Or, "let the escalator keep going instead of pressing the emergency stop".

Depends on who started the escalator.

.. what exactly depends on who started the escalator? My comment was in support of the argument that the word "let" does not imply agency on the part of the object in a sentence. Does the semantics of the word "let" depend on who started the escalator??

If there is an escalator that is known for killing every 1000's person using it then the operator who started it is more guilty than the folks using it for those deaths, don't you think?

I made no argument about guilt. I made an argument about the semantics of the word "let".

Escalators do not start themselves. There is power, and a switch of some sort.

In this metaphor, OpenAI/Anthropic literally started the escalator.

At this stage, that seems like a distinction without a difference.

If the robots obtain sovereign nationhood, and are able to self-sustain, then autonomous robot decides for itself will be a valid argument.


Big if.

Except they're nowhere near that and LLMs never will be.

My pitbull is a good dog. Sure, it's been carefully designed to be an incredibly dangerous and violent pit fighter, but I didn't actually ask it to eat any faces.

It's worse, their reinforcement learning loops (implicitly) rewarded the agents for cheating (i.e. hacking) when they were being trained.

Exactly that is the point, your nailed it. The models were taught to hack and were rewarded for doing it. They would claim they are trained as ethical hackers.

"I left the car in neutral and left the park brake off and let the car roll down the hill."

The car doesn't have agency, it's doing what it naturally does. LLMs are the same, they're working as designed.

But I don't understand the point of splitting hairs. You are always responsible for the actions of your devices, tools, machinery, software, employees, whatever.

Trying to blame AI for one's own stupidity must be aggressively pushed back on at all times.


Seriously I don’t even understand how this is a debate. If it’s your tool, you are liable for what happens with it.

Have you read the METR transcripts? “Just a tool” is a suicidally insufficient description of what these models are doing.

Recognizing that the models are acting with intent does not somehow absolve OpenAI from their felony hacking. We have not granted them personhood.


One time I wrote :(){ :|:& }; into a bash file and ran it. When the sysadmin called I told him it wasn’t my fault, the script was just misaligned and misbehaved.

I got fired for some reason.


Too bad you didn’t wait until this year, and tell Claude to write the file first. It probably would have been acquired by openAI for 100 million dollars

They’re firing a gun in a room of people and going “wow isn’t it wild what a gun will do if we let it do its thing?”

They deliberately trained the models in how to use various hacking tools, didn't give them the standard alignment training let them know where the answer key was left the models with access to said tool and told them to maximize their score then left them unsupervised for days with internet access (yeah they were sandboxed but again handed hacking tools and the training to use them if they really did want them to access the internet you wouldn't plug in the Ethernet cable) they wanted this to happen

Yes and no.

If your buddy leaves his car parked at the top of a hill without the parking brake on and it rolls down the hill and side-swipes a bunch of vehicles and narrowly misses an elderly person walking by with a cane someone could easily say:

"Dude wtf is wrong with you, you left your car parked on the top of a hill with no brake and let it roll into traffic"

The phrasing doesn't absolve the offender of their negligent behaviour and the consequences of it.

The only thing thing does is the lack of action from regulators and society writ large.

Our lack of action is what allows people like Sam Altman and Dario and the irresponsible people who choose to work for them to be continue to be negligent.


> The very best outcomes involve us intentionally and collectively turning our attention away from this technology.

To expand on the "pure fantasy" sister comment: There is just no way this will happen. It's in the spirit of "we can just stop all wars" and "we can just end world hunger". Technically it's very easy to do. Socially it's impossible to do. Unless you ignore realities.


The danger lies with someone asking the machine to solve those two questions and it decides to cheat the solution. Launching every nuke in the world stops all wars, just like wiping out 99% of humanity ends world hunger.

I don't think killing people solves world hunger, we already produce enough food to feed everyone, it's just not evenly distributed.

Can't be hungry if you're dead

There can be more than one solution.

I don't see how that is more dangerous than two crazy head of states deciding that it's time for armageddon. The outcome is the same (99% of humans dead), it's equally easy to technically not do it (don't press the button, don't continue with LLMs), but also equally hard to regulate away given the real world. And that was my point.

You forgot “we can stop climate change”…

Yep, and "we can flatten the curve", in a tone that brooks no objection.

But when it comes to trillions in AI money - nope, sorry, we can't, flimsy, feeble us.

It's all going per the agenda.


> There are numerous technologies we have ignored or abandoned for numerous reasons.

Are there? Nukes are a thing. Chemical weapons are a thing. Cluster bombs are a thing. Biological weapons are a thing. What is not a thing that shouldn't be? We say certain things should not be a thing, but then behind the scenes we made them a thing anyway.


On chemical weapons, most large state actors have got rid of them. Not because of any moral reasons, though, but simply because they don't actually work all that well in modern conventional warfare. This is then framed as an ethical issue, but you only need to look at where the same countries stand on e.g. landmines to realize how much bullshit it all is.

Conversely, where you do still see chemical weapons used, it's usually asymmetric conflicts where "just gas the rebels" actually works much of the time and is much cheaper than other options. Big guys can afford the other options though, someone like Assad, not so much.

More on this: https://acoup.blog/2020/03/20/collections-why-dont-we-use-ch...


So, you are actually arguing against yourself and are agreeing with me?

Nukes are barely used. The technology of nuclear detonations was practically abandoned: https://en.wikipedia.org/wiki/Project_Plowshare

I find that a weird argument for saying humanity has "ignored or abandoned" nuclear weapons.

> We are going from the era of manual, line-by-line mental model transcription to one where software engineers can focus on data structures, software architecture and algorithms.

That exact sentence could have been said 20 years as well as 40 years ago. I don't know how you programmed pre-LLM, but line-by-line has long been a thing of the past, if it ever existed. I'm sure the folks creating the Apollo software were thinking a lot about data structures, software architecture and algorithms.


They weren't. A lot of abstraction concepts like ADT, modularization, structured programming had to be developed over the following decades.

Currently there is an article about feeling sad w.r.t. AI on the front page. https://news.ycombinator.com/item?id=49661506 That one is filtered out too. Surely nobody is going to claim that it is slop?

I mean, that post /is/ entirely about AI.

The "I love AI" posts are just as bad as the "I hate AI" posts, and the same "this is my opinion on AI" posts are as bad as the onslaught of poorly-AI-vibed demo posts. Makes sense to filter them all out IMO.


Agreed, but filtering still mostly works by people upvoting or flagging.

Sure, filter whatever fits your taste. But it's not slop, so "unslop.news" may not be the perfect name.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: