Hacker Newsnew | past | comments | ask | show | jobs | submit | rsfern's commentslogin

I find this troubling, there are so many potential ethical issues with this application of language models. Maybe it would be more appropriate for models to issue safety refusals and help the user learn how to get help from a licensed mental health professional

In the United States at least there’s a huge lack of mental health professionals, they’re expensive, insurance goes out of their way to make it inaccessible, and wait times for appointments can be months out. For psychiatrists in some regions there are zero that accept insurance and there can be a six or longer month wait. Many regions have no inpatient beds or any sort.

My parents are clinical psychologists with 40 years experience, and they both agree LLMs could be a great boon to many many unserved people if made well. They think this is a wonderful avenue of research. But agree in the moment it’s not ready. I think the ethics of allowing people to suffer due to economics while ignoring a technology that can help them is worse, no? Telling them to seek a health professional when none are available is worse ?


I agree that accessibility is a big problem, and take your point about ignoring technology that could help, but if the technology if not there yet then I think more research and discourse is warranted before deploying it widely. As a society we need to decide what kind of requirements we expect of these systems, then agree on standards of performance

Is it ok for models to offer specific mental health interventions, or just provide neutral information? To what degree is that technically achievable? Does the answer depend on the scenario?

Health care workers have mandated reporting responsibilities (in the US, I assume other countries are similar) for some situations. Should models be bound by this responsibility as well? (I think yes but it is a thorny question). How to achieve that while preserving patient privacy for issues that don’t fall under mandated reporting rules?

We also have standards of care for health workers. Benchmarks are good, but that’s similar to a licensing requirement, and we have accountability mechanisms when people don’t adhere to their ethical and professional duties. What should be the liability situation when a model responds out of standard resulting in harm (to the standard of evidence we would hold a human health worker)?

Maybe I am not plugged in to this space enough, but my impression is these questions aren’t really in the foreground


I think the expectation that ChatGPT will be your therapist is wrong, however the expectation that you can build a therapeutic system using LLMs is not necessarily wrong. You are right that there are regulatory aspects, and liability, as well as guard rails built into our system (but less for mental health than medical as it’s largely a neglected aspect of our care delivery).

Also, it’s a different mode of delivery. Some people work well remotely, others better in person. Likewise some people open up more easily to a machine than a person. One doesn’t replace the other.

Further, LLMs can be very useful for diagnostic work in addition to a psychologist to help accurately classify mental and personality disorders. This is my fathers speciality and he finds modern LLMs to be fairly adept at classifying.

A lot of psychologist work in a clinical setting is case notes and an LLM can be quite skilled at producing standard expected case note structures from clinical notes, improving the time spent by the professional in care vs paperwork.

All of these are opportunities for use of LLMs as specialist models, agentic frameworks around general models, guard rails, delivery mechanisms, etc. I wouldn’t look at the world being OpenAI and Anthropic, but as model providers as providing screws and bolts and it’s up to us to build the products. LLMS are trained with an enormous corpus of witting on mental illness, but also writings by the mentally ill, care professionals, and all sorts of other materials. By eliciting all these dimensions the applications to mental health are enormous - it is a giant learned model of human thought and experience, not just of factual information, but the entire complex.

Here’s an idea I’ve talked to my dad about - how about using a model set trained on the writing of many people of specific mental illnesses. Then you can interact with them and they will produce responses likely of someone with that illness. This can be used to train clinicians, but also to build better inventories and diagnostics since you know the classification of the responses to some reasonable degree. This should accelerate research into mental illnesses.

The possibilities are vast, limited by our imagination and biases.


> Likewise some people open up more easily to a machine than a person. One doesn’t replace the other.

This is a fundamental personal failure. Not being able to open up to a professional whose purpose is to help you. The first step is to be honest with yourself, then honest with the person whose entire professional purpose is to help you.

AI psychologists will only ever be able to do so much, but I imagine it will be a whole lot of “you’re right” with very few harsh truths, criticisms, or personal development.


Dude, in that lens, all mental health is a personal failure.

The fact is a lot of people have trust issues, anxieties, social issues. Some are paranoid, suffer extreme neuroses, or panic in a social situation. Some have felt betrayed by professionals seeking to help them, putting them into inpatient care against their wishes, or medicating them with debilitating medications to sedate them, which is common for people experiencing mania. That’s not a personal failure, that’s just personal issue. And someone seeking mental help is having personal issues more often than not, unless they’re just seeking life coaching.

Ascribing these as fundamental personal failures, and therefore they deserve their personal hells with no relief unless they comply with your view of how people must be, is cruel.

The sycophantic behavior you’re outlining is a result of reinforcement learning and alignment, not a fundamental property of generative language models.


> I think the ethics of allowing people to suffer due to economics while ignoring a technology that can help them is worse, no? Telling them to seek a health professional when none are available is worse ?

Not necessarily. Half-measures—especially those that try to offer technological solutions to social problems—allow the actual root problems to remain unaddressed, fester, and grow. The ACA was a hack for a failing system, and then when that quit working, more hacks were added on top, and on and on. LLMs are just another hack that let the can get kicked down the road once more.

If and until actual meaningful health care system reforms happen, the rich will continue getting their $63,000 14-day concierge service from Mass General Brigham whenever they feel blue, and the rest of us will end up getting Doctor Amazon telling us to practice deep breathing and buy some supplements.


I can’t argue this, but I can build an LLM based care provider today with likely decent results in patient outcomes for those adapted to it and good cost dynamics, and across a variety of critical use cases in the care vertical from providers to end delivery. I can not however change American society, our economic structure, or political dysfunction single handedly, can you? Should I not do the former and let people suffer in protest because the latter is not possible, and hold my breath hoping it’ll change? How many lives must be destroyed in protest hoping everything changes first?

Fine, let’s have socialized delivery of mental health care. Let’s go back in time 20 years and establish a pipeline of capable providers that doesn’t exist. Let’s do those things if we can. But I don’t see it happening in my life time.

I do however see an opportunity to take the technologies presented to use in the immediate present and providing a mechanism to bring help to people who need it today with good efficacy where it’s appropriate. Is it OpenAI that brings that? I doubt it - they’re selling shovels, not digging for gold. Is it a general model that provides direct care? Doubtful IMO. Is it a chatbot interface ? I doubt that too. But I think if we are creative and thoughtful we can use these technologies to address huge swaths of the care providing problem, which is more than talk therapy. And we can do it today, and don’t have to wait for MAGA/MAHA to accept reality on realities terms.


The transcripts of mental health professionals could likely have an improvement in the baseline below-average experiences of most folks with such supports.

In other words, mental health professionals will resist this technology to replace them until they realize they can put it in front of people like a new kind of web form, at which time the same will become amazing and wonderful.

At the very least, a tool like this could serve as an anti-virus and firewall for poor and harmful experiences from mental health professionals towards clients.


I think this is dramatized to the point it’s talking past the article, the math community isn’t really making any of these arguments from what I can tell.

The discourse is (1) models are capable of making really impressive mathematical advances, usefulness is not in dispute, (2) the frontier AI companies aren’t being super transparent about information sources so it’s hard to know exactly how to evaluate the level of capability that was demonstrated, and (3) there are lots of kinds of math that is interesting and there are open questions about how to get there.

In particular this article highlights a particular open question I’ve seen discussed on HN before, which is that the particular proof strategy of finding a counterexample might be more amenable to RL than other strategies of proof that might be needed to resolve the other branches of the Navier Stokes problem (and probably other similar areas of math)


yeah i agree its dramatized, but the situation was quite dramatized by the parties involved as well. i just find it quite funny, that the perceived drama might play out like this now.

Is it really impressive, or just kind of interesting?

If they spent about 10 GWh solving the problem (was it solved?) then that is much much more than 500 lifetimes of a human brain working.


A specific version of the problem has been solved. But not the broader, harder version. The article explains that fairly clearly.

I’m very anti AI and OpenAI, and do think it’s a pretty interesting finding! Very likely not worth their spend, but interesting and novel nonetheless the less


yeah but the problem they solved is not really useful. the article is also quite clear on that

According to OpenAI, but they haven’t exactly been transparent about what information the prompt entailed.

The bigger question is to what extent did expert mathematicians metaprompt the model with fruitful solution strategies through their sessions finding their way into training data. Answering that question definitively is kind of important for understanding the models contribution/capability. But I feel like people want to turn this into a debate about priority and credit which is sort of secondary


I think that’s a lot of risk of anchoring reviewer bias. I’d be more comfortable with a triaged review where the editor’s office uses models to score whether a human editor should evaluate a paper to potentially send out for review, then the editor makes their own assessment, and the reviewers continue to do their job unassisted

For me the convenience of “hey remember this fact” is outweighed by a desire not to get stuck in a search or context bubble.

It might be nice to have better UI to control which bits of history get added to the context of a chat, but then just use a coding harness instead of web UI


Maybe. But I don't just use it for coding. For instance one thing I use it for is helping me track and evolve my workouts over time. It helps that it can remember the weight and number of reps I did the previous time. I could make an app for this, but it works well to do it purely within chat, with memory, and that's convenient. This is just an example, I use it this way for a number of things. I'm not going to use a coding harness for this kind of trivial use case!

It should resonate here. The thesis is that good design principles are transferable between humans and agents. So focus on solving important problems and build well designed tools and documentation to get there, you don’t need special design considerations just because agents

I don’t know about the startup space, but in AI for science people are spending time building MCP wrappers around poorly designed APIs instead of redesigning the API or building an abstraction layer that humans can also use. That seems like a mis-allocation of effort


There’s no reason it has to be like that. We could change course one more time back to a stable and non-partisan science funding landscape and then stick with it if we choose to. Even some of the alarm in this article over the proposed funding cuts isn’t set in stone, the presidential budget request from last year aimed to cut NSF by a similar amount and Congress didn’t go for it

Only to be reversed by the next MAGA politician when the administration from the “other team” fails to immediately address the economic impact of decisions made in 2025.

We vote on vibes and popularity now, not qualifications.


On the contrary, I think the chess comparison is on point. We’re discussing observations that even the strongest models devolve into making invalid moves without scaffolding. For me that raises the question of whether these models are learning the rules and generalizing from them, or of they’re just pattern matching and flailing on this task. Maybe the reality is somewhere in between, but the benchmarks don’t seem to directly measure conceptual generalization, they measure task completion. They can disrupt a lot of people and industries by pattern matching and flailing without being AGI.

I’m sure these models know the rules and can explain them when prompted, but that doesn’t seem to be the way they actually complete this task. Will they get there? Maybe


the discussion isn’t really about whether language models can become strong chess players though, the point is they seem to struggle to consistently make valid moves. Most humans don’t need to read two books to pick that up, just a couple lines of basic instructions

That has not been my experience with new players, they regularly make invalid or incorrect moves even after detailed instructions especially in novel situations.

Maybe it depends on the person? My six year old isn’t great at strategy but they can pretty consistently make valid moves. Sometimes they ask for confirmation on a move which is also not a trait I see in language models (at least unprompted)

Your child never messed up en passant (or had trouble understanding it in a real game), castled through or into a check, didn't see a discovered check after moving a piece, never got confused by how stalemate works?

For numerical code I like einops.reduce more than numpy/pytorch sum reductions because you can reduce over named dimensions. It’s much more readable than having to reason through axis indexing again every time you come back to the code

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: