Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> "The exact words we choose when writing matter."

Then write your own damn text if you care about the exact wording so much



His complaint seems to center on conversations between him and the LLM, not copy being written for publication elsewhere. That is, if Claude is going to teach him a new skill, he wants it to pick the most accurate words possible, not "pretty accurate words, subject to watermarking techniques".


If the exact words we choose when writing matter so much, then why use a non-deterministic LLM that produces slightly different output on every run?


Additionally, LLM are already rolling the dice on different wordings so I don't see how this watermarking makes it any less precise.


You are responding to the original article, authored by John Gruber, who, last I checked, has been writing his own damn text multiple times every damn day for multiple decades as a primary vocation.

I’m willing to venture that Gruber is on the list of folks that get to hold the opinion choice of words matters.


Holding that opinion is fine; holding that opinion while using LLMs to probabilistically choose words for you doesn't add up. You want carefully chosen words? Read a book, or write your own. An LLM is not carefully choosing anything, it's providing words based on probabilities.


He's a professional writer. As the article says, even if you write your own words, this is still a problem with proofreading, copy-pasting references or quotes, and "AI checkers".


I still don't get why a professional writer would ever accept an AI proofreader over a real one, or a "quote" returned from an LLM over one sourced from a reputable reference.


Is the concern that the AI will rephrase the text during proofreading, thus falsifying a copy-pasted quote? I'm not sure how such a change could go unnoticed, unless you blindly publish the AI output without looking at it, which doesn't sound like something someone would do who wants to “write their own words”.


The exact words matter to people who people who don’t use it to write for them. A lot of people use Claude as a friend/therapist/romantic partner/etc. That’s who I imagine would be most affected


>A lot of people use Claude as a friend/therapist/romantic partner

People developing a para-social (pseudo-social?) relationship with a corporate robot have far bigger problems than the word-chooser in their robot "friend".


Very intelligent people need intelligent-others to bounce ideas off of, and the LLM can be that.


Very intelligent people should be able to grasp that a LLM is not intelligent.


Have you tried to "rubber duck" an idea with an LLM. The latest models are pretty damnably good at it.


I have, and I still do this occasionally. I don’t believe the plant on my desk (I don’t have a rubber duck) is intelligent. I don’t need it to be intelligent, either.

I also do it with a LLM every now and then and, while it’s feedback is more useful than a toy’s, it does not need to be intelligent either.


Even intelligent people have demonstrated they are not immune to the damaging effects of AI sycophancy. The most recent example that comes to mind is Hank Green.


Yes, I use LLMs this way, they're not friends, not therapists, not persons, they're tools, like a word/language calculator.


LLMs never are limited to "exact words", because their output is inherently probabilistic. The method Anthropic (along with Gemini, who has been using the exact same watermark for at least a year) uses doesn't bias the output token distribution, just reseeds the PRNG in a way that can be detected after the fact: https://www.anthropic.com/news/claude-text-watermark#which-s...


> The exact words matter to people who people who don’t use it to write for them.

Right. If the exact words matter, using a non-deterministic LLM is a terrible idea in the first place. I hope these people never try putting the same prompt in different sessions.

Also, now I am curious. How would these people interact with other humans? Is there anyone on earth who would provide the exact same reply, down to every single word, if we asked them the same question more than once?


How would you ever know if you've been affected? How, indeed, would you know at all whether the word you get next is different from some different word next?


Affected by what exactly?


OpenAI released a feature that lets you tweak ChatGPT output to sound more human and like yourself and correct inaccuracies:

http://blog.tyrannyofthemouse.com/2026/04/open-ai-strikes-ba...


I mean this is the thing that really comes off hard.

If you want precision and clarity of your writing, then you need to hand write it. Just like when you are optimising, its common to drop to a lower level language because the compiler doesn't express what you want. Sure its hard, but you know, thats kinda the point.

Even if you don't want that, the LLM is an average of the style it was trained to give out. Which is a homogenisation of the language to create a vague padding medium between a few generalised facts. (because a. it makes it less jarring when stuff is wrong, because its smeared over a higher amount of text and b. it looks more 'professional' because American business English is all guff and no meat)

Also yes, two isolated phrases may have subtly different meaning, frankly, the nuance is missed on most people. If you look at the interactions on here, at least 25% of the arguments are caused by people angrily reacting to the things _they_ thought the other person was saying, rather than what the actual person was saying.

So no its not a perversion, the LLM is, if you're gonna be picky about things.


The bigger problem is its effect on the way we use language. The amount of LLM generated or edited media is going to keep increasing and the media that people consume affects their own word choice.


> affects their own word choice.

exactly. in the same way that printed books affected word choice, so did the radio.


I can't remember any radio determining words or adjusting grammar of the person speaking through it. Neither can I recall there has ever been a moveable type press, laser printer, or inkjet which bastardised the words of authors.

These were mediums _through_ which communication happened. Language models, large or small, are not any such medium.


> Neither can I recall there has ever been a moveable type press,

Ah my friend, you are about to fall down a rabbit hole into standardisation of spelling, and the sometimes deadly debates about how to translate latin into the vernacular.

English, as she is written, is a great example.

for the spoken word, BBC/received pronunciation is another. I speak the way I do _because_ of BBC radio. The reason I have the accent I do is because I changed it to fit what the BBC put out, rather than what my local (impenetrable) dialect was.

You have to remember that your language is shaped by those around you when you are young. So if you are in an insular community, it will be reflected in your language. If I was a journalist, or hell, just me, I wouldn't be letting an LLM speak for me. So the bastardisation of my voice is down to me, not the machine.

but again, your argument is against LLMs and globalisation of culture, not finger printing.


This is very poignant. The entities that controlled the content broadcast over radio and TV or what was published in books and newspapers had an enormous impact on both culture and language. Radio and print didn’t just transmit the voice of a single person, the message was shaped by whole teams of editors, bureaucrats, politicians, censors and the necessities and constraints of the technology. Any message that a person tried to transmit through these channels was transformed more than it would be if you ran it through an LLM or more likely not transmitted at all.


There are two different things being talked about here, one of which is "what effect does widespread LLM use have on a culture", and one of which is "what is the effect on a specific text of running it through an LLM".


A radio device at BBC did all this and not humans?

> again, your argument is against LLMs and globalisation of culture, not finger printing.

ABSOLUTELY NOT. Do not ever put words in my mouth. My argument is that radio is a medium through which humans communicate. An LLM is not.


[1] > A radio device at BBC did all this and not humans?

Technology mediates (human) agency.

[2] > My argument is that radio is a medium through which humans communicate. An LLM is not.

KaiserPro's argument is that radio (a one-way medium) mediates how humans communicate (phonetically), thus influencing how people speak.

Now that does not explain how the printing press and radio have influenced word choice (both of which I would like to see examples of!)

[3] > I can't remember any radio determining words or adjusting grammar of the person speaking through it.

Now media themselves did not really have something that looked like the kind of pseudo-agency that LLMs seemingly have. There may be some kind of qualitative leap.


I mean I get your point, but sadly LLMs are a tool by which humans communicate.

In the same way that handwriting conveys more information about the writer than type, typing ones own thoughts conveys more information about the writer than prompting an LLM.

The analogy here is hiring a speech writer to do your speeches, or dictating to a skilled typist.

I understand the vociferousness in push back


What are you talking about? I understand OP has the perspective of writers, but let’s say you’re asking an LLM to explain a concept and it uses green words that are actually more difficult to understand. Or you ask for an analogy to explain a topic and the analogy doesn’t quite land because the description used “gray” instead of “overcast”.


Watermarking or not, LLM are already using RNG to pick a variant between different words/expressions.


I'm sure he honestly appreciates the pushback.


I am sure he is quite happy to have people disagree with him on HN. He usually wears it as a badge of pride. We might even have a follow-up in a couple of days about how these techy weirdos lost the plot.


Indeed.


:D

(For the record, I read DF quite often as a Mac-minded techy weirdo, I just happen to disagree on this particular issue)


why you act so tough bro?


What an absurd take on a very reasonable concern. As much as I hate “obviously AI” writing, I don’t understand why we have to handicap the tech and prevent it from improving. This manipulation of word choices virtually guarantees AI will always sound like a robot.


Have you compared pre- and post-watermarked text to make that assertion? And I'm not sure how writing your own text if you care deeply about the word choice is handicapping the technology?


The entire watermarking scheme is based on replacing a random number generator with a seeded random number generator.

This cannot change the "voice" of the LLM. It was already letting a random number generator choose which adjectives to use. Now that random number generator encodes a tiny signal.

But fundamentally the way it writes has not changed.

It's not making it choose different words. It's a minor change to how it chooses between multiple nearly identical words, where in the current case it literally flips a coin.


I don't think there's been a single version of the major LLM providers that haven't handicapped the tech since the beginning by changing temperature/top-P/frequency penalty/etc. All of those stray from "highest probability" token selection. What's funny is it was done specifically to make it sound less like a robot/deterministic.


> This manipulation of word choices virtually guarantees AI will always sound like a robot.

You make it sound as if that were a bad thing.


Reasonable? A concern which is based on no real data?


With all things going on among AI bros and the AI industry as a whole, are you really that surprised there is a widespread aversion against the tech?


Do people not realize this will apply to ALL Claude output, not just writing you ask it to produce?

> Marks will apply to output from supported Claude models across Claude Platform (API), Claude, Claude Code, Claude Cowork, and Claude Tag, and wherever Claude is offered.

I've never asked an LLM to generate writing I wish to post as my own. I don't understand why we think it is a good idea to fudge all output just so people can't cheat on their homework or generate slop. It won't have any impact on those things because there will always be models that don't do this. All it will do is increase the rate of false negatives.

It is deeply misguided regulation and Anthropic should have just said no on grounds of common sense.


One could also use butterflies to write ;) https://xkcd.com/378/

The problem is, that LLMs worked very well for me to improve my writing. Especially as I'm not a native speaker, it was a great way to improve the legibility of my work.

I want a tool that helps me improve my writing. A tool I can learn from. Not a tool that switches out "bananas" to "airplanes."

I'm was using Claude Opus and now Fabel extensively for editing my texts and I find the recent updates abysmal. Not sure if it's due to the Text Watermark.

Before Claude was great in sharpening the meaning in my writing, it's now close to unusable.


I'll be interested to see if people can actually pick out which text is watermarked and which isn't once they introduce it. It won't switch "bananas" to "airplanes". It'll switch "I really enjoy eating bananas" to "I love eating bananas" or similar


That’s my point of the post … I had the feeling that the English text editing skills of Claude went significantly down in August.

I was frustrated at first not knowing what they are doing.

After I read this post, (seeing they introduced it in August) I think it has to do with the watermarking.

Try it on a paragraph … the connections between sentences feel clunky now.

I will play with it more and see if that’s really the case (the watermarking making the text edits worse).


> After I read this post, (seeing they introduced it in August) I think it has to do with the watermarking.

The power of confirmation bias…

> Try it on a paragraph … the connections between sentences feel clunky now.

We might have a definitive explanation at some point, but there are about a dozen possible reasons for something like this. For starters, is this something really significant and not something you notice because you are looking for it (again, confirmation bias)? A bit like some people still lose their minds when they see a dash, even though statistical analysis showed that they are not a significant marker of AI-generated text?


I’m frustrated since August with the results from Claude for English text editing … I didn’t look for that. It’s just a fact.

This is not about em dashes …


You're hallucinating


To that point, given a corpus of writing from person A and another from person B, I have no doubt it’s easy to train a classifier to determine who wrote what. In fact I believe law enforcement agencies already have these classifiers.

The only difference here is that Anthropic is actively trying to make the watermark undetectable.


They're not trying to make the watermark undetectable, that would defeat the point of a watermark. They're making it detectable, but not make the text obviously watermarked


Undetectable by a human reader. Come on, give the post a charitable reading.


This is an excellent use case. It made learning German much easier. I write what I think is correct, then get a fixed version.

When writing in English though, I use it more like a dictionary. If you want to write past a certain level, an LLM works better as a metaphor and idiom search engine.

I also like to ask it to generate 20 ways to say the same thing. It’s a great way to simplify or smoothen sentences without losing your voice.


Interesting, are you being literal about requesting 20 ways of saying the same thing? It seems pretty excessive and I'm actually impressed that asking an LLM to rewrite a statement 20 times yields output that isn't excessively redundant. Does the LLM do a pretty good job reading your mind, or do you still find yourself manually piecing together pieces from the 20 suggestions into a satisfactory sentence?


Yep, I literally ask it to say it in 20 different ways. Sometimes a sentence structure or a combination of words will just work better. The goal is usually to simplify a sentence without losing meaning.

I don't expect the LLM to read my mind. The unit of work is too small for intent to matter, and I'll just steer the next recommendations in a direction as needed.

Most of the suggestions are crap, but they can contain the seeds of a good sentence.


And this beats just writing the piece yourself?


I can’t speak for them, but I do this occasionally. Every now and then among the propositions there are one or two that did not come to my mind and that are actually quite good. I still write the whole thing myself, I just ask for advice occasionally.


How did you understand this as not writing the piece myself?


As a writer, Claude’s metaphors are trite and obvious 90% of the time. It also has a terrible penchant for an immediately recognizable emphatic voice that makes even the best outputs super cringe. But people do not notice and do not care. In our bubble and John Gruber’s bubble we really overestimate how much people must care. Truth is, we’re in this predicament because statistically speaking the people who hive a fuck are a rounding error.


They really are. This is why I prefer the volume approach. I might not accept any of the ideas it spits out, but it often guides me in a direction I was not considering.

I know that most people don't care, but my online presence is a search query for interesting people, so I care about what I put into it.


> improve the legibility of my work.

Nitpicking here in a way that I would usually avoid, but it is relevant to the conversation being had and it seems like you might appreciate the information... "Readability" would be the more correct word to use here instead of "legibility".

Legibility is close enough for me to know what you mean based on the context, but it really applies to the visual presentation and how easy something is to read at a symbolic level (whether someone's handwriting or font choice is good or bad impacts legibility, whether someone uses good grammar or not impacts readability).


> One could also use butterflies to write ;) https://xkcd.com/378/

This comparison is frankly absurd.


LLMs are no more than pen and paper at this point. Especially for those who aren't trying to create slop. We all would want our pens to accurately reflect the strokes (well in this case thoughts) rather than adding tiny watermarks to identify that it is generated by a particular pen or a user.

Watermarking per model is just the start. The method is cheap enough to distinguish individual users.


That is an insane statement, LLMs generate swaths of text from almost nothing.

If they are adding so little value as to be as transparent as a pen and paper then why use one at all? Transcription doesn't need an LLM so that's not what you're taking about I assume.


LLMs do a whole lot more than writing your thoughts down. They write extra text. If you just want a pen and paper, use Notepad. Or better, a pen and paper


>LLMs are no more than pen and paper at this point.

Then use pen and paper. It is the same, you say, right?


What? The entire reason I use an LLM is to be able to avoid thinking about a topic.

That's their whole damn value prop: outsourcing thinking and producing without understanding.

I don't need to read emails in detail to respond any more.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: