Hacker Newsnew | past | comments | ask | show | jobs | submit | vodou's commentslogin

What an empty and depressing world it would be if that was true.


If you let a manager, accountant, or politician decide what 'unnecessary', 'wrong', and 'bad' are. Then I would agree. But if you let an artist decide those things, then I think it would be fine.


I think it is great that people point out LLM generated articles here on HN. Sadly, it feels like I am slowly loosing my skill to identify LLM speak. Maybe I am getting worn out of all LLM content... So, please, list the indicators and telltale signs from the specific article or blog post (like others have done here already). At least I would appreciate it a lot.


> It is not age verification. It is identity verification.

> You can change a password. You cannot change your face.

> This is not a popularity contest, and refusal is not a vote you are trying to win

These were a couple sentences that were immediate flags to me. There've been countless articles written on this (I can dig them up if you want), but IMO there are pretty clear semantic rhythms you start to notice.

It is not foo, it is bar. You can zip, you cannot zap.


If anyone uses a couple of these red flags to dismiss the entire article and the underlying idea, that says a lot more about them than the author.


I personally don't like reading text written by LLMs, I'd rather just read the prompt itself. I suspect you feel differently. I imagine this says a lot about the reader :)


At least the article makes you be curious about what the prompt was. That's already better than people dismissing the entire thing.


I agree that these are signs of AI, but they're also the way that people write. I use the "it's not X it's Y" framing a lot of the time because it's a quick way to get my point across. It's probably the sign of a bad writer because I can't come up with a different/better way to say the same thing, but I'm not AI.


> I agree that these are signs of AI, but they're also the way that people write

Certainly. If these were one-offs in an otherwise human written piece then we wouldn't even be discussing it.

In this case, they're just the simplest examples in an objectively LLM produced piece. I'm happy to point to more if you're skeptical; I'd happily wager cash the prose here wasn't written by human fingers.


I think AI was used to help write this but I doubt it was 100% AI generated. Kind of like how AI could be used to write code, but if a human's thoughts go into how the code is structured, and a human does a final pass to edit/polish the code, I'd still consider the code human written.

AI slop in my mind is anything that's completely written by AI with the only human input being "hey it would be nice if this article existed in the world"... If the human doesn't verify or edit anything the AI wrote, then I might as well just prompt the AI myself if it's a topic I'm interested in.


I could tell it was AI, but it was interesting nonetheless. FWIW, I don't think AI enablement is inherently bad; this can be a game changer for an individual with an interesting thought yet difficulty expressing themselves in written word. (Obviously, it's also created a real problem with the ability to create near infinite well-written content, especially in the case of propaganda.)

I do agree with you that the quotes cited out are literary constructs used by humans, and there's a risk we get trigger-happy in calling out AI-slop. Still, those are just the most obvious tells — there were absolutely other, less notable mannerisms that confirmed it for me. If you interact enough with an LLM, you can become quite good at detecting their output through subtle subconscious cues that are hard to put to words.

I do wonder where some of the tropes came from. Claude tends to say "____ is doing a lot of work in this sentence", yet I don't recognize that as a common construction for humans overall or even a specific community (e.g. journalists). Perhaps I'm just unfamiliar with some vernaculars found in training data. Yet sometimes, it legitimately feels like they've actually developed a lingo of their own — an emergent property.

I find it all truly fascinating (along with other feelings…), and I never expected computers to be able to "understand" language anywhere near the degree we see today. Will it soon plateau, requiring another breakthrough? Or is there plenty of juice left to squeeze?


> was AI, but it was interesting nonetheless

I suspect the prompt given to the AI was 3 sentences, maybe interesting sentences,

and everything else was AI expanded, a waste of everyone's time? (Those who read everything)


That would imply that all of the knowledge held with these models is an uninteresting waste of everyone's time. I don't believe that's true.


Yes I guess there exists such blog posts. And that it varies from person to person of course (what they already know or dont).

This blog post in particular though, I got so bored so I couldn't finish reading it. I gave it another quick try now because of your comment :-)


These are normal patterns in US English. The zeal to accuse limits the scope of possible responses. There are all sorts of things humans do when writing that LLMs mimic — em-dashes for instance — that are entirely legitimate ways to communicate but get shouted down for… reasons?

Surely you’re aware that LLMs were trained on the ways humans write specifically to mimic them? Yes? So what’s the gripe? Someone cranked out a “thought piece” with no effort or actual thinking on their own?

But thats the promise of AI.

So are you advocating doing away with AI tools and research? Maybe we start asking “should we” not “can we”? Now that is a position I might get behind.

But really how the hell am I supposed to write at all when nearly ANYTHING I write could be interpreted as AI-generated and then shouted down in some quasi-ad-hominem attack on me while not engaging with any points made?

It is the utter end of written discussion.


it's like AI-generated images, though those tend to be more obvious than AI-generated text due to the "average" effect being visually obvious. Yes, these are normal patterns in English, but it is not normal to use them with the frequency that triggers people's llm-lese senses.

I don't see a way out of it until either generated content is either completely unable to be distinguished from authored content (unlikely?), or there's a silver bullet for identifying generated content (also unlikely?)


I'm advocating that if people present a written text for others to consume, they should specify whether an LLM or a human produced it.

I think that's super reasonable, you don't?


Having thought about this for a bit now, I am honestly not sure. I think that obviously the end goal would be to not need to include such a disclaimer. But in the interim while AI is being constructed? Maybe that’s a good thing?

Somewhere in this conversation we need to make room for the following idea:

Mostly we seem to take in information, check who it’s from, and either digest or discard it. With AI acting as a “force multiplier”, so to speak, we can no longer do that simple check — we have to actually engage with the text, which is often longer than anticipated, and really attempt to understand it, whereas in the past we could lean on authority and reputation. And it could be pure bullshit! Reputation gets around that neatly. But it also makes it possible to disregard the thoughts and opinions of everyone without authority or reputation really easily.

All of society is based around that authority and reputation mattering. With AI, we all have to disregard authority & reputation and actually engage with the ideas presented. This is uncomfortable for many who have authority and reputation presently because they simply aren’t used to having to acknowledge the thoughts of people they consider themselves “better than” — or perhaps simply “lower status socially, economically, etc.”

Now if you want all points of view represented in society and for the socially and economically advantaged people to have less of a voice overall, perhaps AI furthers that goal? Perhaps the desire to label stuff as “AI or LLM generated” really is us just clinging to an old familiar way.

Honestly, I’m still not sure.


Are you saying that people are being tricked to being anti-intellectuals? If even the smart people are anti-intellectual, society loses.


No, I don’t think anyone is being tricked. I think people aren’t used to accepting information from all parties, and are looking for a quick way to determine if an article (or whatever) is valid and worthy of their time. I think this applies equally to all levels of intelligence, from the dumbest to the smartest. We all want a quick filter.

I do find it more troubling when people who are very intelligent dismiss things out of a sense that “AI did it” so it isn’t worth engaging at all.


Tons of people write that way.


Right, it's either bad writing, or an LLM that's been trained on bad writing.

I'm giving the author the benefit of the doubt and assuming the latter.


Did people used to do it 5 times every 10 sentences?


Bad writers do worse


I didn't ask about other worse things, I asked about this specific bad thing.

Since you dodged the question, I guess your real answer is "no, people never did this" and therefore "yes, this is AI slop"


in human conversation "seen worse than that", "do worse than that" usually means the thing in question is not a concern in given context. "worse" obviously includes the thing. Are you AI?


Although these are indicators, real people also use these sometimes.


My brain skims the entire blog before reading it and if I see two short sentences with dots and negation or even one single em dash, I ctrl+w out


Straight quotes were my first clue, followed by “it’s not this it’s that” and subheads.


Straight quotes? I thought an LLMs thing was always using the curly ones?


Claude uses straight, at least for me. I assume they’d use straight so that the curly ones can’t ever be wrong.



I know, right? It used to be easy: just look for writing in a very long-winded style, almost as if the author is being paid per word, in a place where that sort of writing didn't belong. I think it was because that type of writing represented a disproportionate fraction of the tokens in the training data due to the long-winded-ness. Somewhere around a year ago, they figured out some way to deal with that problem.


I think it will become pretty simple. Does the author say somewhere on their site that they don't use AI for writing? If yes, then great. Otherwise, it's very likely LLM-generated — particularly if the author spends a lot of time writing positively about AI. (And what if they lie? They will be discovered eventually, at which point they can be blacklisted.)


>>Sadly, it feels like I am slowly loosing my skill to identify LLM speak. Maybe I am getting worn out of all LLM content...

Its quite simple, at least by HN rules. If you dont like an article, post "This is LLM" and move on. Within 5 minute your post will be voted t o the top.

Easy karma farming!


The em-dashes


The multiple uses of "it's worth X" made me question the authorship, for one


> Sadly, it feels like I am slowly loosing my skill to identify LLM speak.

Just read a few fiction books written in 2026 in Amazon Kindle Unlimited. Your brain will be trained to recognize AI-Slop Speak in No Time.


I guess this just shows how divided the world is right now (in a lot of ways), but for me this sounds like one of the creepier episodes of Black Mirror or Twilight Zone.


Was this a joke? I must know!


It was true! I really do have Turbo C, Unix V7 C and Sun C compilers in my CI workflow, alongside modern GCC for C23 and C++26, Clang, MSVC, Fil-C and Tinycc and others.


It is still used for operations procedures in, at least, European space industry. E.g., in mission control systems from Terma (CCS5 and TSC).


I am pretty sure there are people here qualified enough to edit that Wikipedia page in a proper way.


An "explorative" hex editor where you can do "fuzzy" searches, e.g., searching for a header with specific values for certain fields. (I thought ImHex should be able to do this (and still think it might), but haven't really figured out a good work flow...)


Then you are part of truly strange circles, among people who don’t understand human behavior.


Almost 16000 lines in a single source code file. I find this both admirable and unsettling.


Does it really matter where the lines are? 16,000 lines is still 16,000 lines.


Even though I do find your indifference refreshing I must say: it does matter for quite a few people.


If you want recognize all the common patterns, the code can get very verbose. But it's all still just one analysis or transformation, so it would be artificial to split into multiple files. I haven't worked much in llvm, but I'd guess that the external interface to these packages is pretty reasonable and hides a large amount of the complexity that took 16kloc to implement


If you don’t rely on IDE features or completion plugins in an editor like vim, it can be easier to navigate tightly coupled complexity if it is all in one file. You can’t really scan it or jump to the right spot as easily as smaller files, but in vim searching for the exact symbol under the cursor is a single character shortcut, and that only works if the symbol is in the current buffer. This type of development works best for academic style code with a small number (usually one or two) experts that are familiar with the implementation, but in that context it’s remarkably effective. Not great for merge conflicts in frequently updated code though.


  > but in vim searching for the exact symbol under the cursor is a single character shortcut
* requires two key presses which is identical to <C-]> or g]

https://kulkarniamit.github.io/whatwhyhow/howto/use-vim-ctag...


... yes.

If it was 16K lines of modular "compositional" code, or a DSL that compiles in some provably-correct way, that would make me confident. A single file with 16K lines of -- let's be honest -- unsafe procedural spaghetti makes me much less confident.

Compiler code tends to work "surprisingly well" because it's beaten to death by millions of developers throwing random stuff at it, so bugs tend to be ironed out relatively quickly, unless you go off the beaten path... then it rapidly turns out to be a mess of spiky brambles.

The Rust development team for example found a series of LLVM optimiser bugs related to (no)aliasing, because C/C++ didn't use that attribute much, but Rust can aggressively utilise it.

I would be much more impressed by 16K lines of provably correct transformations with associated Lean proofs (or something), and/or something based on EGG: https://egraphs-good.github.io/


On the other end of the optimizer size spectrum, a surprising place to find a DSL is LuaJIT’s “FOLD” stage: https://github.com/LuaJIT/LuaJIT/blob/v2.1/src/lj_opt_fold.c (it’s just pattern matching, more or less, that the DSL compiler distills down to a perfect hash).


Part of the issue is that it suggests that the code had a spaghettified growth; it is neither sufficient nor necessary but lacking external constraints (like an entire library developed as a single c header) it suggests that code organisation is not great.


Hardware is often spaghetti anyway. There are a large number of considerations and conditions that can invalidate the ability to use certain ops, which would change the compilation strategy.

The idea of good abstractions and such falls apart the moment the target environment itself is not a good abstraction.


I was not expressing an opinion on giant files, I was just postulating on why people dislike them.


I find the real question: are all 16,000 of those lines require to implement the optimization? How much of that is dealing with LLVM’s internal representation and the varying complexity of LLVM’s other internal structure?


What would it mean to implement the optimization without dealing with LLVM's internal structure? Optimizations don't exist in a vacuum.


I do too, but I'm pretty sure I've seen worse.


Modern autogenerated C code from Simulink is rather effective. It is neither garbage nor spaghetti, it is just... peculiar.


It’s also much, much more resource intensive (both compute and memory) than what a human would right for the same requirements.


For control systems like avionics it either passes the suite of tests for certification, or it doesn't. Whether a human could write code that uses less memory is simply not important. In the event the autocode isn't performant enough to run on the box you just spec a faster chip or more memory.


I’m sorry, but I disagree. Building these real-time safety-critical systems is what I do for a living. Once the system is designed and hardware is selected, I agree that if the required tasks fit in the hardware, it’s good to go — there’s no bonus points for leaving memory empty. But the sizing of the system, and even the decomposition of the system to multiple ECUs and the level of integration, depends on how efficient the code is. And there are step functions here — even a decade ago it wasn’t possible to get safety processors with sufficient performance for eVTOL control loops (there’s no “just spec a faster chip”), so the system design needed to deal with lower-ASIL capable hardware and achieve reliability, at the cost of system complexity, at a higher level. Today doing that in a safety processors is possible for hand-written code, but still marginal for autogen code, meaning that if you want to allow for the bloat of code gen you’ll pay for it at the system level.


>And there are step functions here — even a decade ago it wasn’t possible to get safety processors with sufficient performance for eVTOL control loops (there’s no “just spec a faster chip”)

The idea that processors from the last decade were slower than those available today isn't a novel or interesting revelation.

All that means is that 10 years ago you had to rely on humans to write the code that today can be done more safely with auto generation.

50+ years of off by ones and use after frees should have disabused us of the hubristic notion that humans can write safe code. We demonstrably can't.

In any other problem domain, if our bodies can't do something we use a tool. This is why we invented axes, screwdrivers, and forklifts.

But for some reason in software there are people who, despite all evidence to the contrary, cling to the absurd notion that people can write safe code.


> All that means is that 10 years ago you had to rely on humans to write the code that today can be done more safely with auto generation.

No. It means more than that. There's a cross-product here. On one axis, you have "resources needed", higher for code gen. On another axis, you have "available hardware safety features." If the higher resources needed for code gen pushes you to fewer hardware safety features available at that performance bucket, then you're stuck with a more complex safety concept, pushing the overall system complexity up. The choice isn't "code gen, with corresponding hopefully better tool safety, and more hardware cost" vs. "hand written code, with human-written bugs that need to be mitigated by test processes, and less hardware cost." It's "code gen, better tool safety, more system complexity, much much larger test matrix for fault injection" vs "human-written code, human-written bugs, but an overall much simpler system." And while it is possible to discuss systems that are so simple that safety processors can be used either way, or systems so complex that non-safety processors must be used either way... in my experience, there are real, interesting, and relevant systems over the past decade that are right on the edge.

It's also worth saying that for high-criticality avionics built to DAL B or DAL A via DO-178, the incidence of bugs found in the wild is very, very low. That's accomplished by spending outrageous time (money) on testing, but it's achievable -- defects in real-world avionics systems overwhelming are defects in the requirement specifications, not in the implementation, hand-written or not.


HN is a very poor platform for good conversations so we'll have to agree to disagree, as I'm not willing to go further in this format


Codegen from Matlab/Simulink/whatever is good for proof of concept design. It largely helps engineers who are not very good with coding to hypothesize about different algorithmic approaches. Engineers who actually implement that algorithm in a system that will be deployed are coming from a different group with different domain expertise.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: