Hacker Newsnew | past | comments | ask | show | jobs | submit | book_mike's commentslogin

What I care about is whether the model is capable of the tasks I give it at the lowest cost. Right now I'm using Kimi-K3/GLM-5.2/Minimax. Sonnet is great but I burn through the tokens too fast. Opus 5 set to max is amazing and more intelligent than all of us. .998 of the time I don't need that kind of intelligence. I just need the job done.


How do you define intelligence? I encounter that kind of sentiment all too often, and I have to assume we go by wildly different understanding of what that might entail.


My definition is that I can be much less precise with AI the more intelligent it is. It can extract the intent from my fuzzy description of the problem. Which means I can offload some of the thinking effort.

It wasn't possible a couple years ago. I used to make fun of people who were trying to get ChatGPT to think about the problem when all it could do was write code from the pseudocode you provide.

But now I can say: "Look at the latest log and make a plan to fix". And it takes it from there.


I think this is the best and most useful way to measure model intelligence. In my experience it's what really sets apart the capable models from the best. A small model can be RL trained to be extremely good at programming or narrow problem solving for its size (eg 5.6 Luna, DS4 Flash, Qwen 3.6 27B), but even Luna is IME comparatively awful at understanding intent and making good decisions with limited guidance.


I'm not sure if you are aware as to the extent certain processes and functions are being anthropomorphized.

These systems are not "intelligent" if you follow the dictionary definition. Hence the question posed to get a better understanding of how it is being used in this context.

They also do not "extract intent". There is for sure some intent behind your input to the service. What follows is a predictive text that uses your input, together with a LLM trained on a corpus with similar relations, that ultimately gives you a series of words.

That isn't to say a service like this cannot be useful. But I'm often wondering if the people who rely on these, and are particularly enthused by them, are actually aware that the terms they used are in fact anthropomorphized. I start by giving the benefit of the doubt, but it rarely lasts. 'Reasoning', 'agent', 'skill' 'hallucinate', 'know', 'think', 'train', 'learn', 'understand', 'harness', 'attention', 'context', 'prompt'.


Sounds like a good working definition in the context you are using it in.

> But now I can say: "Look at the latest log and make a plan to fix". And it takes it from there.

I usually tell the agents to first work on reliably reproducing the problem in the log, and only then even start thinking about a fix.


If you read Opus 5's output, it is beyond the comprehension of virtually all engineers and developers. That is what I mean by intelligence. Math, science, and engineering are all contained in one model. We may be experts in one field. The model is an expert in everything that humans know.


I'd have to ask for you to be more specific, otherwise, to take your answer at face value, it comes across as a contradiction.

> [Opus 5's output] is beyond the comprehension of virtually all engineers and developers

That would make it pretty bad? The key defining quality of good software, is clarity, and the ability to simplify a complex problem to the point of it seeming trivial.

> Math, science, and engineering are all contained in one model. We may be experts in one field. The model is an expert in everything that humans know.

The bar here should absolutely be to judge this against the expert level within each domain. I have time and time come across LLM output being woefully underwhelming in every single request where I am an expert. For all areas that I am not, it sure seems plausible. It is far more likely than not, that it is equally inadequate in the areas I lack the necessary knowledge to tell.

If the AI is being subpar in every field and category compared to an expert in said respective field, then, what a strange gauge of a tool's usefulness. Are we attributing higher value because a single model is "attempting to solve all knowledge and fields at the same time", why is that of any importance, or excuse?

We should not define "intelligence" as how effectively it can convince a non-expert of something being plausible. That sounds like the absolute worst tradeoff. You'd have to waste the experts time in filtering and refuting incorrect postulations that are cheep to generate. The perfect storm for bullshit asymmetry.


I disagree with points 1, 2, and 3. Point 4, AI is better than average, and sometimes it's better than excellent. Point 5 is irrelevant.


That seems awfully self deprecating. Surely, you expect better of yourself in at least some area, than the average competency of humans across all areas?


Careful, you may have a bit of psychosis. They are very, very far from incomprehensible, and also very far from the top at least of my field. The best in my field are produce far higher quality results, and I think that's true for all fields. It's just an incredibly good 85% quality machine that experts all use because they can guide it to be up to their quality faster than doing it themselves.


You could take that even further, to the actual danger of reliance of these tools when you lack the expert knowledge. That is, when you assume it took you 100%, but missed the 15% it got very wrong, or perhaps even worse: subtly wrong. This compounds with the next similar task, and either you've made the actual experts quit their job as it has become to babysit LLM output, or you end up with an unusable mess, deleted production databases, etc.


I don't think that's because of its "intelligence". It speaks obtuse techbro-ese: stringing together words that sound smart to obscure the simplicity of the thing it's describing. In many ways it's the opposite of intelligence.

Opus 5 and Fable 5 in particular suffer from this issue at worse level than most models in the same class.


This is a "load-bearing" issue recently.

I think the idea is packing more information into fewer words, but the result is a word salad that is somehow simultaneously very dense in adjectives and adverbs, and still way too verbose.


  > stringing together words that sound smart to obscure the simplicity of the thing it's describing
Laser targeted at LessWrong posters


yep, it is so bad i had to create rules to cut down on the techbro language and domain slang.


I just canceled my Claude subscription outright. The models are all gairly fungible, it's easy enough to just switch to another provider.


You mean Fable 5 right? Opus 5 makes lots of stupid mistakes about anything that requires any domain knowledge.


If you read the many, many complaints about opus 5 on anthropic forums, the sentiment is that opus 5 output is poor and people are back to 4.8 and 4.6.

You may want to re-evaluate and compare to the older models.


Opus 5 is too verbose.

I'm using Sonnet 5 on a large porting project and it's good. I switched from Opus 5 to Sonnet 5 on a project of another customer and I didn't notice a decrease in quality. I concede that it's very difficult to assess a difference in quality unless one uses both models on the same task and carefully compare the code, not the output in the terminal. I really don't have the time and the tokens for that. Anyway, Sonnet is still doing a good job.


so?????


Opus 5 fucking sucks to talk to and read compared to 5.6 Sol though. I’m fully done with Claude models until they figure this out


My god the Jargon is so hard to parse. Everytime i resort to cursing it , it understand. Even putting `ASD-STE100` or simplified english in the claude.md doesn't work. It gets the job done but is an anti social asshole


Check out https://unbiased.ai

(Disclaimer: I’m a co-founder)


how are you paying for tokens with this setup + what harnesss + how many tokens/day are you consuming?


Technology and information migrates from one form to another. Let me get my papyrus.


We will see if this project has legs. This is the kind of efficiency we desperately need. Now if we can address efficiency with llm training.


Good. Perhaps the administration should follow the law.


Sematic ablation... that's some technobable.


Going off search results, it seems to be a new coinage. I found mostly references to TFA, along with an (ironically obviously AI-written) guide with suggestions for getting LLMs to avoid the issue (just generic "traditional" advice for tuning their output, really). The guide was apparently published today, and I imagine that it's a deliberate response to TFA. But FWIW the term "semantic ablation" does seem to me like something that newer models could invent

At any rate, it seems to me like a reasonable label for what's described:

> Semantic ablation is the algorithmic erosion of high-entropy information. Technically, it is not a "bug" but a structural byproduct of greedy decoding and RLHF (reinforcement learning from human feedback).

> ...

> When an author uses AI for "polishing" a draft, they are not seeing improvement; they are witnessing semantic ablation.

The metaphor is very apt. Literal polishing is removal of outer layers. Compared to the near-synonym "erosion", "ablation" connotes a deliberate act (ordinarily I would say "conscious", but we are talking about LLMs here). Often, that which is removed is the nuance of near-synonyms (there is no pause to consider whether the author intended that nuance). I don't know if the "character" imparted by broader grammatical or structural choices can be called "semantic", but that also seems like a big part of what goes missing in the "LLM house style".

Bluntly: getting AI to "improve" writing, as a fully generic instruction, is naturally going to pull that writing towards how the AI writes by default. Because of course the AI's model of "writing quality" considers that style to be "the best"; that's why it uses it. (Even "consider" feels like anthropomorphizing too much; I feel like I'm hitting the limits of English expressiveness here.)


My word, it doesn't have to be that way.


BBC, nice PDF. Fossils.


LLMs are useful if you use them properly and they are getting better everyday. Arguing against LLMs is like arguing against a shovel. Just use it right.


A lot of arguing "against LLMs" is not arguing "shovels aren't useful," it's arguing "maybe shovels aren't actually going to replace all human labor, and sinking so much capital into it we're starting to conceptualize it in terms of 'percent of global GDP' might not be such a great idea."


That's the theory, if the vast majority of people use it wrong the problem is the tool, not the user.


I haven't noticed them getting any better in the last year.


You absolutely have not been paying attention then. The difference in quality between September 2025 LLMs (GPT-5, Claude 4/4.5) and September 2024 (we were still on GPT-4o) is huge.

For one thing, last year's LLMs were nowhere near winning gold on collegiate math and programming competitions. That's because the "reasoning" thing hadn't kicked off yet - the first model to demonstrate that trick was o1 in ... OK that was September 12th 2024 so it just makes it to a year old now.


LLMs are powerful tools but they are not going to save the world. I have seen this before. The experienced crowd gets chuffed because it is a new pattern that radically changes their current workflow. The new crowd haven't optimised yet so they over use the new way of doing things until they moderate it. The only difference I can detect is that rate of change increased to an almost uncomprehensable pace.

The wave’s still breaking, so I’m going to ride it out until it smooths into calm water. Maybe it never will. I don't know.


> The only difference I can detect is that rate of change increased to an almost uncomprehensable pace

This is a pretty seriously bad difference imo


I fully support developing newer cheaper alternatives to housing and cut open the jugular vein of this rotten market. Let the investors and creditors burn.


What are these alternatives you speak of? Genuinely curious.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: