Hacker Newsnew | past | comments | ask | show | jobs | submit | claudeIsDown's commentslogin

I couldn't agree more. This articles describes how frustrated has been OpenRouter experience. At the end of the day, I ended up configuring to use the owner provider of each model I needed to use.

On OpenRouter the pricing is: Input $0,075/M - Output $0,25/M - Cache Read $0,015 /M

How is the business model of Anthropic/OpenAI will sustain?


They're obviously in a pickle, nobody is going to continue to pay $15-50 a mm tokens here soon. There's a reason OpenAI stopped training large models last week, and it's not because of "saftey" or "alignment" they know these gigantic models are not worth the squeeze.


I think anthropic is behind but Luna on a Jalapeno seems profitable


This is a bad model. Worse than Luna in every way; slower, dumber.


OAI/Anthropic shareholder? Speed and intelligence are not "every way". Cost is essential. Hence the Pareto boundary illustrated in TFA.


It literally cannot complete tasks that Luna can do easily. It doesn't matter how cheap it is.


It actually can. I have been using it regularly for past 4 days. It is on sol-low level. I deliberately tested it on a moderately complex task. glm-5.3-flash one-shotted it correctly. Luna max couldn't achieve parity even after 3 total attempts.

It was creating java bindings for this project: https://github.com/jeffhajewski/latticedb

And here is the binding one-shotted by glm-5.3-flash: https://github.com/jeffhajewski/latticedb/pull/5


Sounds like discrete propaganda


Again, again and again.


I would love to see a more descriptive review from simonw instead of just SVGs generations.



He is not an ML researcher or engineer, he is a passionate AI enthusiast blogger. He mostly does SVGs and other low effort checks (sometimes with major flaws, as people have pointed out a few times in the HN comments). Properly evaluating the model across all fronts requires a deep understanding of LLMs, how they work, the trade offs behind new architectures and the relevant research papers. It also takes a lot of time to build a proper evaluation framework so basically you can't just vibe code that if you want something that is solid.


He created Django, what do you mean he's not an engineer? Also 'low-effort??' his posts are extremely in-depth, clearly very thought through with a significant amount of time and energy. Additionally he does perform multifaceted checks across LLMs in many of his other blog posts.


> ML researcher or engineer

The charitable reading is that they meant “ML researcher or ML engineer” with the latter meaning, I guess, an engineer who works on developing LLMs not just using them.


Yes, thank you.


> He created Django, what do you mean he's not an engineer?

I specifically said that he is not an ML engineer (emphasis on ML), so I'm not sure what Python web frameworks have to do with anything.

> Also 'low-effort??' his posts are extremely in-depth, clearly very thought through with a significant amount of time and energy

And yes, low effort. Pelican was low effort, his Fable test was low effort, his HN filter etc. Read the discussion in the comments under the Fable test, it's not just my opinion. There was also another example a few months ago. You can search for it, I don't keep track of these things.

I discussed this with him directly after he called himself an "ML expert" in comments.

This is a classic case of the Gell Mann amnesia effect. I read ML papers and work with ML, but to people outside the industry, his writing can look "extremely in-depth" even though it really isn't. People I work with have the same opinion.

> clearly very thought through with a significant amount of time and energy. Additionally he does perform multifaceted checks across LLMs in many of his other blog posts.

I have never seen an article by him about any model that I would describe that way.

And the most revealing sign that he is not an expert is the type of questions he asks and the mistakes he sometimes makes in the comments here. They show why he is not capable of doing any technically in depth evaluation (at least with his current knowledge level).

If you actually want to learn something as a layperson, read articles written by ML PhDs like Sebastian Raschka or watch Stephen from Welch Labs etc. that are directed at general audience.


We at HN: https://xkcd.com/2501/ to basically say that I think you might be considering low-effort what’s actually an attempt at simplifying - which is arguably higher effort


> you might be considering low-effort what’s actually an attempt at simplifying - which is arguably higher effort

I'm not saying that simplifying complex topics is low-effort, good simplification can obviously require a lot of work and I fully agree here.

What I meant is more that some of these tests feel methodologically sloppy, they are too shallow, miss important technical context, do not control for enough variables etc, yet the conclusions are sometimes presented lets just say... too strongly, as I don't want to be too harsh.


Oh, i see. That’s entirely correct. I think the pelican test is more of a meme at this point, similar to Ethan’s Otter on an airplane for video models


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: