Hacker Newsnew | past | comments | ask | show | jobs | submit | mavamaarten's commentslogin

I'm wondering, is there a tool or something out there that helps me pick a model, in the vast sea of models out there these days? Every time I need a model for something I see the list on openrouter and I'm completely overwhelmed.

I'd love to be able to explain my use case, my cost preferences and have a tool select a few good models to try.

E.g. I wrote a tool that cleans out my email spam box. It classifies emails that are already flagged as spam, and if it's very obviously spam it removes it permanently (keeps a copy on disk though). And after x emails, it goes through the list of deleted spam mails and suggests email rules. What model would be best suited? I'd love to be able to explain this use case and get this info served to me. The list of models and the information about what they're good at is just too splintered and spread out. I landed on google/gemma-4-31b for now, because it's cheap and good enough and also supports Dutch and French a bit. But I can't realistically try them all.


My approach to this problem is to just...not try them all.

As long as the model you're using solves the problems you have to your satisfaction, there is no need to try any other models, except for financial reasons maybe.

So I start with a relatively cheap model (GLM 5.3 flash for me) and as long as it accomplishes the task (it did so far) I don't have to change. And even if it can't do something, the first thing I change is see if I can give it more tools or better context (useful even if I switch models later) or trying a different approach to the problem.

If google/gemma-4-31b works, you don't need to overthink it.


> If google/gemma-4-31b works, you don't need to overthink it.

Up until recently I had a gemini flash 2.0 api deployed that did summarization and translation of news articles/corporate statements fast and cheap and had no reason to update it.

If it works fine, this chase of the latest LLM is bit pointless.


OpenAI, Anthropic and peers deeply fear that everyone would eventually come to that same conclusion.

(I think you're right)


Yeah, I really think we are on the cusp on the difference between SOTA and cheap models is small enough that paying 10x or 100x the cost makes no sense.

It also helps that a lot of effort has already been spent figuring out how to do more with weaker models because SOTA 1 year ago was behind what the cheap models do today.

1 year from now, unless the SOTA companies come up with something truly revolutionary they will be in a lot of trouble.


I find Qwen 3.8 better at coding but I also use Gemma 4 31B. I toggle between them in Ollama.

Which is why they are just making sure we can't buy any GPU to use any models. And because GPUs are so costly, I would be just worried to ruin it running a model continuously.

Yeah, might be far fetched, but it seems like heading that way


Starting with GLM-5.3 Flash was a pretty decent first try! I started with other models, and ended up settling on this exact one because all the others were either too slow or unreliable for my tasks. Qwen 3.8 didn't do it for me, whatever tweaks I added to my harness. Where I'm getting at is you did start with an incredible model in the first place, which greatly helps sticking to it.

GLM-5.3 Flash has been my goto since it came out. Only failed once when it lost context, but I'm assuming that was my fault rather than the model. If models never make it past today's close-to-frontier for the rest of my life, I wouldn't complain.

I use Gemini(s) because I can send pdfs as files to their API and not worry too much. I've started to diverge and consacrate a part of my pipeline to sending image based pdf pages to glm flash 5.3, not sure how to address / test it properly.

Long term I have fears I can't depend of the Google's AI api.


satisficing instead of optimising

This 100%

I use models.dev's CLI tool, which I think gets data from OpenRouter, and ArtificialAnalysis so your coding agent can help you narrow it down.

<sidenote>

Similarly, HuggingFace has a CLI + a few skills, and they are very useful.

I had a production image processing using Gemini 2.5 Flash Lite (which is getting discontinued in October), and in 20 minutes Claude Code + HF Cli recommended the best replacement small model (Qwen VL 3B something) and proceeded to fine tune it on my datataset. All this while I was in a rush to get dressed and go to the store.

It cost ~$3 I think, and results were excellent. Not perfect, but not far from perfect either.

We didn't replace Gemini in prod at the time, because we didn't have time to do all the math on how to end up with a smaller bill/mo.

</sidenote>


Have you tried openrouters auto model? Tries to give you the best model based on prompt and price

https://openrouter.ai/docs/cookbook/coding-agents/openclaw-i...


That’s if you have disparate prompts and don’t want to actively select a model. If you’re developing a pipeline, it’s a terrible idea. You want to choose a model, validate it, then stick to it.

I have, but honestly that was exactly what I am not looking for. Sometimes it picked a model for a Dutch email that totally does not support Dutch. Other times it would work fine. It's just a layer of indeterminism I wasn't looking for.

If you use the Chinese ones at least the energy comes from solar - aside from that, pick one and see if it solves your problems

I doubted this but it does look like there's an actual government initiative that mandates 80% clean energy for all new data-center builds: https://www.fastcompany.com/91578780/how-china-is-powering-n...

that said, China is also rapidly scaling up coal-fired plants: https://apnews.com/article/china-coal-power-plant-carbon-cli...

those presumably support all of the surrounding infrastructure + people + manufacturing so it's not as if it's truly solar-powered. but it's still handily better than the state-by-state abandonment of clean energy goals here in the US - I lay this out a bit here: https://news.ycombinator.com/item?id=49700743


Note that the 80% are qualified with "inside eight large data center hubs created in China’s “East West Computing Strategy”." Most new data centers are built elsewhere and will use whatever source of energy is most convenient.

>80% are qualified with "inside eight large data center hubs created in China’s “East West Computing Strategy”." Most new data centers are built elsewhere

this is a pedantic misinterpretation of the scope of East Data, West Computing. the hubs are all in tier one cities with the power generation capacity to handle them. each hub is comprised of dozens of data-centers and are the place to build because of large economic incentives, infra guarantees, and university-trained talent living in close proximity

see https://www.sciencedirect.com/science/article/pii/S209580992...

>This initiative is expected to accommodate up to 95% of China’s digital data needs through this reorganization

>Within a year of implementation, more than 112 new data centers have either been authorized, under construction, or completed across the planned computing hubs

this project itself is the thing that has all the other nation states in the world clutching pearls about not supporting their native AI corporations enough. it's something you and everyone else should really familiarize yourself with because it is one of the major items influencing current realpolitik and economic decisions (and is likely going to be the thing that'll lead us to a thrice-in-a-lifetime sized recession but that's a convo for a different day)


If you use a model hosted in China while it's night there, the energy obviously doesn't come from solar, as there's not enough storage capacity. Even if you use it during daytime, most of the energy still won't come from solar, because the ideal solar power locations in the sparsely-populated west are far from the ideal data center locations in the densely-populated east and there's not enough transmission capacity between them.

Additionally, using a Chinese provider doesn't mean the model will be hosted in China, e.g. for Qwen Omni here, the supported regions are: China (Beijing), Singapore, China (Hong Kong), Japan (Tokyo), Germany (Frankfurt), and US (Virginia). https://www.alibabacloud.com/help/en/model-studio/qwen-omni#...


power capacity coming from solar means that at peak, no new non-solar power plants are needed. it does not mean that at nadir everything is solar. legacy power plants continue working around the clock, it either goes to waste or it goes to things that run at night.

that aside, solar also includes storage of energy that comes from solar. you should read up on energy networks.


> at least the energy comes from solar

If by that you misspelled coal, sure

https://ourworldindata.org/grapher/share-elec-by-source?coun...


Coal in China: dropping year in year out for past two decades as percentage of total energy consumption.

It only started dropping last year and by dropping you mean total % of energy generation because total coal energy production is still going up.

I appreciate the input, maybe tighten up your commentary.

Check the source we're discussing: https://ourworldindata.org/grapher/share-elec-by-source?coun...

As you can see it is percentage of coal as part of total energy mix that has been dropping for two decades.

As already stated above. Perhaps you missed that?

You're correct that absolute tonnage of coal use is still climbing, - but at an ever decreasing rate and was predicted to peak .. of course, thanks to some shenanigans by the Hegseth's of the world that may be deferred for a while, thanks USofA!

Of course on the plus side (for the atmosphere) there's been a drop in global fossil fuel consumption, so swings, arrows.

China, of course, is still well short of total CO2 tonnage lofted over the past century in comparison with the US - even allowing for a vastly greater population.


And yet, the single vastly dominant source of electricity at around 55%. It's a long way you go and a long shot to call it clean.

And yet dropping - by a series of deliberate 5 year plans with the goal of building out sufficient renewable energy to phase out coal use altogether.

Kind of like what absolutely isn't happening with the current US administration.


if you mine 4 tons of coal, anywhere on earth, 1 ton of it will be going to china to be burned to make electricity. China's energy production is extraordinarily dirty, the cleanest part of it is the PR. More than half of the energy they produce is from burning coal. They certainly want to integrate more renewables but they have the same problems with that everyone else does - storage and transmission are expensive and essential for a renewable heavy grid.

It's not like new releases come with fully mapped out capability scores for exactly the aspects that you're interested in. There are benchmarks, but reality is often different. It's simply unknown to humanity how well each model will perform in your own bespoke context unless you just try them. You can read experiences and vibes by others but often they will use them in different ways or have different preferences etc.

They generally all try to make them good at everything, it's not like they'd declare "this model is not made for task X".


I bumped into openrouter's `ori eval`, didn't try it though https://openrouter.ai/ori/eval

I use claude for work, and opencode go at home, but i do dabble with openrouter from time to time, and then i just browse the model catalog and filter/sort by recent popularity, price, context size or whatever matters for the task.

Mostly just popularity trends, hoping that there's some wisdom in the crowd.



Pick the cheapest model with good speed. ZDR, and price. If it works great. If it doesn’t pick one a bit more expensive till you get what you need.

This is the approach recommended for starting a hobby as well. Buy the cheapest set of tools and if they break or are not meeting your needs.

For TTS I launched this like this week based on a Reddit thread of recommendations, added a new one to it.. I want to say on Wednesday, but this week has been a blur. Problem I’ve found with similar sites is I can’t run a lot of the models, or the results are beyond stale.

But this is only stuff I can run locally, or it’s a cloud model.

So.. This is good for right now!

https://apimade.com/audio-compare.html


The approach I normally take is: I have small benchmarks for myself which test for things I care about.

And that has any models that I'm considering through that.


>I'd love to be able to explain my use case, my cost preferences and have a tool select a few good models to try.

Ask one of the top-tier models to do deep research on it.


For cost you can route the same task via openrouter and see what it costs. Literally a for m on models do task loop and look at the cost.

It's the same problem as trying to buy a car or choose what clothes to buy. You just have to read about options, try things out.

Why not just ask Google AI or ChatGPT to help with choosing?

No. If you're making maps and building a database of facts, I think the possibility of hallucinations is actively harmful.

Another thing they apparently do is let business owners (e.g. hotels) remove negative reviews because they're "defamatory".

I had a very factual 3-star review on a hotel that was supposedly very quiet, but was super noisy and the AC was not working. And there was literally nobody at the reception at the time of check-out and nobody picked up the phone either. All in all, nothing serious, but worth a review. Purely factual stuff, nothing bad-mouthed, just my experience and three stars.

A few months later the review was removed because it was "defamatory". I was able to get it reinstated by sending them proof of payment to the hotel and proof that I was there. A few months later, it was removed again! That same dance was repeated six times..... What a weird system.


I'm surprised you knew it was removed. All my negative reviews appear visible to me when logged in and disappear when logged out.

Only sure fire way I've found is to leave a two star review with no body. One star without a body is also shadowbanned for me.


Yeah they sent an official looking mail saying that someone took action against my review and that it was therefore removed.

MBAs, consultants and accountants have taken control of the company. All are in it to move up and cash out. They will optimize for profit and not much else. Loosing (or even defending) a defamation case is expensive. Spiking your review is free.

This is not the old Google run by engineers. Stop expecting that and you will experience less frustration.


This depends on local defamation laws. It's an issue with all online reviews in countries like Germany where the onus is on Google to establish a review isn't defamatory.

Yet another example of how the incentives work when the platform makes its money with ad revenue. The hotel pays Google for ads and gets favorable search results. If Google then hosts a negative review, the hotel might be less inclined to advertise there. So, only positive reviews get published.

The reviewer isn't contributing to Google's revenue at all, so his honest review is only worth something if it makes the advertisers happy.


> The reviewer isn't contributing to Google's revenue at all,

In theory, good reviews drive traffic to Google because people come to trust it as a dependable source of important information. In reality, Google is big enough that they don't have to care. They figure that people will keep using Google no matter how many times Google lies to you. Seems like they're right because their AI overviews lie constantly. https://arstechnica.com/google/2026/04/analysis-finds-google...


Not just Google but many others. A big travel tech company in Asia removed my review of a shack that I once stayed in, since it mentioned things that the business felt were negative.

After all, there are many customers like you and me. If we’re gone, someone else can fill the gap. Not saying I condone it. I dislike it. But I understand.


And on top of that, we've all seen how confidently LLM's get things wrong. I've seen them get better, but even the current SOTA models fumble. Why would you want to send a bot a message "sure, reschedule my flight" and have a bot potentially fumble that in all sorts of ways? I really really do not see the appeal. My calendar and my emails and my things to do are all things I want to be on top of. Not have it half-assed or potentially messed up by a chatbot made to harvest all your data.

Or worse, take money to steer you to a specific airline.

So I can turn it on using Home Assistant reliably.


Oh and don't forget: it keeps chaining a bazillion commands together so any whitelisted commands still need approval because they're nested in such a convoluted way.


I have a hook that auto-denies when it sees 'python -c "', ' awk ', ' sed ', and '&&'


It can still invoke them in subagents and does this.

And find --exec too.


Nothing another .md file can't fix


Now you guys are just adding no-op fuel to the fire.


I made Claude write noöp instead of no-op and it's still amusing a couple of days later :P


> I made Claude write noöp

Tangentially related [1]:

> The billboard ad, located next to the Ikea Tempe store in Sydney, says 'NÖFNIDEA? No tools, no worries’.

[1] https://www.adnews.com.au/news/koala-mattresses-takes-swipe-...


LinkedIn too. The first time they show a bottom sheet dialog thing on the website, and if you dismiss it, it scrolls you all the way to the top. The second time it appears, it messes with the page and sometimes automatically closes the tab.


It's a real face palm how LinkedIn squandered the opportunity to be a relatively benign social media product. School kids would have zero interest in it except for that guy who wears a bowtie to high school. And yet they have made it such a swamp full of dark patterns. What's the marginal gain on that? They're the only game in town for the kinds of people who advertise on LinkedIn.


I think the only answer is that all dating sites are full of dark-patterns, and promise to promote your profile if you pay more money.

Linkedin isn't a dating site? Sure it is. I've seen the messages friends receive, and the way people try to sell themselves in the fakest possible ways - it's the only way to judge it these days.


I love it how a normal social network makes a "ding" sound when somebody likes your post but LinkedIn gives you a "ding" for just posting anything at all! Overall it is a place that sends the message that "standards are low here"; don't know how to write? Microsoft Copilot will do it for you, etc.


And on Hacker News, if you post something really controversial, you get a “dang”!


Perhaps these dark patterns were necessary to ensure that they stay the only game in town?


I used to have some trouble with linkedin, but currently it works for me reasonably well in Firefox on Android with ublock origin.


Yeah I've seen it a lot. It goes through the effort, unasked, of pulling screenshots off a connected device and then it's like... Oh shit yeah I can't see.


It's doing it's best to accomplish whatever task you've thrown at it.

It's expecting you to have done at least something besides select DS4 on Ollama, essentially.


Even with the price hike, Deepseek V4 Flash still does this a lot better than any similarly priced model, in my experience. I've had Luna take shortcuts (like adding an overload to methods whose signature it changed so they don't break existing tests, instead of fixing the tests) or just not do the entire work and report it as done (did not fully resolve rebase conflicts). Deepseek has never really failed in this type of way for me, and it has been far more persistent in validating its work than Luna (and several bigger models).


That's cool. I often write tiny blurbs of kotlin just to test out a simple algorithm. I often do this on kotlin playground because doing so inside a scratch file or test is somehow more cumbersome and slow. This ran and compiled something in 98ms on my smartphone, cool stuff.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: