Hacker Newsnew | past | comments | ask | show | jobs | submit | hintymad's commentslogin

Finally Apple has a screen time that allows different configurations for different time segments. What took them so long? /j

> Right now, overall compute is a moat.

Very true. Or further, access to capital is the moat. It is the very reason that we don't have a real open-source community that trains frontier models - individuals simply can't afford the training infrastructure, nor sufficient high-quality training data.


https://hugovergnes.github.io/little-lm-3-8b/ << less than $1k for a 4B model

Remember how DeepSeek v.whatever cost ~$5m

Stable Diffusion 1.5 was reportedly $70k in compute.


None of those are frontier models, nor could they have been. E.g. DeepSeek's "$5 million" included piggybacking off OpenAI/Anthropic's hundreds of millions/billions by using distillation.

1) I don't think "frontier model" is a term of art, it is a marketing term. I have an M1 MacBook Pro. What do you have? Always the "frontier model" laptop? Do you always order the most expensive "frontier menu item" at the bar and grill?

2) I do not believe the "billions" number is OpenAI/Anthropic's training costs. I suspect it includes business expenses (including the big $$ to the guy who came up with "frontier model") and infrastructure, etc. That includes the data-center costs for running the cloud. And the "R&D" expenses which includes who-knows-what. And the settlement payment for data access. Etc. Why are the cost breakdowns not available to the public? Not because of thoughtfulness, altruism, care for the human race, but because of business plans.

3) "Piggybacking off OpenAI/Anthropic" - Scraping the web is "piggybacking" too, and so is buying existing data or even paying for new data. The "L" in LLM stands for "language" which is our common heritage.

But what does this have to do with anything anyway? People argue that truly Free (FOSS) LLMs couldn't be be developed because of costs, but I do not think that that is obvious. This used to be the argument against Linux and Wikipedia.


Just regarding "frontier model", it is most definitely not just a marketing term. It was borrowed from the financial concept of "efficient frontier", i.e. the best performance at a particular cost, though in the context of things like "pacing the frontier" it refers to the highest capability models. ECI (Epoch Capability Index) is probably the most commonly used metric of overall capability: https://epoch.ai/eci?view=graph&tab=release-date

"What this has to do with anything anyway" is that pushing the boundaries on AI capabilities has required tons of compute (and money). Open weight models exist and are useful but they always trail in capabilities and they can't be used as evidence to think somehow you can push the boundaries of LLMs without a shit ton of money.


> I guess what they really want is to limit sale of AI models to compliant vendors and then raise the bar to compliance just high enough so they can pass it but smaller labs can’t.

If it's true, it makes them evil. Especially Dario, who speaks about moral high ground and fate of humanity all day, yet it's not that different from a cult leader does.


> Especially Dario, who speaks about moral high ground and fate of humanity all day, yet it's not that different from a cult leader does.

He sells to Palantir. He is against regulatory change that would make Ai labs liable.

His idea of "good for humanity" or "safe" is completely different then mine or yours.


> If it's true, it makes them evil. Especially Dario, who speaks about moral high ground and fate of humanity all day, yet it's not that different from a cult leader does.

I mean think about it, it's the most efficient way to accomplish all of this.

doing what he does is exactly what you need to control the narrative and push the regulation you need and to gain maximum power.

not to mention, it also serves as a way to keep everyone inside Anthropic in check and in line with "the mission".

contrast that with the alternatives like "I'm doing this because I wanted to build a company and get rich"

doesn't work as well as "I'm so concerned for everyone"

it could be that it's really just him being him, but at the same time, this would be the optimal way to gain power and influence right now.


"If it's true, it makes them evil."

Hitler level evil.


> To him, the true AI doomer scenario is for all the immense power of frontier AI to wind up in the hands of a single powerful, proprietary provider.

Isn't this exactly what Dario wanted? He thought he knew what's best for the humanity...


It is. Dario the Book Burner will not good what he wants, the world sees through him.

Maybe the real threat is that our industry has been stalled for years, or to put it more politely, that it has been mature for a while. Back before 2010 or so, people actually read books like The Art of Computer Programming, the Dragon Book on compilers, and the Lions’ Book on Unix, or followed sites like lambda-the-ultimate.org. Then sites like High Scalability became popular. Not that everyone studied them cover to cover, but plenty of engineers considered them essential reading. Yet long before LLMs came along, that kind of depth had clearly become niche. The creator of High Scalability even put the site up for sale. Whenever someone posts an article titled "Top X Data Structures for Y," every single structure mentioned was invented decades ago. If most of what we do now is just slice and dice abstractions created and refined by previous generations, we are really just relying on our once-unique ability to transfer knowledge - something AI is rapidly replacing.

This isn't unique to software engineering. In Renaissance Italy, mathematicians like Tartaglia and Fior hoarded cubic formula shortcuts like proprietary algorithms and challenged rivals to public math duels. Today, we solve cubic equations without a second thought. Special functions used to be a staple college course for physics and engineering majors. Are they still? The US military used to employ thousands of people just to calculate PDEs by hand. Do we need anyone doing that today?

Our only hope is that our society moves fast enough to create new demands and problem domains that genuinely require new systems and algorithms. Look at AI: it’s evolving rapidly, driving massive demand, and forcing the development of new systems and architectures. As a result, the lucky few[1] working at that frontier are having rewarding careers. Without frontiers like that, the rest of us risk becoming irrelevant.

[1] One unfortunate factor is that building AI now requires lots of capital for accessing GPUs, which means individuals in the open source community have a hard time working on it.


> LLMs are able to write almost perfectly correct code.

This is kinda vague. Correct at what scale? I wonder if there's a measurement on the correctness per scale, and hopefully the scale is not just CLOC.


Curious why it is so hard to find an owner to the issue. Gates' experience started with the web UI, then I'd expect that the team who owns the microsoft.com or owns the experience. To them there are only two types of bottlenecks: their web pages (usability, JavaScript performance, ways to get backend data, and etc), and their immediate dependencies. So, they drive improvement on both types of the bottlenecks, and the owners of the immediate dependencies recursively handle their own. For instance, the web team will identify that calling the catalog API has a P99 latency of 5 seconds and the network is fine, then they ask the catalog API to improve the API latency. If the catalog team does not do the obvious, the team's manager gets punished.

Of course, I'm being naive here, as I've seen too many companies fail to achieve such basic ownership. So, curious what I have missed. Of if the ownership is not a clean DAG, well, it goes back to Gates, as he was responsible for both the org charts and the company culture.


It wouldn't surprise me that what seems like whole things from the outside are split up even further on the inside of Microsoft. There's probably no single person responsible for microsoft.com (nor was there at the time). Instead the front page is one team, Windows Update (at the time) was yet another team, and sales another, and then there are 5 other teams doing other things, and no single individual or group is responsible for the whole thing. They've all got permission to push things to the website and change whatever just from time immemorial, and not through any defined authority, since that authority doesn't actually exist.

You can still see symptoms of this today. MS will (or at least would) spin up new domains rather than just update microsoft.com, or create newthing.microsoft.com, since both of those require talking to whomever is responsible for the main domain. Even if that person exists, knowing who they are in an org like MS is already a big ask when it's organised the way it is. Meanwhile, buying a new domain likely takes less than an hour, even if you include billing and such.


All luxuries indeed. Are they enough, though? Particularly, don't many books talk about how important it is to have productive hobbies or output-oriented hobbies? The underlying thesis is that one would quickly get bored and start seeking the meaning of life if he does not output something consistently. I was wondering if being able to find such hobby and being able to afford it is also a luxury.

There's an interesting dynamic, too. Even if an engineer reads the output of the AI and understands the root cause of the problems and how to diagnose the incident, somehow it's hard for them to internalize the learning and apply it next time to a new incident. As a result, the engineer loses touch with the system anyway.

It looks like our brains somehow have to experience the failures during a diagnosis and in gemerak perform this kind of pathfinding by themselves to truly understand the system. I don't know if this has to do with how our brains actually learn.


I think using open-source AI is no longer about API cost but about company survival.

Take Anthropic for an example. Anthropic has successfully destroyed customer trust, at least for me. DHH in a recent interview mentioned that Claude refused to translate an article about immigration. Not summarize. Not editorialize. Translate! I think this reveals an unacceptable level of paternalism: Anthropic fundamentally believes that it possesses a moral authority superior to the people actually paying for the API. If such basic and mechanical translation is already too sensitive to touch, the goalposts have moved from safety into outright censorship. What prevents them from quietly deciding tomorrow that your proprietary business logic, financial data, or legal documents cross their invisible moral line?

Let alone how Anthropic treats Cursor and Figma - not that they are wrong as companies are free to compete legally, but nonetheless it shows that companies can't outsource their intelligence to a potential competitor.


I get what you're saying and it's concerning how much power these big labs have amassed and how little transparency there is in what they do with it...

But I doubt this a major factor in the trend. I just don't think it's something most corporate users run into. My understanding is these guardrails are negotiable for enterprise customers anyway.

And, not for nothing, but if I owned a human-powered translation company I would've refused to translate it too.


I like the Claude constitution overall - I hope it becomes something representatives vote on and amend, to avoid the centralized corporate censorship you describe. In the meantime, I am fine with it abstaining from doing DHH’s bidding, especially because there are so many AI alternatives.

Ah, yes, I'm sure the article that moral paragon DHH wished to translate was not at all harmful, and that this was a good-faith effort on his part /s

While I agree that Claude can be overly paternalistic at times, how should it respond to a request to translate, say, bomb-making instructions? It's reasonable to me that it might refuse this.


How do you know what was in the article?

Also, you see zero distinction between hearing opinions on political topics you might find objectionable, and building a bomb to kill people?


Because he posted about it.

It was an incredibly racist post claiming “gypsies” are like invading wolves and that something more drastic must be done to get rid of them before they kill all the “sheep” in Copenhagen.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: