Hacker Newsnew | past | comments | ask | show | jobs | submit | j_maffe's commentslogin

I'm not sure I agree. More is different. Being able to execute logic at a higher scale and speed would make some previously infeasible intelligence tasks possible, resulting in a new level of intelligence.

That would be the best case scenario. I honestly wish things would finally slow down a bit. I don't see it happening.

Would you argue the same for allowing everyone to have guns?

(1) Are guns open source? (2) Far more people - order of magnitude higher - have access to guns than they do to open source software - I count access as in actually being able to do something with it.

i would, i support the 2nd amendment.

the same logic applies well to the freedom to own and use ai.

empowering the people prevents tyranny.

such freedoms make society less safe. it is better to face the harms caused by a free public than to vest power in the elite, those like dario amodei.

the elite belief that they have a right to paternal rule makes them blind to their fallibility. no individual should have absolute power.


Yeah the Chinese fear-mongering falls a bit flat when coming from a point of maintaining US supremacy

> in a fraction of the time

Well if you do the math, the number of agent-compute time in total, given the insane number of agents thrown at the problem, might end up being comparable in time, if not for the budget.


I think if you use an LLM just to proofread then it'll not be able to insert a strong enough watermark.


The tasks are the thing to really look at here:

https://github.com/harbor-framework/terminal-bench-science/t...


I am hoping someone with more free time than myself can contribute some things in the RF engineering domain in the 'engineering-sciences' section. There's some problems out there that will definitely stump even a smart LLM.


Looks like most things definitely stump even a "smart LLM"... Best score on this is 30%. Which is what you should assume for tasks you give an LLM if they aren't exactly the same as an existing benchmarked task. They're just not that good for the purposes people seem to think they are. Very limited application space.


> I worry this doesn’t check correctness

Then it's not a valid benchmark. I agree though they're not reliable enough to just put results in a paper.


The Purpose of a System is What it Does.


The Word Purpose was Invented Precisely to Distinguish Between What a System Does and What it Ought to Do.

Less catchy, but damn, I hate that slogan.


No but a provider with more amicable terms can.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: