GLM 5.3 looks strange, because of this Chinese labs benchmaxx moto. So rather they have emergent abilities or...
Also a lot of questions to benchmark because opus 5 is completely useless model right now.
I think that the main problem with opus that they try to solve context size optimization problem, and that is the main reason why it speaks like alien with only one technical dictionary at hand. So why it is so good?
Opus 5 generates really good code and terminal commands though. It's just bad at the accompanying text it tells you. These benchmarks don't grade the text generation of the response I don't think, only the task outcome.
Terminal Bench 4.0 did not introduce new questions. All tasks were public for a while. If you look at GLM 5.2, which is using the same base model as 5.3, but was released prior to most tasks, it does extremely horribly on terminal Bench 3.0 (4 to 8 times worse than every other model) - source: https://benchlm.ai/benchmarks/terminal-bench-3
It is rivaling Astra, on their own benchmark that they made (FrontierCode), that they ran themselves in their own closed-source ecosystem that isn’t reproducible by anyone.
The last time I tried Minecraft, from the point I dropped into the world, I could start punching dirt and trees and placing stuff around and building things.
In No Man's Sky I got dropped in some world that felt empty with tutorial on how to collect some stuff and how to get to the ship.
The first impressions weren't the same between these 2 games.
I made this point earlier in the comments, but I think the main difference between Minecraft and No Man's Sky is that Minecraft does not push the players towards specific gameplay systems, while No Man's Sky does, making No Man's Sky feel much less like a box of legos and more like a collection of specific tasks/activities.
I've seen cards with 1% chargeback in the EU though. How does that work?
And it wasn't just a temporary marketing promotion. I've used such a card for many years.
(It was issues by a big bank that had almost no presence in my country... so maybe they were eating the cost just to build up a bigger presence and potentially enter the country?)
I don't know it has as much to do with scale of the system vs the general architecture. E.g. the system I primarily work with these days has millions of lines but most PRs are for a small changes which are well contained in scope by the overall architecture.
It's a cultural thing, but you can do incremental PRs towards a large goal. Giant PRs that are expected to be reviewed never really seemed worth it imho.
I don't think a PR needs to be tens of thousands of lines for a review to take more than 20 minutes. I've worked on projects where you can't reasonably get a branch set up and running somewhere to test it within 20 minutes if you need to set up dependencies, peripherals, external systems, etc.
reasonably you should already know what the code and project is supposed to do before they even start writing code, so you can get directional feedback in.
then you are maybe reviewing 1 out of 7 PRs that implement the agreed upon change
I'll be honest I read it a long time ago, and not even in its entirety, so I don't remember exactly all the different ways it suggested AI could take over, but I remember being convinced it was definitively a possibility, and that there were many ways in which this could happen.
What I remember clearly is how surprised I was to see these(what felt at the time) unlikely futures being expressed in such clear detail and with so many references to existing literature, just to give an idea there are 56 pages of small-font "footnotes".
Also just to clarify, the book talked about different types of AI, was not focused just on LLMs.
>I’m sort of surprised that the EU isn’t stepping in to support them.
On a similar note: Why does it have to be the EU to step up?
Why doesn't EU rather speed up making VC investments more attractive, so that EU and banks don't do the majority of investing?
*I don't have answers to these questions. It just frustrates me how many investments here come from politicians and banks, rather than from investors, people, and companies.
It’s a different culture with different rules, if it was that easy it would have been done already. See my other comment about lack of budget at a federal level.
A general AI classifier that can be set up easily and used to classify anything… but with probably lower quality than a purpose built one.
reply