This is the takeaway here: That's how they have been serving it at scale as Ox-Alpha. This is a definitional moment.-
Further quote:
"Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale."
They are high end really expensive Huawei ascend GPUs.
It is kinda bruteforcing the performance on a older semiconductor processing tech, so total production is pretty low.
By this they probably mean RTX series GPUs? If so, then they are not comparing the hardware efficiency with the A100 / H100, etc. that are commonly used for training models
Another self-inflicted own courtesy of US government policy.
While I think China would always get to hardware self-sufficiency eventually, all export controls have done is (1) accelerate China's development, and (2) divert revenue that would've otherwise gone to NVIDIA/AMD/etc instead.
Long term it's irrelevant. The only relevant thing is that there's lots of money in chips that can do high performance inference. You see all kinds of competitor products in development or already on the market even here in the US where there are no such restrictions. Cerebras comes to mind. It's natural and expected that eventually Nvidia will either have to keep way ahead or competition will catch up with specialized products.
That doesn't mean by any stretch of the imagination Nvidia will disappear. But the entire stock market valuation, not just tech, has had me scratching my head for a while.
Cerebras "competes" with Nvidia in the same way a Vespa scooter competes with a Ford F-150. Groq and Tenstorrent are in a similar boat, ASICs don't really threaten CUDA.
Curiously, there is not a single real CUDA competitor anywhere in the world. We almost had one with OpenCL, but all of the American stakeholders abandoned it right before the crypto/AI takeoff. All of which means that Nvidia sets their own margins, exploiting American investors and taxpayers while letting China avoid their dominance. So the American economy subsumes the bulk of Nvidia's arbitrarily-priced debt, and the Chinese economy can direct SOEs to pour billions in liquid cash into real GPGPU research.
I'm an American and I'm pretty fond of Nvidia, but Jensen was right about this policy; it gives China everything they need to actually replace CUDA. It's reminiscent of America's attempts to deprive China of ARM and Texas Instruments IP, only to end up swimming in unlicensed clones after refusing to sign an IP deal.
Revoked or not, just ever having those controls signals to the Chinese ecosystem that you're not necessarily a reliable supplier (Would you trust US export policy to remain stable for the next ~decade given the state of US politic?) and to the Chinese government just how strategically important you see these components.
This isn't the kind of thing you can hash out in public and go back and forth on. Once you put it out there, the other party will take steps to make sure they don't have to rely on us in the long run.
The export controls were not revoked, only reduced, and not before, but after China refused to buy low performing chips. Top gear was and is still sanctioned, as is any EUVL equipment.
And to add to the above: by building their own supply chain for chips, China is helping the unprivileged, those who can't front-run the market with long-term contracts. If China wasn't producing their own chips, the prices for us would be even higher.
Similar to the war-pricing of oil, China's reduction of imports is actually helping to keep our inflation from going even higher.
I don't see a situation where subscription payers move outside American LLMs (chatgpt, claude, gemini)
And I don't see a situation where serious API payers are OK with handing the Chinese state all their data. Like manufactures of decades past did and learned a hard, even existential, lesson for it. The state mantra has been "Collect and Copy" for a long time now, tech just hasn't had that moment to experience it yet.
So that leaves local hosting/leasing, but one of those has totally non-practical economics and the other doesn't have enough compute to meet any kind of real demand.
I also have yet to meet a single person who isn't neck-deep in the tech space mention a Chinese LLM. It's 100% the big American three.
If anything it's custom chips from the labs that threatens Nvidia.
These open models serve as price / performance pressure. Not all tasks require frontier models and cheap open models can be quite good for in-app assistants, if you're building that sort of thing. We also aren't sure the subscriptions will continue to be sustainable. They're currently subsidized to the tune of 50-70x. As someone who is hitting limits weekly that would easily cost me over $10k month per sub.
Google probably serves more tokens then OAI and Anthropic combined, even if many of those tokens aren't from explicit gemini requests, but from AI overviews and other service integrations.
xAI is already selling spare compute, and basically exists just to gas spacex's perceived valuation.
Casual consumers are using American models because their usage is low. As usage scales, the economics heavily favor open weight models. The API pricing from American companies is absurd. This is particularly true in an enterprise setting.
Open weight model hosts don't have the compute to meet enterprise demand. A large part of why these models are so cheap is because overall demand for them is incredibly low. Back in May, Gemini alone was doing about a month's worth of Openrouter tokens every day.
I disagree totally. DeepSeek raised prices because they couldn’t serve the demand. But there are tons of American vendors ready to fulfill it. Many enterprises, including the one I work for, are swapping to open weights.
I am not ok with handing all my data to American companies that are best friends with the American surveillance state. I still remember the Snowden revelations. Chinese companies are a much better option in that regard.
you don't have to hand them your data, the models are available so you can run them on bedrock yourself (or use another US housed inference service). and for what it's worth in my job i have access to data that gives a picture of the way companies are doing inference, and they're using a lot of chinese models (deepseek-v4 is a huge percentage of inference requests for example)
Lead doesn't really matter anymore. I just ported a very old cuda library to rocm, so it can be run on MI300s. 2 years ago this would have been a nightmare. Today it was an afternoon.
Presumably the efficiency numbers they're quoting are for the high concurrency state they were serving.
RAM was probably the bottleneck for the amount of context they were offering.
I assume it would run a little faster with lower concurrency but "RIP nVidia" is a little premature. The cutting edge inference hardware is amazingly powerful
It's also in this very announcement, in the first paragraph:
> Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips.
Most US companies that have anything to do with government, finance, medical, etc. already have contractual or regulatory obligations which prevent them from using Chinese hardware or services, even before the AI boom. That's a huge market.
Nvidia will do just fine. (Disclaimer: not a shareholder. At least, not directly.)
>> "They are already there on open weight models and Jensen knows that it is only a matter of time until China catches up with GPUs or other AI accelerators."
It is also why Nvidia becoming a bank for other AI companies who are unable to find VCs to fund them isn't really a good thing and that is bearish.
Not really. Chinese AI companies were never using NVidia AI chips.
This announcement doesn't really mean anything at all. It means the very few people who are already using Z.ai's API will continue to do so, but the vast majority of money going to Nvidia is through the massive amount of business going to Anthropic, OpenAI, and other western cloud providers and inference providers, who are mostly using NVidia chips for inference.
Also, NVidia chips are still sold out and supply constrained.
What is there to phish? These are simple vibe coded websites providing information for a certain topic, nothing else. In this case, that 2nd url is a website with information regarding the new model as well as a broken chat interface to try out.
You do realise there are hundreds of these types of 'sites'.
This one that is listed is designed to rank on Google as an informational source (although unofficial and not from z.ai which is why I said it is phishing)
Assuming you are technical you are able to discern this, imagine the average person.
> You do realise there are hundreds of these types of 'sites'.
Yes, and its these sorts of websites I am asking about.
> This one that is listed is designed to rank on Google as an informational source (although unofficial and not from z.ai which is why I said it is phishing)
Even if it weren't slopped together, 65% vs 80% on 10 tasks just isn't a significant difference. For 80% power to distinguish at a significance level of 0.05, you'd need more like 140 samples, if those were the true success probabilities.
The number one problem in LLM benchmarking is that people try to draw conclusions from sample sizes far too small to conclude anything but "it works sometimes, it fails sometimes, hard to say which is better." (The number two problem is that people run benchmarks blindly without checking that they measure something meaningful.)
They should have tested something else, like a raw H265 4K video or something.
Can't imagine "pro" users would be amazed by the battery life of watching netflix.
They write both. They write x86 repeatedly in the article and title, then show an instruction matrix that doesn't include, for example, the 468 CMPXCHG instructions or the crypto extensions PCLMULHQHQDQ instruction. Best I can guess, they mean 8086, which they think is equivalent to x86
Why is the 8086 not equivalent to x86? PCLMULHQHQDQ is from the CLMUL extension, which only began appearing in CPUs in the early 2010s - are CPUs from before then not x86?
x86 is an overarching group. Each processor is backwards compatible, I believe, so a 486 can run 8086 code, but they are not equivalent. If I download an x86 version of a program, I don't expect it to be written only in 8086 instructions
When you download an x86 program you're making a lot of other assumptions too, such as what the target operating system and hardware are. Even 8086 MSDOS software won't directly work in this emulator because it's not emulating DOS nor an IBM compatible, it has it's own addresses for the I/O. It's still x86 though.
You could also swap to an distro where apt ugprade can't brick things, and where if you manually mess up you can rollback cough cough nixos cough cough
Parameterized queries have been a thing for decades, which mitigate SQL injection attacks.[1] This is true of the examples in the post too, they used this:
query = """
SELECT * from tasks
WHERE id = $1
AND state = $2
FOR UPDATE SKIP LOCKED
"""
rec = await self.db.fetchone(query=query, args=[task_id, TaskState.PENDING], connection=connection)
Parameterized queries fail to protect from SQL injection for decades, because database engine developers fail to listen. What could work instead, if any parameter could be safely injected:
SELECT $1, $2($3) FROM $4
WHERE $5 $6 $7
GROUP BY $1
ORDER BY $8 $9
but at that point SQL loses its point and turns into MongoDB query language.
Porsager’s Postgres package does a great job of letting you feel like you’re writing raw sql, but avoids the attack vectors.
Anyway, I agree that ORMs are pretty terrible. I like writing SQL or using a lightweight builder like Kysely. Was a huge Dapper fan back in my C# days.
There are plenty of reasonable alternatives to ORMs that don’t open you to SQL injection attacks.