Benchmarks and cost don't really help me understand how good a model actually is. Anyone have hands on experience working with the current newest models? What work did you do, and how did the model performance ?
I have been using it since it was available in opencode Go plan.
It replaced all other models for me.
It's much less chatty than DeepSeek V4. And it feels noticeably smarter. I also use Claude at work and GLM 5.3 feels like Opus 4.8 (which I consider better than Opus 5).
It tends to be more proactive with suggestions after a task is done also.
DS4 Flash is awesome but GLM 5.3 is better despite being a bit more expensive.
I'm on 0.139.0 on Linux and my visible ~/.codex/logs are only about 129 MB.
This makes me even more conservative on upgrading these tools every time they prompt me to. Better to let them get a few miles on them and see how the community responds.
In this case it sounds like 0.142.0 reduced the issue but didn't fully settle it. I'll wait for 0.143.0+ and see if that version is more acceptable.
Cursor was my first hands-on experience with AI. I didn't know much about getting set up with specific providers via API, and Cursor made it easy to pick any model, ask a question about some code, and get a clear suggested answer easily viewable in the IDE with an 'accept' or 'reject' button. I think they answered this question well: "How do normal developers want to interact with AI?"
I moved away from Cursor when I noticed the responses from specific models were not as clean or accurate as when I'd prompt the models directly, which was something I didn't know how to do early on. I hypothesized that they had some boilerplate prompt sitting atop of my own, causing less precise or desirable results.
I would assume Cursor is still one of the best options for normal developers to get started with AI, but with Copilot forcing their foot in the door at many companies, I wasn't sure how well it would fare on its own. Being acquired by SpaceX should help, and I'll be interested to follow along and see how things develop.
My opinion is that the publicity can only help Cursor. I don't necessarily think SpaceX would make Cursor better. Copilot (which I view as a direct competitor to Cursor) has a huge structural advantage. I have several friends in various American companies where Microsoft products are all they are allowed to use. They get "free" Copilot access as a part of their Microsoft plans. Developers aren't having Cursor placed right in front of them in the same way Copilot is, and from my experience, when developers have the choice to pick one, they pick Cursor. So, I just feel the SpaceX/xAI publicity could help Cursor get more visibility in these general American software companies more so than they could on their own.
Given how easy it was to get our own version of an AI service tool set up, I'm shocked to see a huge, capable company buying it's way into that functionality as opposed to using their large existing datasets to fine-tune an internal tool. Is 3.6B really the better option?
PAX ERP. An AI-assisted ERP for small regulated/job shop manufacturers that have outgrown QuickBooks and can't afford (or take the time) to go through a traditional ERP implementation.
PAX is the easiest ERP to pick up. Our core idea is that ERP should not take weeks of formal training and implementation. We have no formal training, no implementation fees, and are a complete ERP+CRM GAAP accounting system. It's intuitive, well-documented, AI-assisted (never steered), and people can pick it up and start using it well with no questions to our team.
reply