I tried this on M1 Pro today with 16GB ram and it worked!!!
I was using vscode and it seemed to interperet the system prompt right and then started actually inspecting and doing stuff.
Unfortunately the vscode system prompt is 24000 tokens, and I was getting 100 at beginning, 69 by the end of it, but honestly I'm super impressed. Great work team
1
All right there is so much hate for this rust rewrite, but I'm going to tell everybody that people are way way underestimating how sucky the Zig codebase was.
Basically all the time people were having horrid memory leaks and stuff that they couldn't even figure out.
So yeah, I get it, people are frustrated, but stop pretending Bun was this battle tested mature thing. It was a prototype, and still is.
Switching it to Rust is understandable since they just couldn't fix the memory leaks. It was and still is in beta, because anything with memory leaks is beta.
So yeah, I understand the concern but nobody would ever use it in production anyways for a server now because of the leaks so treat it like what it is - beta software and a cool experiment.
Ok heres the thing you will nevwr be able to truly do this due to logic.
Logically five people pooling their resources beats one guy.
therefore datacenters will always win because they get higher time utilization.
so forget it.
I always wonder the same but i let logic tell me its a fantasy, on average you cant outspend a whole group of people making better use of the hardware.
you will get better hardware though, cutting edge will always be cloud
Laptops/desktops are cheaper per flop than any datacenter hardware by a good order of magnitude.
The problem is that expectations rise in datacenters, hardware/power/security/availability guarantees cost real money. Then the operator providing these guarantees expects some margin.
You can see this most clearly with "developer desktops", a gcp instance costs about 10x a hetzner instance which costs between 5 and 10x the same hardware sitting in the back of an office somewhere. While all of these premiums matter for 24/7 systems under active development, they don't really matter for ephemeral small scale workloads.
Paying 2k for something that you use 100 hours of is quite expensive. Having the capability built into your existing silicon which you would buy for 1k is cheap. Paying 200 dollars a month for 2 years give a present value of $4200 dollars. Meaning that that paying 2k upfront cuts your overall spend in half.
I spent 6k on codex last month which, if repeated, implies a present value of ~144k.
They just mean this part: "where I upgrade hardware in order to upgrade my ai as an alternative to an expensive subscription."
Upgrading local hardware will remain the more expensive alternative to the subscription regardless what the relative cost of running the models themselves are. If the local hardware to do so becomes affordable then the subscription will be even more affordable, not expensive.
At least for these kinds of mega tasks. For more micro task we will always end up with unutilized local compute we already purchased which will be "free" since we already paid for non-AI reasons (e.g. a gaming GPU while not gaming).
Yes its a huge deal because these are starting to get bound by memory bandwidth not compute. therefore one bit wirfhts stream way faster leading to substantially better results. At least thats what Id guess!
I for one am glad they did it because I felt like it was a little bit... I just felt the company as big as Google, it felt wrong for them to be piggybacking off VS Code.
I know it was MIT license, but honestly, I feel better knowing that they're making their own product. So, you know, it goes both ways. I think they did the right thing to migrate off Microsoft's code just out of respect.