Ok, at first I didn’t get it and thought it would make sense if stripe just wants to build payments for agents, but on a second thought your idea about “once buy tokens — use everywhere” is very good one!
I still don't understand it, yes it's a lot of data and presumably they're already shunting it to cpu ram instead of keeping it on precious vram, but they could go further and put it on SSD at which point it's no longer in the hotpath for their inference.
I don't think you can store the cache on client given the thinking is server side and you only get summaries in your client (even those are disabled by default).
If they really need to guard the thinking output, they could encrypt it and store it client side. Later it'd be sent back and decrypted on their server.
But they used to return thinking output directly in the API, and that was _the_ reason I liked Claude over OpenAI's reasoning models.
I assume they are already storing the cache on flash storage instead of keeping it all in VRAM. KV caches are huge - that’s why it’s impractical to transfer to/from the client. It would also allow figuring out a lot about the underlying model, though I guess you could encrypt it.
What would be an interesting option would be to let the user pay more for longer caching, but if the base length is 1 hour I assume that would become expensive very quickly.
Just to contextualize this... https://lmcache.ai/kv_cache_calculator.html. They only have smaller open models, but for Qwen3-32B with 50k tokens it's coming up with 7.62GB for the KV cache. Imagining a 900k session with, say, Opus, I think it'd be pretty unreasonable to flush that to the client after being idle for an hour.
Yes — encryption is the solution for client side caching.
But even if it’s not — I can’t build a scenario in my head where recalculating it on real GPUs is cheaper/faster than retrieving it from some kind of slower cache tier
Next time you use true real independently audited e2e communication channel, don’t forget to check who is the authority who says that the "other end" is "the end" you think it is
I agree that this is still the most important thing, and I don’t try to challenge this.
At the same time we have quite adopted bumping our dependencies when it does not incorporate breaking changes (especially if there are know security vulnerabilities) — and my point is exactly about it, why even simple renames, extraction or flattening or other simple changes have to be treated so differently than internal changes that do not touch public interface?
My thoughts exactly. JSX provides the best templating syntax I have seen - it's just JS, and it uses curly braces to delineate JS. Putting JS, or worse, custom syntax in strings is terrible, and every other delineator choice is less idiomatic and uglier than curly braces.
I see your stance! There are two ways to this: JS-first (React) or HTML- first. Hyper takes the latter: purely focusing on the semantic HTML structure when assembling interfaces. Focusing on pure structure (like React 1.0) and delegating design and logic to concerns that master it the best.
> allow to select the purchase price within the last 2 years
I don't think that's true. My reading of that is "you lock in the price on your start date and can keep that for the next 2 years going forward". That doesn't help anybody joining at >$1k / share. :D (and that's only ESPP, not standard stock compensation).
ESPP is a very small amount vs RSUs. You’re limited to buying $25,000 per year (that you still have to shell out for even if it’s at a discount) vs just being given several hundred thousand (or more) in RSUs.
while (true) { askModelToBeginConversationIfAppropriate(model, previousContext, thingsHappenedSince); sleep(concisenessTick); }
reply