Agree, except the probabilities for outcomes in the structured output. I don't think you can get those for most frontier LLMs (logprobas). You can get it for open source models but not frontier LLMs.
I've seen them talk about it a bit on Twitter -- it seems to be fairly well-calibrated in general, but obviously you need to test it on your use case and dial it in comparison with known data for best results.
Yup. LLMs can do almost anything. Can they do it at the speed, cost and confidence of a model like Jev is a different story. In Jev output tokens are straight up free because it's not doing text token generations.
People are already used to getting *exactly* what they want from an agent. e.g. in our agent (https://www.definite.app/) we often see people take pictures of a sketch on a piece of paper or a screenshot from a few other tools (Stripe + Excel).
Our agent has templates to start from, but ultimately writes a react app to give the user what they want. It'd be hard to get that experience in a framework like this.
> Tinybird is not affiliated with, associated with, or sponsored by ClickHouse, Inc. ClickHouse® is a registered trademark of ClickHouse, Inc.
Yeah, it sucks they need to do this. If I was a visitor to their website, I'd immediately want to know what ClickHouse, Inc. is and you'd realize ---> it's managed clickhouse, direct competitor... why would I use the one that needs all the ®'s
my car has an AI assistant. I don't know what the wake word is, but it will wake up and be useless with no obvious way to turn it off. I'm sure I could search for it, but it happens while I'm driving and I forget about it. It happens just infrequently enough to not remember, but often enough to be really annoying.
My wife turned on a meditation feature on our very old Alexa. I guess the app was deleted or something, but every morning Alexa turns on and says "Starting your meditation times ... Sorry, this is no longer available". If I tell Alexa to turn this off, it has no idea what I'm talking about and it shows up nowhere in the iOS app. Instead of meditation, I got a nice dose of minor annoyance every morning.
I'm not going to care about open models until some blend of the below becomes true:
1. the labs stop offering max plans
2. really smart open models can easily be run on my mac
3. TPS (token per second) AND intelligence are gpt5.6 level
on #1, it's nearly impossible for me to run out of codex tokens right now (I have 4 resets banked) and Fable 5 seems to be sticking around for the foreseeable future. I have virtually unlimited token usage for $400 a month, so open models being cheaper doesn't appeal to me.
on 2 and 3, benchmarks are showing some of the open models at around opus4.8 levels, which is incredible! But running them locally at anywhere near the TPS of cloud inference is far off. I can run a smaller (dumber) open model locally and get good TPS, but see #1, whats the point?
This is far cooler than how I've been using my Pi.
I've started building a box for managing agents. I hate sitting in front of my Mac all day flipping between terminals. I wanted something that was audio and whiteboard / paper first.
The Pi has a mic, camera, and projector hooked up. The Pi is always listening, so I just say most commands (i.e. "check logs on backend service, customer xyz said abc is wrong"). I can tell it "look at the board" and it can see things I've written or drawn and can project on top of or alongside anything on the board.
Not sure yet if it's more efficient, but it's definitely more fun.
They did it interactively with Claude, it’s possible that it played up the significance and humor of the findings in a way that the interaction left the user feeling like they were really on to something.
each "question" is answered in parallel instead of a sequential (like an LLM). so if you have an input like:
it answers is_it_hotdog and is_it_apple in parallel and gives a probability.reply