Benchmarks are useful but only on a log2 basis. One model performing at 50% and another at 75% is just as impressive as one model performing at 78% and another at 90%. Confoundingly, a benchmark becomes useless once a frontier model scores over ~95% on them.
I think that's definitely the right way to understand benchmark saturation, but there's a separate problem where the benchmarks are just not representative of real workflows even when they don't seem to be saturated.
If I get a better Gemma 26b-a4b MoE model out of it I'm all for it. Really wishing Qwen 3.8 would release a MoE variant of their 27b model size, but I wont complain if google beats them to it.
I think it'd be pretty easy to fix, too. Have a point on the line that represents "10 seconds from now" and orient your camera along the line between that and the point representation of your car, with slow, smooth panning so as not to get jerky transitions. If "10 seconds from now" is within a certain range of the car, it ought to just default to whatever its doing now.
Because I spent $15k in AI costs last month doing real work at my day job; I'd like to do the same thing for myself but I don't like lighting cash on fire.
reply