Strict determinism is a different, but related, issue.
E.g. if I've written a role playing character using a specific model I may want to pin the character to that model until I've been able to test the model being "better" doesn't affect the feel of the character before switching. That doesn't mean I need the character's responses to be completely deterministic, but that doesn't imply I'm fine with the character having a different quality or feel of response just because the new model is out.
It'd be nice if there was a more explicit way to signal in the request "I want what you think is best per dollar for this class of answer" vs "I want this model to answer".
the fact that this author cannot get qwen3.8-27b run at the same speed as qwen3.6-27b, says the article is not worth reading. the author does not know anything about how to run local AI. 3.8 and 3.6 are the same model with different weight.
both tg and pp speed are so terrible on author's machine.
Cerebras is an uncut whole wafer. Each wafer gets you 44GB SRAM (not a typo, SRAM, not DRAM/VRAM). A few years ago, leading process node wafer from tsmc is $20K/ea without guarantee on yield.
A single full wafer likely can run qwen3.6-27b alone. But won't be enough to run bigger models, which are pretty much all popular models.
on one side, deepmind makes a lot of advancement in science related application. but, on the commercial side, they struggle to compete with other major LLM providers.
if you want deterministic returns, you should set the temperature to 0 to get the best possibility of deterministic.
reply