Hacker Newsnew | past | comments | ask | show | jobs | submit | nshotton's commentslogin

you could try using a local vision model, like Mage-VL from microsoft. Its only a 5b model so its quite small for the capability it has.



That is for 2.5T parameter one, this is the correct countdown for. the 27B one https://modelscope.cn/models/Qwen/Qwen3.8-27B


this seems to be a countdown for the 2.4t model - which is gigantic and is not as exciting for those running models locally


it looks like it says the 27b will be included in the release too, we will find out in 30 minutes lol


Nope 27B is for Friday: https://modelscope.cn/models/Qwen/Qwen3.8-27B

counting down to:

Friday 2026-08-14 17:00 CEST

Friday 2026-08-14 15:00 UTC

Friday 2026-08-14 08:00 PDT


I guess this one is for Qwen3.8-2.4T-A95B from the URL


Thanks for posting this, it made a huge difference tweaking the recipe.



I might be missing something here, but could this concept not also be used in the same way for harm? For example if a model is trained to replicated a sucessfull scammer?


Definitely! In fact that seems to be a central aspect of the proposition:

> Above all, a GA should amplify the principal, and not simply substitute for them for someone else’s purposes or benefit.

[…]

> A GA must be aligned with its principal. It should not be designed to manipulate or control or guide the principal in any way which does not derive from the principal themselves. “Constitutional AI”, “Terms of Service”, “social harmony” etc. may all have their place, particularly for widely deployed superintelligent systems—but inside the privacy of a GA, the principal must have freedom from optimization pressure.

…I read this to suggest that it should amplify a scammer’s scamming, a thinker’s thinking, a tinkerer’s tinkering, a cop’s sleuthing… and I’d imagine it implies amplifying a person’s capability to avoid being scammed, too…

One man’s scam is another man’s “pro-social nudge” and another man’s “attractive opportunity” and another’s “advertisement for a delightful consumer wonder” and another’s “patriotic duty to sustain demand to prop up the too-big-to-fail ideas we bet the whole economy on.”

When you fix and operationalize all values centrally, universally, and externally to the principal… that’s current-gen frontier chatbots, not Gwern’s GA concept.



Oh, it is. I was looking at the Huggingface repo which listed the lower number at the top of the page, looks like that's wrong.


Same, I was relying on that. Weird!


This model is shockingly small for how capable it is. its a little bit bigger than deepseekV4 flash but around as capable if not more on some benchmarks than V4 pro, i wouldnt be surprised if this becomes a popular local model.


I've been wondering about that. GLM-5.2 is also half the size of DeepSeek V4 Pro. (But costs roughly twice as much.)

I looked into DeepSeek's architecture a little bit and the main focus was how can we save as much money as possible. They did a lot of cost cutting with the attention mechanisms. This allowed them to offer an insanely cheap price even on massive contexts, but seems to have come at the cost of performance?

At least, that's my guess, when I see smaller models costing more and outperforming, I think, "they must have denser attention?"


The current Deepseek V4 Pro is still just their initial preview AFAIK, with the "real" model release rumored to come later this month. GLM-5.2 might be outperforming simply because it's had more post-training on top of the GLM-5 base.


If the "final" release of Deepseek V4 Pro outperforms GLM-5.2 while maintaining the current Deepseek price, then it's going to be a marvel.


hardly, its still quite big unless by "local" you mean people that spend many thousands on rigs :)


Yeah i shouldve been more clear, a model of this size could run on 2 dgx sparks so out of the range of a lot of the typical consumer sure, but I think there is definitely a market for that size


> Hy3 has 295B parameters in total. To serve it on 8 GPUs, we recommend using H20-3e or other GPUs with larger memory capacity.

I would.


This looks really great! Quick question, is there any support for using local models?


Yes! You can use local models through Ollama and LM Studio. We do some special handling when you use local models, such as suppressing background agents when you chat, so the model is not overwhelmed.


sweet, is there any plans to implement VLLM in the future?


vLLM works today - it exposes an OpenAI-compatible endpoint, and you can point Rowboat at a custom base URL in model settings.


I'm struggling to see how this could lower the cost to ideate here, almost all of the designs I can see on your page would require re drafting/designing it would be more beneficial for people to look at homes that have already been built compared to an AI generation with no real world grounding. Also who is paying for the compute to generate these floor plans and renders? This would have to turn into a paid service eventually im assuming.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: