I might be missing something here, but could this concept not also be used in the same way for harm? For example if a model is trained to replicated a sucessfull scammer?
Definitely! In fact that seems to be a central aspect of the proposition:
> Above all, a GA should amplify the principal, and not simply substitute for them for someone else’s purposes or benefit.
[…]
> A GA must be aligned with its principal. It should not be designed to manipulate or control or guide the principal in any way which does not derive from the principal themselves. “Constitutional AI”, “Terms of Service”, “social harmony” etc. may all have their place, particularly for widely deployed superintelligent systems—but inside the privacy of a GA, the principal must have freedom from optimization pressure.
…I read this to suggest that it should amplify a scammer’s scamming, a thinker’s thinking, a tinkerer’s tinkering, a cop’s sleuthing… and I’d imagine it implies amplifying a person’s capability to avoid being scammed, too…
One man’s scam is another man’s “pro-social nudge” and another man’s “attractive opportunity” and another’s “advertisement for a delightful consumer wonder” and another’s “patriotic duty to sustain demand to prop up the too-big-to-fail ideas we bet the whole economy on.”
When you fix and operationalize all values centrally, universally, and externally to the principal… that’s current-gen frontier chatbots, not Gwern’s GA concept.
This model is shockingly small for how capable it is. its a little bit bigger than deepseekV4 flash but around as capable if not more on some benchmarks than V4 pro, i wouldnt be surprised if this becomes a popular local model.
I've been wondering about that. GLM-5.2 is also half the size of DeepSeek V4 Pro. (But costs roughly twice as much.)
I looked into DeepSeek's architecture a little bit and the main focus was how can we save as much money as possible. They did a lot of cost cutting with the attention mechanisms. This allowed them to offer an insanely cheap price even on massive contexts, but seems to have come at the cost of performance?
At least, that's my guess, when I see smaller models costing more and outperforming, I think, "they must have denser attention?"
The current Deepseek V4 Pro is still just their initial preview AFAIK, with the "real" model release rumored to come later this month. GLM-5.2 might be outperforming simply because it's had more post-training on top of the GLM-5 base.
Yeah i shouldve been more clear, a model of this size could run on 2 dgx sparks so out of the range of a lot of the typical consumer sure, but I think there is definitely a market for that size
Yes! You can use local models through Ollama and LM Studio. We do some special handling when you use local models, such as suppressing background agents when you chat, so the model is not overwhelmed.
I'm struggling to see how this could lower the cost to ideate here, almost all of the designs I can see on your page would require re drafting/designing it would be more beneficial for people to look at homes that have already been built compared to an AI generation with no real world grounding. Also who is paying for the compute to generate these floor plans and renders? This would have to turn into a paid service eventually im assuming.