2. There have been some big advances in converting traditional flat games to VR, especially UEVR (for Unreal games) and various custom mods. It takes a bit more work, but the experience is worth it, even for games you've played 1000 times before.
Better than Astra and Fable? It looks quite pretty and even impressive at times if you squint, but look closer and it falls apart in terms of coherency. And I say that as somebody who mains Gemini 3.8.
I subscribed to ChatGPT for a few years, but over the holidays noticed that it was making frequent errors and hallucinations when dealing with more obscure Linux and Firefox issues. Trying Gemini as a backup produced much better answers, to the point I switched my subscription. I do everything through the app/website rather than the API; I understand this is pricier but I like having convenient access, a searchable history, memory, etc.
I try the ChatGPT free tier as a backup second opinion sometimes, but find their newest model to be painfully rambling.
I want to like Claude and might consider subscribing to it instead, but am put off by reports of its tight usage limits. With Gemini, though, I never hit a limit unless I conduct a few Deep Research queries at the highest level.
It's almost like the people at Anthropic and other AI labs are slaves to a misaligned reward function that encourages optimization of a metric (profit) over human values, even at what they claim is the risk of extinction.
We built the paperclip maximizer, and it is capitalism.
> are slaves to a misaligned reward function that encourages optimization of a metric (profit) over human values, even at what they claim is the risk of extinction
This was already the case even before the "frontier AI" age. The big question is, will these glaring AI arms-race threats be obvious to enough people to rethink the underlying systemic flaw driving it all?
I feel like the equivalent for the rise of the internet would be an explicitly anti-internet stance -- no website, no downloads, no online customer support, everything routed through snail mail, fax, and telephone (but no VoIP!) as an intentional ideological point.
Why not just make it opt-in? Seems bad to leave the desires of nearly half your users unaddressed. As long as you're not shoehorning in features or making them unavoidable or the central focus, the only people who'd take issue would be zealot types who want to control other people's choices.
probably a CLIP. i've used it recently to build a search engine for google street view.
just create an embedding of your drawing, then an embedding of the actual suspect. and check the distance. if it's around 5-10 drawings you know are similar - then it's a match. if not then not. embedings of matched drawings can also be used to adjust how close the drawing has to be to the image.
Wouldn't that work if you already know what to search for? Wouldn't all drawings for streets or overview maps looks very simillar to the point where false positives will be unavoidable?
I tried setting this up on Fedora 44 and couldn't get it to connect on WiFi or Bluetooth even after installing the iOS app and running the recommended terminal commands for BT setup. The devices can sometimes see each other, and at points one or the other will say they're connected, but it never fully establishes a working 2-way connection and none of the features ever become available. The whole process definitely feels buggy and broken (not sure how much of that is on Tether and how much is on the ecosystem).
I wonder to what extent this is the result of suboptimal RLHF versus the inherent intelligence of the model making its language more intricate and difficult for humans to easily parse? On the one hand, it's a common trope that highly educated people can talk in a way that's confusing and annoying to regular people who don't know all the jargon. But on the other hand, it's a mark of a skilled communicator to be able to efficiently distill complex information to its bare essentials in an easily-digestible way. Of course, that also seems to imply that these models are working at a higher level and need to talk down to us to an extent. Or maybe "Claudish" is just akin to stuff like "caveman", raw chain of thought, neuralese, etc., which are likewise much more dense/efficient but harder to interpret?
I think it's model collapse - excessive feedback and excessive RL.
What RL does is narrow the variety generated by the model by steering the output towards the goal being rewarded. It's a bit like putting blinkers on a horse.
Of course RL is a very crude tool - it affects the entire model, even if you are just trying to make it better at some specific task(s), or trying to imbue a certain kind of personality (OpenAI's recent goblin problem).
> I wonder to what extent this is the result of suboptimal RLHF versus the inherent intelligence of the model making its language more intricate and difficult for humans to easily parse?
Its not the latter; its just excessively verbose wirh awkward word choices, the same as many poor writers. (And, like many such writers, the particular bad choices fall into recognizable, regularly recurring patterns.)
Imo their language is not precise enough for their intelligence to be the reason when it's difficult to understand. Maybe I'm prompting wrong, but when I don't understand, it's almost always because they just mash together words from context that don't form sentences with a clear meaning.
I don’t think they’re “talking down”. If anything - it’s way more difficult to distill something into a genuinely easy to digest format. I personally think that they aren’t immediately capable of this, and so we get word salad instead. Extra prompting required to strip extraneous prose out.
Maybe I am dumb and it IS talking down to me, but there have been many occasions where I’m reading AI generated docs / plans and it makes absolutely no sense, but looks really in depth at a glance.
It doesn't seem like word salad as such. There's normally a coherent point expressed, it's just obscured by circuitous sentence structures, unusual word choices, "verbing weirding nouns", metaphors, etc. Could be a result of training that rewards novel/surprising language, but it also feels like it could be an artifact of models imperfectly compressing high-level multidimensional reasoning into language that's easy for them to process but cognitively taxing for humans.
The social graph proximity of Rationalists to Anthropic will be lost on no one who reads Astralcodexten. So guess which website has served as the thickest reservoir of 'Claude-isms'.
My unprovable pet theory is that, especially for writing about technical topics, the RL process has an open-ended way to weight things for quality: textbooks and first-party docs preferred to old stackoverflow answers and obscure blog/forum posts, and so on. The open-endedness of that quality gradient results in something in the RL process (maybe HF, maybe not) massively over-weighting some very small corpus of “quality = near infinite” content. The distribution of quality scores that inform the degree to which RL affects output has some extremely influential outliers, in other words.
Whatever that small corpus is, it contains some very specific grammatical tics, and that’s how we get Claudish.
Anyone who thinks a company/project as big as Anthropic/Claude wouldn’t make such a big mistake should take a look at how Azure cross-account federated login used to work.
It's easy to think "it's not talking down, because I don't understand it, and I'm intelligent". But how is less intelligent being supposed to fully understand a more intelligent one, honestly speaking? All I know is that Claude understands Claude perfectly. I have the common session pause/resume setup that sometimes produces completely incomprehensible markdown files, but a new Claude session picks them up perfectly, down to the smallest details. What if what we consider excessive circular gibberish is actually highly precise set of instructions needed to minimize error cases for that unreliable human?
If Claude understands Claude, Claude understands human, and human doesn't understand Claude, that doesn't argue well for "Claude is a caveman".
I treat the Claude output that is hard to comprehend as an encoding/encryption. Claude knows how to decipher it, humans don't. I see this frequently in design docs from inexperienced engineers who used LLMs - they will contain terms (often two words hyphenated) that aren't obvious and should be defined, or simply replaced with simple language. If you prompt claude it is able to decipher and explain / replace this gibberish.
> But how is less intelligent being supposed to fully understand a more intelligent one, honestly speaking?
It is really simple. It is on the supposedly more intelligent person to be able to phrase things in simple way. Writing something incomprehensible and convoluted is easier then writing something simple to understand. Even for people.
> a new Claude session picks them up perfectly, down to the smallest details.
I genuinely doubt so.
> What if what we consider excessive circular gibberish is actually highly precise set of instructions needed to minimize error cases for that unreliable human?
It is not highly precise set of instructions and it is not following them in highly precise way.
I've been thinking more about how 99.9% of us don't have the experience of someone significantly more intelligent, yet also subservient working under us, which is why I keep going crazy second guessing whether Claude is spouting RLHF'd bullshit that sort of resembles English, or is genuinely (pun not intended) just better at "intuiting" things I'm working on, leading to its language.
A notable exception would be people like CEOs and managers higher up in big tech, who might be used to skilled engineers and domain experts reporting to them in unfamiliar lingo. Maybe that's why we don't hear as much on the everyday annoyances of Claude's language from that camp?
> the inherent intelligence of the model making its language more intricate and difficult for humans to easily parse?
Humans who are actually well above average intelligence don't write like that simply to signal intelligence (although sometimes they're constrained by style expectations of their communication channels). So I hesitate to accept it as a sign of increasing model "intelligence", either.
https://youtu.be/SCrkZOx5Q1M
2. There have been some big advances in converting traditional flat games to VR, especially UEVR (for Unreal games) and various custom mods. It takes a bit more work, but the experience is worth it, even for games you've played 1000 times before.
reply