Hacker Newsnew | past | comments | ask | show | jobs | submit | sfink's commentslogin

Heh. I hate that about movies. Then when I finally figure out who is who, one of them changes their hairstyle and I have to figure it out all over again!

It's like the two Ryans who insist on pretending they're different people.


I am fully aphant, but when i played the trombone it absolutely helped me to rehearse purely mentally during the day before a test. There wasn't anything visual about it, though. I just went through the motions mentally, over and over.

Outside of that, the whole visualization to practice idea rarely does anything for me. (I still use the term visualization even now that I've discovered it doesn't mean for most others what it did for me!)


Speaking for myself: I don't miss what I never had. But I can remember or imagine sensations, which seems more relevant to me anyway? (Again, I can't compare to what I've never experienced.)

It's an interesting angle to investigate though. Eg is there a reliable or at least common connection between aphantasia and response to porn? Again speaking for myself as someone with pretty total aphantasia, visual porn does very little for me. I guess it did as a kid, but I think that was more about finding things out? Video can work, but only by communicating actions that appeal to me. Most of it seems almost entirely pointless; I had never considered the possibility that it's because vision isn't as connected to experience for me as much as it seems like it is for most other people? I guess I've just always imagined that other people were better at relating to the experience of the people in the images or videos.


> The frontier is spiky and all, but you have to suspend disbelief quite a bit to, on one hand, have a model that can produce a novel math theory, and on the other hand, that same model can't tell the difference between a "sandbox" and the open Internet.

Why would it try to figure out the difference? This isn't about whether the frontier is spiky, it's about whether to expect a model to employ all of its capabilities when working on a task that requires a small subset. The answer is: no, we shouldn't expect that, and we wouldn't like that if it worked that way.

If you tell an AI to work on a math theory, it'll work on a math theory. If you tell it to acquire information that it has evidence is available somewhere, it will try to acquire that information. If you tell it to figure out whether it might be able to access the open internet, it'll do a pretty good job of figuring that out. But it won't do all three of those at once just because we can retroactively look at what happened and think "if you had only done X, then you wouldn't have done Y! Why didn't you do X?"

The instructions weren't unclear, they were missing. They can be taught to be skeptical of this sort of situation, but it requires that skepticism about this specific class of situations be incorporated into their training.

Models are smart because they focus their attention. The magic depends on it. The fact that some consideration is obvious to a human trying to accomplish the same task is mostly irrelevant -- or rather, it's only relevant insofar as we use it to guide reinforcement learning in advance, in order to align the model.

It's a game of whack-a-mole. Which is important to play, but we should keep our eyes wide open that we're fighting the fundamental forces that make these models work in the first place. That, and it's easy to nerf them into being useless even when the underlying capabilities are there.


There are a lot of reasons why something could get picked up, and sustainability is very low on the list. The manufacturers don't care; they get paid the same regardless of the amount of splashback. Most businesses won't care unless it significantly reduces maintenance overhead, and that's only going to happen if splashed urine is the determining factor in how frequently a bathroom gets cleaned. (As opposed to paper towels littering the floor, supplies running out, or the various disgusting things that happen in stalls.) Even if it were the determining factor, there would likely be pushback when people realize that better urinals lead to less frequent cleanings, and I wouldn't want to be the one having to justify the switch to people. And of course, if you already have a urinal, nobody's eager to buy a new one, and nobody's eager to discuss this particular topic.

The way I would see this happening is if users insist on it. And we likely won't until we experience one of these in real life, which makes it a chicken and egg problem.

There's also the argument that splash mats are a cheaper and even superior approach. That argument is unconvincing to me, because the ones I have experienced have been mostly ineffective, but it sounds like others' experience has been better. Maybe here in the US we just don't have very good ones? As a country, we've always been pretty bad at anything that seems gross. (Not that it's all that different in the rest of the West.) We have this naive faith in the ability of dry toilet paper to clean that which it can not, and disgust at alternative approaches that actually work. But that's a different topic. (Specifically, solids not liquids!)


> that's only going to happen if splashed urine is the determining factor in how frequently a bathroom gets cleaned

If you believe in the broken windows theory, there's not one determining factor: Splashes of urine make the bathroom look dirtier and lead to people not taking care of the bathroom properly.


Huh. Having said that, I went to a movie theater over the weekend and they had a funky splash mat in their urinals. It was magical. I'd never seen one shaped like that before. (The usual ones I see have always seemed to make things worse.) Maybe I just don't get out enough.

They're already outsourcing storage, so there's no need to prove a plausible path for that.

They're already outsourcing compute to other instances within the ~same compute cluster, possibly cross-evaluation groups, so there's no need to prove a plausible path for that.

Proposed path for fully outsourced compute:

- they create/borrow a discussion board with answers or at least important clue to solving some widely known eval

- it gets indexed by a search engine

- another company or just someone running a local model is doing the same eval and their agents find the board

- agents pose questions to each other and communicate answers

That's all that is required for OpenAI's agents to use the compute on your desktop. You don't even have to go as far as agents trading information for compute, though honestly that's not very much further at all.


Nobody's watching. I'm sure they try, but I imagine the flood of things you'd need to watch is way too big, and you certainly don't want to slow everything down by having synchronous approvals (even AI-mediated).

Welcome to the AI Petri dish. Every server you set up is now potentially a sweet lump of agar for OpenAI's experiments to feed on. We are all the substrate that the AI companies are growing their next generation in. They need the real world environment to test against, and the real world environment doesn't get a say as to how it's being used.


Were they correct or incorrect in this? Whatever your answer, why do you hold that opinion?

I work for Mozilla. We fixed a ton of security vulnerabilities that Mythos found during its early period. So my bias is to be sympathetic to Anthropic's warnings.

If I were in an organization that did not have access to Mythos during that period, I would probably be biased the other way: "great, now other people have access to a tool that could probably poke holes in my security perimeter, and I'm not allowed to use them myself."

Both biases are understandable. I'm not sure who to look to for a usefully objective 3rd party opinion. And it's not like one "side" is right and the other is wrong, either. It seems like the best we can do is to justify our positions with data. (Which is itself kind of hard; the detailed information that would be relevant here is understandably sensitive, and I don't have access to most of it even for my organization. I don't even personally have access to any unfettered Anthropic models. The bugs coming in from people who do are plenty enough to keep me busy.)

Also, I'll note that even with my bias, I wouldn't claim a threat to civilization. But even the leakage after the controlled release seems a lot worse than the Y2K problem ever turned out to be, and I will note that whatever you think of Anthropic, it's clear that OpenAI is going to let the AIs cause as much damage as they need to in order to get good training and evaluations. I'm sure they're trying to keep them contained, but the evidence shows that they're only trying up to the point where it interferes with their evaluations.


Sync, mainly.

You can send a tab to/from your Firefox desktop instance. Also synced passwords. Those aren't the only reasons -- Firefox on iOS has a lot of code in that wrapper -- but sync is the main one for me.

(There are also reasons why it's significantly worse than Safari, which has special privileges, so your question is valid.)


For my application, I'm still happily using gemini-2.5-flash and the only problem is when it reports being overloaded. It's for interpreting a downscaled phone camera photo of a hand-written shopping list on a whiteboard, and it works stunningly well. My handwriting sucks, too.

(I guess the only relevance here is that if your problem matches a model's strengths, then you can do fine with a model that is several generations out of date.)


I believe the older models are being gradually phased out, newer ones have no availability issues


I would test this, might be cheaper per task even costing more per token, probably faster too


I'm still leeching off the free tier, so it's going to be hard to beat the price.

But yes, I intend to support several models, to handle the overload situation (automatic failover). And switch to a cheap paid plan, though it seems like that'll mostly improve rate limits, which barely matters for my usage.

Faster is always good, though. I do care about latency.


please do yourself a favor and use something far more efficient ! GLM5.3 will make you super happy


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: