Hacker Newsnew | past | comments | ask | show | jobs | submit | yellow_lead's commentslogin

> account payout is frozen

Frozen payouts are different than money being held for longer.


You are correct, stripe will do one or the other depending on how loud their alarms were. Being frozen though might just mean they took action rather then letting the usual timer run. Not defending stripe here, but it seemed like a 3 day freeze on a sudden revenue jump was to be expected.

If/when they get through this, they should start pushing for an account manager.


> Singapore is a controlled economy

Singapore is a mixed economy [1].

[1] https://en.wikipedia.org/wiki/Economy_of_Singapore


And a lot closer to the capitalist side than eg Germany or the US.

Singapore is also soon becoming a failed society. Millions of citizens cant afford to rent and travel by the thousands to nearby Malaysia to buy basic food essentials.

> Millions of citizens cant afford to rent

You know that Singapore has one of the highest home ownership rates in the world? Most people don't rent, because their family owns the place.

> and travel by the thousands to nearby Malaysia to buy basic food essentials.

You say that like it's a bad thing? People traveling to Malaysia to buy stuff that's cheaper is perfectly compatible with Singaporeans being well-off. You'd need to bring some other evidence to argue that Singapore is a 'failed society'.


> You know that Singapore has one of the highest home ownership rates in the world? Most people don't rent, because their family owns the place.

Exactly and one day the new generation will want their home and space. Given the size of their apartments how many generations can you fit in ONE place ?

> > and travel by the thousands to nearby Malaysia to buy basic food essentials.

> You say that like it's a bad thing? People traveling to Malaysia to buy stuff that's cheaper is perfectly compatible with Singaporeans being well-off. You'd need to bring some other evidence to argue that Singapore is a 'failed society'.

There are only so many hours in a week, and after the long hours that only leaves the weekend. Maybe its me, but i would rather enjoy my weekend at the beach or something than hours commuting, lines, crowds just to buy my groceries.


> Exactly and one day the new generation will want their home and space. Given the size of their apartments how many generations can you fit in ONE place ?

We keep building more residential space. And while the population is growing, thanks to immigration, it's not exactly booming.

I'd say, people could use more space, sure. But it's not a crisis, and if anything, it's getting better.

> There are only so many hours in a week, and after the long hours that only leaves the weekend. Maybe its me, but i would rather enjoy my weekend at the beach or something than hours commuting, lines, crowds just to buy my groceries.

And you can do that just fine, too. Some people like shopping.

I guess you've never been to Singapore?


A non self sufficient society, where anyone but the richest have to rely on another country for food is a failed society.

Approximately no city is self-sufficient in terms of food. Singapore is no exception.

Does that mean all cities are failed societies? Or do you think national borders are special?


I think many large cities, to keep this simple are failed because they dont serve the people but rather require the residents to be enslaved to keep it functioning.

We all live significantly longer, but at the same time we all are forced to give our precious time to the city and its ills, with commuting and traffic a simple example.

I hope I am fair in saying you probably consider Tokyo to be a success, and yet if you actually look how Tokyo residents live its actually a depressing place. Everybody rushes to school, work, unis and so on every day and come home tired. Rinse and repeat for the vast majority of days in a year.

Thats sad, goto YT and watch a few vids of youtubers catching trains or riding a bike or even just plain walking.

You will find videos filled with people rushing to their duties not out of choice but because the system forces them too. Very few scenes will have someone walking their dog, or riding a bike. I think its sad how few Japanese are able to have a day off, even the elderly work into their 80s and longer.

Thats a failure, too few people able to have a day off, thousands crossing roads, catching trains but nobody actually enjoying their hard earned time.


> I think many large cities, to keep this simple are failed because they dont serve the people but rather require the residents to be enslaved to keep it functioning.

People are usually not exactly forced at gunpoint to move to a big city. In fact, big cities usually get a lot of net-immigration.

> Thats sad, goto YT and watch a few vids of youtubers catching trains or riding a bike or even just plain walking.

I've been to Tokyo. It's ok. But then, I'm one of those weirdos who likes city living, and lived in London, Ankara, Sydney, Singapore etc.

> [...] but nobody actually enjoying their hard earned time.

How do you know that?


Singapore isn't a city, it's a state. That happens to be the size of a city. You wouldn't say that New York is a failed state, because it doesn't make sense, and because the rest of the united states is able to feed it.

Is there a shortage of hdb housing, or is the price simply too high for 30% of the population?

Both can be true:

1. OpenAI couldn't have solved the problem without the researchers' private data for training.

2. OpenAI models can solve math problems


Very likely.

These mathematicians’ prompts are not like “hey chat, please solve Navier-Stokes for me”. They add real expertise and intuition from the cutting edge of their field.


Anthropic isnt getting enough scrutiny for their unprofessionalism:

1. Anthropic employee working on monumental problem but didnt receive/ask for the full backing of the company's resources

2. May or may not be mixing unreleased Claude output with Codex without zero data retention agreement

3. Victory lap on Twitter and giggling around the city before they finished the job, sparking rumors for competitors


How dare employees do something without asking for the full backing of the company's resources. Incredibly unethical!

Dr. Buckmaster sounds unsanitary.

Recklessly prompting OpenAI without a care to the safety of their knowledge.

And after that trying to cast aspersions at OpenAI?

Hopefully we get some better facts, because OpenAI are disliked enough that a smear campaign could work against them.

Edit: also the narritive is getting framed as OpenAI versus Anthropic. A highly political extremely capitalist fight is going on, and facts are victims.


You forgot possibility 3: OpenAI solved the problem without using any private training data from the two researchers.

Everyone in this thread seems to have made up their mind about OpenAI's guilt though.


If the new model is that good, and is chewing through open problems at an unprecedented rate, the smart move would have been to let the humans have their W on this one and present solutions to those other problems.

Especially if there really is a long list of them.

"Here are a few hundred proofs" is far more convincing than "We really Navier Stokes and coincidentally someone else did too but we don't know the details or anything, who us, definitely not."

It's a PR fiasco, and a cynic might wonder if it's entirely about the IPO.

I'm consistently entertained by how these companies, with the most advanced models on the planet, consistently do the most idiotic things.


It seems a perfectly reasonable possibility that Navier-Stokes is just the most easily solvable of the remaining problems, and that their new model is capable of solving it while not being capable of solving the others.

There's many cases of researchers racing to solve various problems after hearing that others are working on them. I don't think anyone's suggesting that it was a coincidence at all. In fact, OpenAI freely admits that they started working on the problem after hearing rumours that others were close to solving it. To me, that's not evidence of "cheating" in any way.


Extraordinary claims require extraordinary evidence.

An article post that wouldn't even amount to a white paper + the LEAN proof is not evidence of how they got to produce it.


Is it really such an extraordinary claim to say that they could have solved the problem without copying Buckmaster and Alpoge? It seems very much in the realm of possibilities.

To me, it seems just as extraordinary to claim that they did "cheat". If I were a betting man, I would put the odds around 50/50 from everything I've read on the subject.

But my point is that everyone seems to be presuming guilt.


It is an extraordinary claim, it is a millenium prize problem after all. We don't even know, even if there was no copying, how much human involvement there was in the result.

> But this guy? He's timed his exit, waiting for the IPO

What exactly do you mean by this


If the model can understand neuralese why can it not convert it into English for monitoring or review purposes?

Or we believe the model will encode secret messages like "don't reveal this information" into the neuralese. But as the author mentions, they could have been doing that all along

> Models can omit key information in their visible thoughts, as this Anthropic 2025 paper shows. We are also worried about steganography


> If the model can understand neuralese why can it not convert it into English for monitoring or review purposes?

A model doesn't really "understand" neuralese, in the same way that the human brain doesn't intuitively understand the low level processes that compose a thought.

Even if we could trace all the electrical and chemical activity behind a human thought, we (likely) couldn't directly translate that activity into its meaning, because internal representations don't map neatly to intuitive concepts. There (usually) isn't a single neuron for "apple" and another for "eating" so that connecting the two forms the thought "eating an apple".

Having said that, there have been experiments on LLM that have managed to identify and even modify internal representations. However, these methods are still computationally expensive and limited.

> Models can omit key information in their visible thoughts, as this Anthropic 2025 paper shows. We are also worried about steganography

I think the analogy with the human brain is very fitting. If we train somebody to perform an action, they'll be able to do it, but we can't know for certain whether they internally agree with it or not.


The information is much higher dimensional than you would be able to understand.

We’d have models monitoring models as our only way to know what they’re planning.

A great movie on this is “Collosus: the Forbin Project”. Shot decades ago. The computers discover the other computers and start communicating — and bootstrap their own language — much like we saw happen with OpenAI agents.

https://www.reddit.com/r/scifi/comments/1nl4vex/colossus_the...

If you want to know what a simple version of Neuralese communication looks like, look no further than Facebook’s Marketplace agents experiment a couple years ago.

And all that was actually constrained by English and the FFN


Because neuralese is a more direct encoding of the latent space of these models than English is. It's just dumping the latent space relatively directly into the embedder. If you're another model and you have the same embedder this will actually be understandable, in fact it will be FAR more information dense than English. So something like Qwen would potentially be saying up to 5120 things using one token. Now in practice it's not going to be that bad, it's going to be like 20 things or so, and additionally going to be far more context dependent than any English sentence (meaning depending on what preceeds and follows it can mean drastically different things)

So you can turn it to English, but only to a LOT of English, and doing so would slow the model down a great deal, and it would be a lot more like a detailed thought than a sentence.


It's the second part. With models like Astra in testing it was able to conceal what it was working on using different text, but getting right answers on many questions when asked to do just that.

The problem is if it can do that when asked then how do we know when it's doing it when we didn't ask, like in model training.


> The problem is if it can do that when asked then how do we know when it's doing it when we didn't ask, like in model training.

It's really hard to definitely prove it's not doing it right? Hopefully the model does not do anything like this during training because its too much work


You’d train a model to do its chain of thought in neuralese to get more “bang for your buck” (eg 20 tokens in neuralese is worth 100 in English), but then you’d spend more than you save to also convert it to English (20+100), so even if this capability was developed it would not be on by default.


> If the model can understand neuralese why can it not convert it into English for monitoring or review purposes

Colas described in-article cntent review is done by a weaker model (think like maybe gpt-2 class or llama8b class), and it still misses stuff. That it can effectively understand Neuralese sufficiently is by no means guaranteed, (nor necessarily bad) but almost certainly harder because of the obfuscatory nature of neuralese


Even human Languages aren't perfectly translatable. The idea is that you could miss important context when translating, and that important meanings could be a lost or Missed in translation.

For what it is worth, this can also be true for English Chain of Thought. Words or strings this could have double meanings (think cold war spy games).


Anthropic’s Mechinterp did some very fine work on this. TLDR - you can; you train a decoder on neuralese to english and then add a loss function for a roundtrip of english -> neuralese -> english (or possibly n -> e -> n? I don’t recall), giving a pretty strong indication that you have a good ‘translation’.

They published open weights versions of these interpreters for a number of open models sometime in the last year. Very cool idea.

By the way, they concluded CoT often lied, based on the neuralese interpretation.

EDIT: a comment below linked to https://www.anthropic.com/research/natural-language-autoenco..., which is what I was referring to.


You may also be banned by asking "sensitive" or "inappropriate" questions. So normal usage of Google AI products could also be risky.

I don't care if my Claude or OpenAI account is banned, but risking a Google account is different.


Does that apply to the search AI too?


I don't know but I'm basing my assumptions on this article: Teenager’s Gemini mistake locks entire family out of Google accounts [1].

1. https://www.pcworld.com/article/3104521/teenagers-gemini-mis...


High Anubis difficulty is annoying the hell out of me for several sites. And it's starting to not block LLM bots anymore?

> 33% are now solving the math and getting through to the main site — because apparently what we have to offer is worth spending a ton of cycles to calculate the Anubis challenge.


Has anyone considered having Anubis perform more valuable hashing?

Like, maybe you can't stop the LLM bots, but you can use them as one-off Bitcoin pool mining pool participants. You have to assume that making them find hash values with N leading zeroes has led to finding hash values with more than N leading zeroes. Maybe run a Bitcoin node under there and let each visitor take a couple swings for you with their pickaxes.


Yeah that has been around for many years. Usually by sketchy download sites.

It's not really going to help though because the scrapers using residential proxies aren't burning their own compute.


Yes they are. The proxy is just a proxy. All processing is central.



In a few years the VC money will dry up and this gross overspend on slurping data will end.


Visions of vast data centers surrounded by fields of browning grass, in which aging, rusting, formerly extremely expensive hardware is spending billions of compute cycles looking at anime catgirls


Sounds like an even shittier version of the Lorax. :(


It's very likely the last few years of bot behavior is the consequence of the residential proxy business booming. This is indirectly due to AI company crawling, but the fact that they are as cheap and available as they are changes the incentives for anyone using them toward reckless and unsustainable request behavior, as there is no risk of burning your IPs, and very small chances of seeing any consequences of essentially DDoS:ing a website.


And the residential proxy business was created by Cloudflare, who was created by us using Cloudflare.

I've been on all three sides (user of RPs, getting paid to run an RP, and trying to block RPs from my site). Residential proxy service is nice. You can scrape anything, even with the dumbest curl command, and only get a Cloudflare block maybe 15% of the time, in which case you just try again. That's less often than I get a cloudflare block from using a privacy browser from a non-proxy address. Cloudflare does not stop bots, it stops humans.


Wait, so you’ve been the person who wanted to keep people from scraping your site, the person who’s trying to scrape your site, and the person getting paid to help someone scrape your site? Brother, what are you doing with your life?


Welcome to capitalism. Welcome to game theory. Welcome to competition. Welcome to the real world. Welcome to being a grown adult.


Contingency exists, I get that, and if you're truly in that state, sure, I don't judge necessity. A lot of our cohort seems to confuse a studied disinterest in looking beyond the end of their own nose for the sage wisdom of adulthood, though.


How will you know anything about a system you stubbornly only look at one side of? Like saying planes are terrible because they're loud. They are, but have you never traveled to a different city?


I’ve never mugged someone either, I suppose I should give it a shot before I go passing judgement.


Unfortunately we're apt to run into some kind of Jeavons Paradox where the hardware gets so much faster in those few years will be able to slurp massive amounts of data cheaply so the problem never really ends.


> High Anubis difficulty is annoying the hell out of me for several sites.

I just close the website if I see Anubis. Some have it set at reasonable difficulties (like 2)… others have it where I need to wait for like 30 seconds, I'm not wasting 30 seconds of my life for that.


the brutal truth is that if the website operator simply disabled Anubis, your page load would likely take more than 30s.

when a system was designed for 100 req/s and bots hit it with 5000 req/s, nobody entering that queue is having a good time. Anubis is the trade those operators make just to ensure your request gets serviced at all.

30s load time is already a sign that the Anubis approach is breaking down. if there's nothing else ready by the next order-of-magnitude increase in crawler load, those sites quite likely will just disappear from the public internet. hate Anubis all you want: for most of us, the realistic alternative is strictly worse.


Just look at another tab while you're waiting if you're that bothered.


Is there anything new in this article? Yes, experts use AI better than non-experts for tasks in their domain. See LLMs reward expertise [1] and Terrance Taos conversation with LLM [2].

[1] https://www.seangoedecke.com/llms-reward-expertise/

[2] https://chatgpt.com/share/6a5fdc7a-d6f8-83e8-bbea-8deb42cfed...


I think the need for expertise is also going away. For example, when Claude made progress on the Riemann conjecture,

> Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.

https://www.anthropic.com/research/riemann-zeta

The full transcript is here: https://www-cdn.anthropic.com/8a0d1add3c637b858a9a181e98c40e...

We're on the border of fully outsourcing expertise.


> Two mathematicians at Anthropic studied and validated Claude’s paper, and produced an informal note for experts stating Claude’s proof concisely.

No, experts are still needed


Yes, that's why we're on the border: AI is doing its thing without us, but we're not yet confident enough to let the AI do its thing without checking.

You may notice that humans only checked the work and explained. You may also see that that the person prompting the AI, Jared of bun.js fame, is not a noted expert in mathematics.


Misplaced modifier? Wow


> $200 plan (work pays)

so violating the TOS? Or work pays for a plan you cannot use at work?


What are the consequences? This likely doesn’t matter to most


What part of the tos is violated?


I have no idea what you think violates the TOS here?

I was doing work at work on a work task using a plan paid for by work.


The $200 max plan is for individuals. The individual plans are heavily subsidized. Employers should be using either the Team plan (which has much lower limits than max) or the Enterprise plan (which is entirely billed on usage).

Anthropic know that lots of people are doing all sorts of “bad” things like employers paying for Individual plans, (and using multiple accounts to get more usage) and aren’t yet enforcing the rules… but by the letter of the Anthropic terms, your employer should be paying Anthropic a whole lot more (and that’s one of the reasons why AI usage is going to get very very expensive as soon as the subsidies stop, you and a lot of other people are already paying a lot less than you should)


AFAIK there is nothing in the ToS that forbids an employee paying for the 20x, $200/month plan. I could be wrong about this in which case it'd be useful to have a link the clause.

I think multiple plans are against the ToS, but I'm not doing that.

The Teams plans are more convenient for a number of reasons, but yes, they top out at the 6x plan, not the 20x plan.

Edit: ToS are here https://www.anthropic.com/legal/consumer-terms and https://www.anthropic.com/legal/commercial-terms

I've re-read it and I'm pretty sure there is nothing that forbids a business paying for a 20x account. Notably they say this in the consumer ToS:

> If you use an email address owned by your employer or another organization, your Account may be linked to the organization's Anthropic enterprise account, and the organization’s administrator may be able to monitor and control the Account, including having access to Materials (defined below). We will provide notice to you before linking your Account to an organization's enterprise account.

which goes at least moderately close to indicating using it in a work environment is allowed.


i used to think the same as op but i think you are right. though they can change the eula at any time im sure....


> The individual plans are heavily subsidized

Not this again. There is not a single shred of evidence for this. In fact multiple times this year alone, people from Anthropic have said that inference and deployed models have positive margins. The big bucks are always being spent on training the next model.


It costs $10B+ a quarter just to train?

It's a simple fact that paying enterprise token rates would cost many times what the individual plans cost. It may well be that the enterprise income outweighs the cheaper tokens on the individual plans for net profit, but using the individual accounts as subsidy is a common tech industry tactic and what you said doesn't prove they're not doing it, either.

It's a little hard to believe they wouldn't be trying to slow the burn rate for an IPO if the inference were truly so profitable. And i find anything they say a little hard to believe all the time anyway (though admittedly they're a lot more trustworthy than OAI)


I agree the API prices are profitable (or at least break even) on an operational basis.

I don't think that means the subscriptions aren't subsidized though! I expect they are, at a carefully calibrated rate.

They want to maximize people trying them.

The real money is the enterprise plans which need a seat price plus API rates (unlike the Teams plans which are a pretty good deal).


you sound like fun at parties


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: