Hacker Newsnew | past | comments | ask | show | jobs | submit | ummonk's commentslogin

It could go into personalization / memory. Or they could be A/B testing some system prompt tuning and consider the thumbs up / thumbs down as statistical feedback on the particular flags that are enabled for your account.

How would going leveraged on the S&P 500 because they expect long term rapid GDP growth help them with short term cash flow to invest in AI training that will enable that rapid GDP growth?

Because the market is forward-looking, and brings all future cashflows to the present.

If it became clear by 2027 that GDP growth would permanently reach 15% (I didn't make that absurdly stupid chart, they did), S&P 500 valuations would immediately 100X or more.

They could take out a credit line against those gains and have unlimited money.

Wait...are you suggesting those predictions might be so unrealistic that its stupid to even publish them? I'm shocked!


Are you saying that instead of investing in training right now they'll get higher ROI by pausing training for a year, and investing their money into taking out long positions on the stock market, all on the hope that what they label the "Extreme scenario" comes to pass?

If even their modest scenario becomes consensus by next year, the market will 2X or more. They could easily leverage that to 10X returns.

Their business is not going to 10X cash flows next year, so yes, that would be the smart move if they thought even the short term modest scenario was happening.

They don't think that though. Hence why this blog post is stupid to even publish.

It's like me publishing a "What will our future look like?" blog post where I think I'm the center of the universe and draw a chart that says:

Modest scenario: I take a bowel movement so big it opens a wormhole and destroys humanity


The other reply doesn't contradict what you're saying. Someone who was street racing is going to tell the cops / insurance company that they were lost in thought.

Especially because the data that gets fed into training is first anonymized, so they’d need to look for navier stokes related stuff in the anonymized training set and then get make some sort of ad hoc process (with Tristan’s permission and sign off from legal) to compare the training data against his chats / Codex sessions to check if anything matches up. And that assumes his chats / sessions are still there, and not deleted to compare against.

If the method is indeed found in the training set it's not particularly important to de-anonymize it. You have the proof you need.

This is an excellent point...

My best guess is that from the perspective of the OpenAI people, Buckmaster was letting his paranoia about OpenAI training tank his opportunity to receive the Clay Prize (it seems like Buckmaster and Alpoge's result isn't quite the full result required for the Clay Prize, whereas apparently OpenAI does have that full result worked out, using the same approach that Buckmaster and Alpoge had been exploring).

Whether the Codex sessions could have indeed made their way into Astra training data is something I can only speculate on though.


"using the same approach that Buckmaster and Alpoge had been exploring" is imo mealy wording: it seems fairly likely that OA heard Buckmaster and Alpoge were close to a breakthrough, and decided to use their unlimited compute to quickly prompt based on their assumptions about B&As work.

Is that necessarily wrong, so long as the original innovators get a citation credit?

Citation of what? this was unpublished work! This computational blitz really just reads as "might makes right" on OpenAi's part... which isn't surprising, but they should probably be honest about what they've done here.

Yes. Quite simply yes. It's unethical. And it's a dick move

Buckmaster is already an established world expert at these sorts of problems. Declining the Clay prize would hardly dent his career.

Zero data retention, wink.

No looksies, wink.

No trainsies, wink.


Well, Buckmaster says both his and Alpöge's use of Codex was non-institutional, and OpenAI claims the right to train their models on inputs and outputs of non-enterprise users in their service policies [0]. So I'm not sure they were even promised that.

[0] https://openai.com/policies/how-your-data-is-used-to-improve...


It isn't relevant whether they were promised that. Indeed I think the assumption must be that they were not promised that, since otherwise the author asking if they were would not make much sense.

If OpenAI did use the conversations from Buckmaster and Alpoge, then not disclosing it, explicitly, is plagiarism. If they planned to use that plagiarism to pressure the authors to publish, that is even more unethical. What the terms of use say does not make it any more or less ethical.


If you use their consumer subs, you get subsidized tokens in exchange for them having full access to your data. Those are the T&Cs. Have something secretive? Get a commercial sub with zdr.

(And if memory serves, there is also the opt out from training on consumer subscriptions). Its not plagiarism if you make your data available for the purpose of training their LLMs. It is you giving away your IP for some tokens.


That's absolutely right. Why the downvotes? If OpenAI are using but not acknowledging the work of others that's plagiarism. If they don't know for sure, but aren't performing due dilligence to make sure they aren't, that's also plagiarism.

It's not as simple. All our chats are being used by both labs for their future product (unless signed by ZDR). Where should the acknowledgement begin? Who should be acknowledged? The whole world? All the 2B users of AI?

If I know person A is working on problem B.

I am free to work on problem B too. Why should person A be limited to working on it.


Are you free to intercept person A's emails / hack their computer to find their notes on how they're approaching problem B?

Finding Codex session data in the training set that you tie back to these two researchers is like an hour-long task.

You're also free to plagiarise anyone you want. There are no laws against it on most jurisdictions.

Also brain raping* is not illegal in most jurisdictions.

But they're both deeply disturbing.

_________

* https://youtu.be/JlwwVuSUUfc?si=uWl4-LCHAeI7qtb3


I can imagine excuses for unknowing plagiarism in this case. What is described in the article seems much more serious: a research program that was only initiated following reports of the author's similar program. In this case no excuses of "I didn't know" can apply, it is not like this revealed some obscure work from the 1980s nobody could reasonably have foreseen. And as far as I can tell this program was only really initiated to apply pressure to the researchers, without their knowledge/consent. It looks very weird.

> Why the downvotes?

I think there was only ever one. Not sure why.


Your comment was greyed out when I saw it earlier, maybe you missed some downvotes?

About the plagiarism issue, I model it as OpenAI being an advisor and their AI a PhD student. If the advisor puts their name on a paper behind that of their PhD and it turns out the PhD copied the text of the paper from somewhere else the advisor is also responsible of plagiarism, not just the student. The least the advisor can do is withdraw their authorship from the paper.

But, yeah, point well made: it could be much worse than that. Like an advisor instructing a student to copy someone else's paper.


I think "greyed out" just means "0 points or less", so if you get 1 downvote without any upvotes it'll be greyed out. For instance your initial reply to me is now greyed out, and I have since observed a few upvotes and downvotes on my original comment (the downvotes apparently from people who aren't willing/able to justify why).

Personally I don't like thinking of LLMs like a PhD student, because most PhD students remember where they learned things from, while LLMs essentially cannot. I think of it a bit more like someone using a search tool carelessly. Although in this case it is apparently more like deliberate misuse than carelessness.


OpenAI is willing to credit so plagiarism is not the right framing here.

It's not clear to me that they would have credited the authors if they had not got in touch with OpenAI first.

Only for ChatGPT, if the user hasn’t opted out. Would mathematicians be using ChatGPT for this kind of work? Genuinely asking, I know nothing about this!

Yes, mathematicians are. And yes, most of my colleagues did not even know the opt-out was an option.

Question is what does that button do.

I bet a lot of lawyers are salivating at this question too.


If I am reading your question correctly you are asking about chat interface Vs Codex/Claude code? If so, in my experience Codex/Claude code use is widespread for mathematicians who are seriously using these tools.

Can you call paranoia a fear of something which is happening? OpenAI uses user chats for training and they are "open" about it.

He doesn't seem to be after the prize himself. In this statement he credits the approach of another two researchers:

> I believe Luis Mart´ınez-Zoroa deserves a Fields Medal.


No that would be an argument to return to the original definition of planets (highly visible celestial objects that migrate across the sky, i.e. the Sun, Moon, Mercury, Venus, Mars, Jupiter, and Saturn).

If you believe in the "standard 9" you're just insisting on returning to a scientific definition of planets but one that is frozen on our scientific understanding from a few decades ago...


The argument is to not make any changes to what people call planets but to instead have taken the opportunity in 2006 to create a separate classification system. Use something vaguely Latin sounding like planitus or whatever.

Yes, but again, what people call planets was stable for over a thousand years before changing rapidly in the past few centuries (first with the Sun and Moon being removed during the Copernican revolution, then Ceres, Pallas, Juno, and Vesta being added before being removed when they discovered a bunch of other asteroids, then Neptune being added, and finally in the last century Pluto being added before being removed after they discovered a bunch of other dwarf planets).

If you don't like changes to what people call planets, you should be unrolling all the other changes that have also been made, and go back to the original definition that stood from ~200 BCE-1700 CE.


As you can see, there are many valid ways to go about understanding what a planet is in a cultural sense.

> The fix was to create new exceptions so that str.lower() would behave as if it was using Unicode 3.2.0 for only particular function. So, we go through each Unicode codepoint and record when the behavior of str.lower() is different when comparing the Unicode version shipped with Python and Unicode 3.2.0

This sounds like a really hacky solution compared to implementing a separate frozen Unicode 3.2.0 lower.


The first sentence sounds as if they modified the implementemention of str.lower(). That would be bonkers, but that's not what they did.

https://github.com/python/cpython/commit/7e109d084d55e7eb

The important part is:

   # B.3 is mostly Python's .lower, except for a number
   # of special cases, e.g. considering canonical forms.
  +# To enforce Unicode 3.2.0 behavior of .lower instead of
  +# whatever Unicode version is included with Python we
  +# add unassigned or newly case-folding codepoints to
  +# the exception map, too.
   
   b3_exceptions = {}
   
   for k,v in table_b2.items():
       if list(map(ord, chr(k).lower())) != v:
           b3_exceptions[k] = "".join(map(chr,v))
  +for cp in range(0x110000):
  +    ch = chr(cp)
  +    # Assigned in current Unicode version
  +    # and supports case folding, but not
  +    # explicitly in B.2 or B.3 tables.
  +    if (unicodedata_current.category(ch) != "Cn"
  +            and ch.lower() != ch
  +            and cp not in table_b2
  +            and cp not in table_b3):
  +        b3_exceptions[cp] = ch  # Identity.


That fragment doesn't mean much in isolation. You've just said that they didn't modify str.lower (because "that would be bonkers") but you've posted a fragment which, for all we know, is part of the str.lower implementation.


To be clear (because the snippet is non-explanatory).

For encode("idna") what they did is use lower() except where it would produce a result different to 3.2.0 and then instead use the result from 3.2.0 instead.

Essentially they've frozen the IDNA encoding to be based on 3.2.0 by overriding any changes.


Yeah. I can understand the confusion, though. The title claims the issue was in lower(). Though the problem was actually in encode('idna')'s usage of lower().

The article would probably get far fewer clicks if it were named "when encode('idna') is a security vulnerability"


Yeah, that's what I figured, but my worry upon seeing this is "what happens if lower() changes again and people forget to update the list of exceptions?".

Unless they have unit testing on the entire Unicode code space to ensure what they're doing is always identical to 3.2.0.


I feel like there are a lot of use cases where I’d opt for SQLite and a lot of use cases where I’d opt for Postgres + PgBouncer. I’m curious what kinds of features push towards using Postgres alone over SQLite.


All of my personal apps use Postgres because:

- types are lovely. We love types. SQLite’s default of non-strict typing is, to me, bananas.

- SELECT DISTINCT ON is my ride-or-die

- most importantly, I’m very comfortable in Postgres and the setup cost is basically zero (like SQLite) because Claude does it.


The setup cost is basically zero anyway. Add apt repo, apt-get install the correct version. Easy to run different versions at the same time too. Upgrading is annoying.


Concurrency, data types, scalability, centralization, or replication push for Postgres. The push against PgBouncer is that you don't need it, unless you do. If you have an app-level connection pool, you probably don't need PgBouncer.


I thought those were just allegations they made in the lawsuit, not necessarily something they had written evidence of in the form of messages.


If you check page 16 of the filing here it certainly seems like they have specific messages with quotations: https://9to5mac.com/2026/07/10/apple-sues-openai-trade-secre...

This is one of those times where I wonder if I'm a genius for knowing not to generate evidence of a crime by talking about it in recorded media, or if we're just seeing the bottom 5% of wrongdoers.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: