> but it is undoubtedly and objectively accelerating research.
Part of the point of the letter is that it is quite possible to act in a way that is a net negative to research. The most obvious case is when the companies violate ethical standards in research.
The subtler case, the one for maths in particular, is what happens when you fail to follow well-established patterns for making maths research productive. Tao himself spelled out how that can look in https://mathstodon.xyz/@tao/117207856734787448 (which notably came before any of the news on Navier–Stokes).
> The declaration literally opens by writing that AI companies (or really anybody) saying, “Hey, let’s see if this powerful reasoning engine can solve an open problem in mathematics” is “detrimental” to the “science” of mathematics. Full stop.
So let's just quickly agree that the actual quote is “However, the push by AI companies to solve mathematical problems as a benchmark is detrimental to the science of mathematics, and to the mathematical community.” And that in this, "as a benchmark" is load-bearing.
We wouldn't see nearly the same amount of contempt from researchers had OpenAI picked a research-friendly approach.
What they did: Hear a rumour about the problem being solved by other researchers, then rush to scoop them (unethical), then, when they actually go talk to them, they try to oust an author (also unethical), and when they finally decide to share their own work, do so in the least useful way possible.
What they could have done: Upon hearing the rumours, connect with the researcher in question and propose that they join efforts instead; set up a joint project to test if the machines are useful in any way, and if that's not appreciated, back down again. And instead of dumping only an undigested paper* and a Lean proof, do the digestion prior to publishing anything (as Buckmaster was in the process of doing). If their own lack of competences was keeping them from digesting it, then again, reach out to the researchers to understand if anyone would be willing to do so.
In the second of those two worlds, we wouldn't be seeing nearly the amount of outrage that we are seeing right now.
*: Here, “digestion” is the process of turning an AI slop paper into something humans can read. LLMs can indeed sometimes (if much more rarely than marketing material from the large LLM companies will suggest) produce correct proofs, but they are often written in bizarre ways – they'll use lingo that doesn't exist, seem overly pretentious, dwell on extremely easy steps while glossing over the hard ones. Currently, a real researcher will take that output and transform it into something that others can understand, use, and build upon. This is not so different from what happens when using it to write software, although as someone who does both, I will say that the amount of digestion needed for proofs tends to be orders of magnitudes larger than for code. This meme is quite accurate: https://mathstodon.xyz/@tao/117068266071803252
There are several cases of this already. Bubeck himself had to retract earlier claims of novelty, and more recently, the provenance of the non-sofic group result was brought into question. Most recently, it turn out that the construction used for Anthropic's counterexample to the Jacobian Conjecture had appeared in an unpublished but publically available draft: https://news.ycombinator.com/item?id=49657499
This has a few practical implications: First of all, if you are in the target group of the marketing material, be wary. While these things can do non-trivial stuff, the amount of magic is being grossly over-stated. But also, when several of the big results have indeed been reappropriating the work of others; when the companies fail to provide proper attribution (the NS case in particular is laughable) and present the results as the models' own work, that's plagiarism.
And just to spell it out, since it looks like HackerNews is flooded by people who are new to science these days: even if a result doesn't come with a price, scholarly peer review is the norm across all of science: https://en.wikipedia.org/wiki/Scholarly_peer_review
And chances are they never will publish it in any kind of useful format. Right now, the scientific community is outraged at OpenAI for going about their announcement in the least productive fashion they could have. It really does seem like they have no interest in progressing our understanding of maths outside of mining it for marketing material.
Boo hoo. OpenAI got the result only days ago. It makes perfect sense for them to take the win in marketing, and it's fine if they take a few months putting together the paper and present it more productively later. The scientific community didn't get the result themselves, so it isn't theirs to be bossing everyone else around about.
I don't care much for AI myself, or smart phones either, for that matter. I would be content if NS remained a mystery for another 100 years - or forever. But goodness, does the "scientific community" need to take a deep breath and count down from 10.
Was it a marketing win though? My takeaway is: if you're doing groundbreaking work with openAI's models and they find out, at best they'll outspend you and scoop you. At worst they'll steal your chat history.
That's a completely different matter. And does it matter to my point if they were successful or not? That's ex-post analysis. It seems clear that ex-ante, they wanted this to be a marketing win. The original poster complained that they wanted a marketing win.
My point is: why shouldn't they want a marketing win from this. What obligation does a non-academic institution have to follow the traditions of academia? Its result doesn't belong to academia. And if academia wants to subject OpenAI to their own internal processes and give them marching orders, it just isn't going to work and maybe - who knows - it'll even further erode their own legitimacy. Does anyone actually believe that NS would have been resolved in the 2020's if we lived in a parallel world where LLM's were never invented? Would Buckmaster have gotten as far as he did without LLM's doing a lot of the work for him? We can complain about AI companies contributing to mathematics, but are we complaining about Terence Tao using AI in his research? When Tao publishes something are we all going to go to war against him because maybe other mathematicians' prompts went into training the AI that Tao used?
It's the same on Reddit, and if you look up a few comment histories, you'll find that it's mostly /r/singularity users and /r/accelerate posters that are spamming. So I assume that people are just deliberately misreading the letter and trolling here as well, and I wouldn't read too much into it, but on the other hand, if you never had to think about what scientific misconduct looks like, it's harder to see why the behavior of the AI companies is problematic.
> To be honest, I feel like the difficulty of reading AI proofs is due to the fact that we are on the verge of being beyond human comprehension.
I can see where that's coming from, but I really don't think it's the case. Even with Astra, the proofs you get are just off in a way that doesn't signal superhuman comprehension. As 9question1 says, a common theme is that they dwell on insignificant steps. Another one is that they'll often be full of terminology that either doesn't exist, or has this weird quality where it looks like it is trying to make some minor insight seem much greater than it is. At first glance, that'll often make it look like it knows more than you, but when it's really just doing the same thing but in a more complicated and worse fashion, that to me isn't a signal of comprehension at all. The bizarre thing is that despite all the "stochastic parrot" style nonsense you'll get in individual proof steps, they still often combine to something valid.
In either case, what all of this means is that the working mathematician still needs to go through, and generally completely rewrite, any proof output by an LLM. Otherwise you are passing the burden of unreadability onto the reader.
Yeah, that mirrors what I've seen throwing some of the leading models at a set-theory problem that's stumped me (https://mathoverflow.net/q/511601): in this case, the problem does not easily yield to the standard tools, but the LLMs do not recognize it as a major open problem they should give up on. So they seriously try it, but typically end up in a loop of inventing certain classes of simple solution or counterexample attempts, defeating them, and trumpeting each one as a major result, each time inventing some new terminology.
It's definitely quite curious that the AI labs are able to push these results through seemingly with pure brute force. Perhaps it's largely a function of how many monkeys you have attempting various constructions on top of the known results and strategies the models have memorized.
If frontier labs had chosen to go the path of offering to assist in existing endeavors, helping to build knowledge alongside researchers in ongoing projects and following ethical and professional research standards, we wouldn't be having this discussion at all; everyone would be stoked. Instead we have companies that disgracefully try to scoop researchers and fail to properly attribute earlier work and instead rebrand it as their own (what we normally call plagiarism) to make marketing material.
reply