Yeah, in small examples it's pretty neat: Define a "sorted" property, write a law that checks that the list is sorted and contains the same elements and someone else implements the sort and together with a proof for correctness. However larger software systems have an exponentially larger surface for reasonable and unreasonable edge cases. When I hit ctrl+s in my editor I expect that my cursor does not change colour, that the window does not minimize, that the program does not crash if there is no disk space left and so on. I don't see how this does not devolve into "negative space programming" where the user would have to anticipate and constrain every possible unwanted behavior of his software.
>> When I hit ctrl+s in my editor I expect that my cursor does not change colour, that the window does not minimize, that the program does not crash if there is no disk space left and so on.
In my experience, when that happens it's most likely because you drew the wrong boundaries. Iterating on the boundaries also becomes quite cheap when developing this way though, you should never expect to get them right the first time unless it's a very common problem you're solving _or_ you've done it before.
A lot of math is extremely specialized, to the extent that only a handful of other experts in some field have any experience with those mathematical ideas, with most of them not even yet present in the published literature. It's really not a stretch to claim that it's pretty dubious when the AI decides to use these highly specialized tools after it has trained on chat logs where these techniques were being discussed.
Also, if we just take "high-quality" input data, which these chats would certainly be classified as, then the models are more than large enough to memorize everything verbatim. Spitballing some numbers, research literature suggests that LLMs are optimally trained with around 20 training tokens per parameter (fairly confident on this figure), that a DNN parameter encodes around 4 bits of data (less confident here) and I found sources in the 1-4 bits of information per token range (least confident here). So, fairly conservatively I would estimate that a model has the capacity to fully memorize around 5% of its training data, presumably high-quality data is a lot less than that.
If there's one thing I'm absolutely confident in, it's that Sam Altman personally goes to great lengths ensuring that ethical standards are upheld at his company.
Presumably because this was something Levent did in his spare time and because it was not obvious that this work would eventually lead to a breakthrough.
> Why did Tristan use OpenAI's models when it should have been known was a potential outcome?
I'm sure in the past he had less cynical feelings about OpenAI and their penchant for academic fraud.
> I understand they wanted a normal math collaboration but presumably what Levent brought was his resources (as far as I can see Navier-Stokes is not his speciality)
I think you're not giving the guy enough credit in saying that his contribution came down to having an API key for Anthropic models.
> Normally these things are hashed out formally beforehand to avoid the sort of thing now happening.
How would that have helped? That agreement (which may well still exist) would not have involved OpenAI.
So what do you think his contribution was? His preprint record shows no research on fluids - and the statement says that the first LLM-generated proof Tristan received from Levent was 'the most horrendous I have ever read.' Levent is out for mathematical scalps whether it is in his field of expertise or not, and he has the resources to do it. And I am not saying he is not a very clever person, but the idea that you can bring yourself up to the forefront of research in PDEs, in particular NS, and contribute new ideas in less than a year is implausible.
I guarantee you most of the comments regarding this aren't real humans. The homepage is full of crap meant to distract from what OAI did here, the comments are full of OAI employees. Dead internet theory pushed to the max
reply