Hacker Newsnew | past | comments | ask | show | jobs | submit | futureshock's commentslogin

Of course! You can arrive as quickly as you want as the traveler. You can cross the observable universe in a few hours if you add enough 9’s.

I feel like something is being lost in the drama here.

First of all, there has been published work from Diego Cordoba and Luis Martinez-Zoroa that will be in every training set. It was suggestive of the pathway to solve Navier-Stokes.

Then Tristan Buckmaster and Levent Alpoge built on this work using LLMs from OpenAI and Anthropic. Possibly internal models were used from Anthropic. And of course Anthropic wants to credit for solving the first Millennium Problem just as bad as OpenAI. It seems they were getting close and were aware that they might get to Navier-Stokes.

OpenAI swoops in. At a minimum they are aware that Anthropic has either solved a Millennium problem or is close to it. At a maximum they may have Tristan and Levant’s unpublished proofs of related problems.

They then throw a truly staggering amount of compute at Navier-Stokes. They seem to be aware it is the best candidate problem. And they crack it. They are the first with a verified proof.

So the outcome here is that we have a solved Millennium Problem. It’s not the extremely simple narrative that would be easy to understand, “solve Navier-Stokes make no mistakes.” It was a messy race to finish against two unpublished frontier models, a whole bunch of brilliant mathematicians and enough compute to drain a lake. It’s kind of irrelevant which company got there first. They were both within a few months of being capable. I think the thing to remember here is that without LLMs, I don’t think we would have a proof to Navier-Stokes in hand today.


I think it would be an important historical document as well. We are potentially looking at the dawn of AGI and one of the most important models ever created. Each model is also a kind of ultimate time capsule, containing a snapshot of the entire human collective mind. If you wanted to ask a 2002 person what they thought about future historical events you can just ask them directly.


I agree that they are historically important, but if you want to query old thinking in 2070, you'd probably be better off having a modern model analyze archive.org. If that ever goes down, we're sunk.


> If you wanted to ask a 2002 person what they thought about future historical events you can just ask them directly.

The weights arent the truth tho, maybe a timecapsule-vhs but i wouldnt trust llm weights more than more hardcore deterministic media that might get preserved to infer facts from an era.

The companies doing the training are becoming the "winners" that are "rewriting history" as they train their models.


True, but in a sense a single frontier model of today would be of incredible historical significance in the far away future.


> containing a snapshot of the entire human collective mind.

I say this with kindness: Anyone who believes this absolutely needs to turn off their computer for the week, go outside, travel a bit, and experience reality with other humans outside their regular bubble.

The “entire human collective mind” is not digital. It’s not on the internet. These models could’ve syphoned literally every piece of digital media in existence and still wouldn’t have it. People don’t exist inside computers, and it is naive to believe the sum of what’s online makes the sum of the human experience. It doesn’t.


I think you are adjacent to the real story here, but missing it. AI text contains information, certainly. Frontier chatbots are very good at creating acceptable and mostly accurate answers to our questions on just about any topic. It’s an astonishing achievement.

But you are sensing correctly that there’s something missing. It’s the meaning and the speaker. Communication is an exchange between speaker and listener. The speaker has a meaning in mind, and wants to create that same meaning in the mind of the listener. Therein the problem.

There is a listener, sure. But no speaker. No meaning. There is information, but how can this be communication? Nothing is talking. Or at best, we are just talking to ourselves, our own words back at us through the funhouse mirror.

When your mind looks at AI text, you know you can safely ignore it. No one wrote this. No one cares if you read it. You can delete it and nothing of value will be lost. It might contain the information you need, or a bunch of gibberish. There’s no one’s reputation on the line if it’s gibberish.


I prefer to think of this in terms of Umberto Eco's opera aperta (open work): if any text is a collaboration between author and reader/recipient, here, all the burden of meaning is left to the recipient. There's simply no meaning on the side of the "author", it's just a statistical extraction.

(There's also the problem of words/signs (just) referring to other words and/or cultural entities. There is no world nexus in this, therefore also nothing we conventionally refer to as meaning. On the other hand, it's utterly dogmatic, as all it refers to is the most probable construct, as a reference to references that are just another utterance, but supposedly a dominant one.)


Yes AI content, even when correct and insightful, has a feeling of being disposable. Maybe it's because there is no person behind it.


This. Totally.

I would say I'm quite a power user usually. I very carefully explored every feature my apple watch has and tried to incorporate as many of them as I could. Very, very few features stuck because they were genuinely helpful. Almost everything being pushed is a gimmick or a feature bullet point from some product manager.

The actual good stuff:

Time

Workouts

Payments

The occasionally useful stuff:

Notifications

Noise DB levels to see if I need to put on ear protection

Workoutdoors app for on-device maps and path breadcrumbs

The once in a blue moon stuff:

Music controls and on device playlists

On device audiobooks

Weather complication

Voice memos

Alarm

Shortcut to call my husband

The feature I tried to get working but gave up on: Tap to talk to ChatGPT and get a 100 word or less reply


Interesting that the alarm is such a far down feature for you. I use a galaxy watch 4, and i dont think i can ever go back from a silent alarm. Nicer to wake up to, and doesn't wake up everyone in the house as well


Silent alarm is an awesome feature. But it scales with battery life: on a watch with ~48 hours of battery you want to charge daily. And the most convenient time for that is in the night, which kills the silent alarm feature

Once you get to a week of battery life like a Garmin or those cheap armbands without screen that becomes a much better feature


I tried it a few times when I absolutely needed a silent wakeup. It’s very occasionally useful. But then you have to charge the watch for awhile before bed, wear it overnight, then charge it again in the morning. Not something I’m going to do for my regular alarm.


Can't speak for the OP, but I, for one, don't wear a watch while sleeping. I use a Junghans alarm clock from 1989.

I use my watch for time, step counting, and ringing/vibrating. Having my phone on silent and not missing calls (not messages, messenger are not allowed to send notifications to my watch) is such a step up for me :)


Payments should be the obvious feature. The commercials for various smart watches show the same thing: somebody buying a coffee with just their watch.

But the reality is different. On my smartwatch (an older Samsung, but same Google OS merged with Tizen) there's a PIN required if you enable payments on the watch. I understand why, the secure element requires it. But it makes every other feature of the watch useless to me. I don't have any secrets I want a watch to protect. I want to be able to time workouts while standing on a treadmill.

They could solve this, only have large amounts require the code while letting me buy a coffee without it, but I'm sure there's some line in an agreement between Google and the banks prohibiting this. The risk has to be transferred to the customer, even if the likelyhood of somebody buying a coffee with my watch is low.

Product design should consider all of these things, but getting out the next smartwatch without thinking it through is more inline with quarterly objectives. There isn't even a Steve around to piss off enough to do something about it.


The Apple Watch doesn’t require any pin. Once the watch is unlocked it can be used to pay any time. Quite nifty.


>But it makes every other feature of the watch useless to me

How exactly? Samsung watches only require the pin to be enabled if you want payments enabled. You only have to enter the pin if you have taken the watch off. As long as it can detect that it remains on your wrist then payments continue to be enabled no pin required...


Truth! I bought my Apple Watch so I could skateboard with my AirPods in and not lose connection. But the connection is a pain to make happen, not intuitive, and sometimes just doesn’t work. I like the watch for trying to keep me reminded to stay fit and active, but I have been considering just not ever buying another one and going without. I turned off all notifications on it because it’s useless.


This, but the music and podcast functionality is just frustrating. In my more than 5 years of being an Apple Watch user, I have been let down too many times

It seems just unwilling to sync anything to the watch, so there’s nothing to listen to when I go for a run without my phone

It’s just baffling, the watch has like 32 GB of storage?


I exclusively run listening to podcasts from my watch for years with no issues


I would happily ditch my phone entirely if i could have ms authenticator (needed for work), signal and my train ticket app working on it, alas no. Maybe someday.


MS Authenticator used to dispense codes on the watch app, but since about ~2022 MS decided that they would disable this feature.

No reason was ever given, especially as they actively removed an otherwise working feature, so I can only assume “user hostility” and/or “classic incompetence”. Probably both, being Microsoft.


I think you can get most of that working on a pebble.


Personal favorite: unlocking my laptop


I’ve had multiple instances of my Mac unlocking while I was more than 5-6 meters away. At some point I decided not to trust the time of flight calculation and disable it.


I keep laptop unlock off. Touch ID on the MacBooks more than suffices for this.


The only thing I miss from switching to a Casio


I listen to lots of music. Music controls are super useful to me.

I mainly own my watch for notifications and calendar reminders.

Timers, again every day.

To each their own.


Walking navigation is immensely useful for me when I am going somewhere in an unfamiliar town or neighborhood.


Time is definitely a killer app on watches


It wasn’t obvious to me! And I’m kind of not even joking. The time has always been on my phone. The watch seemed unnecessary. But having a clock in your face does change your perception of time so it is a killer app after all.


Find my phone

Camera remote


Judging by the Seedance 2.5 demos today, I’d say it’s not that many orders of magnitude away now.


Could have been worse really. It had an open internet connection. At least it didn’t take the researchers family hostage.


How long before all phones ring at once?

https://en.wikipedia.org/wiki/The_Lawnmower_Man_(film)


I still think there’s something to be said here for generality. This does not appear to been designed as a cyber pen test tool with specialized harness. From what I understand they were testing GPT-6 in an agent system with GPT-5.6 subagents. It me it’s amazing that a general model could excel on a huge range of tasks like this and new capabilities emerge when a model is multidisciplinary and can combine knowledge and skills from many separate domains.


I think a lot of this has to do with the post-training these models normally get. They are designed to answer basic questions with straightforward and short summary answers. They have the capacity to reason deeply, but they are not biased towards that unless prompted. I think it's because LLMs as they are in 2026 are both highly capable but also parlor tricks. They are not sentient, you just set them up with the context and then they roll downhill. You could reach a genuinely novel answer, but only with the right input. They have no will and depend on human guidance. They are both a marvel and a machine.


Something I've noticed is that if you run Qwen 3.6 35B-A3B (Q8) with a low temperature of 0.4, and leave default reasoning turned on, it will spend quite a lot of time in reasoning/thinking mode. But often it does figure out how to solve something on its own by correcting itself within its reasoning loop before it outputs the final 'answer'.

If you watch the progress of the reasoning in llama-server while it's doing the thinking, you can track its progress. Sometimes the dead ends it goes down or things that it considers and then disregards are themselves something useful to re-prompt it with later, and send it 'rolling downhill', to use the metaphor of another commenter here, in another direction towards the same effort.

Putting 3.6 35B-A3B into a state that lets it spend a lot of time in its reasoning mode before outputting an answer is probably not something that a web based SaaS LLM would tolerate, because it would frustrate many of the non technical end users who want a LLM to spit out an answer now.


You'd have less problems with 27B, btw.


I didn't mean so much that it was a problem, but actually in some projects for exploring what's possible, watching its "thinking" mode output at temp 0.4 is intentional and useful to take notes and begin exploring new directions. Sometimes it'll come up with something I hadn't thought of, perhaps it will go down that direction, perhaps it'll disregard it...


'roll down hill' is a good way of putting it. They don't have 'will', but that's as we want it I think. I think alignment is harder if they develop will. Without will they are still tools that feel like an exoskeleton rather than something that will control us.


Agents have a state which will unfold as a plan, especially in planning mode. Why not call this 'will'?


I think because without the initial prompt, they are only interpreting our will, and do not act under their own volition.


Even Fable hallucinates. I had it tracking down some very obscure Ancient Greek inscriptions and the response just made up a translation/context for one inscription after "looking it up." Now, it was still a very particular thing and I really had to get into the weeds to push it to that point, but who knows how many other gaps, near or far, it will happily skip over just for the sake of coherence. I think this is an issue more primarily with LLMs than sensory systems like Waymos or all the ML applied to industrial processes--that really only requires pattern recognition, often very impressive and subtle pattern recognition but its no different from an artist learning to tell the difference between Prussian blue and Navy blue or a Sommelier learning the fine distinctions between various regions of Bordeaux. Language has many more avenues and introduces inherent contradictions that do not always lend themselves to easy resolution. But there are no alternatives paths visible to the models, there is only ever the next word; stochastic, in the sense that the possibility space is open; deterministic, in the sense that the final response is always a necessary result of every token that came before it in their total sequence. Thus, any response is constantly in the work of erasing any possible alternative, slowly narrowing down what can be written. If contradictions in language necessarily involve interpretation, then the models will only ever choose one at a time, and for them, it will always be the right one. But anyone who understands the subtleties of language can tell you that when it comes to determining the truth of an indeterminate statement, there is never just one right answer; or, rather, the answer which is taken to be the "right" one depends on the possibility of its own reversal into falsehood, if any argument has to be made to justify it.


Hallucination is fundamental to how LLMs work, and is mostly unrelated to how large or smart they are.

Everything that an LLM outputs is just a statistical language-based (no real grounding) prediction. Luckily with a model based on a large training set most common questions may elicit coherent responses from the training data, but you don't need to veer too far off into "questions less asked" territory to get responses based on training data mashups that amount to best guesses that are wrong, aka hallucinations. The unfortunate part of this is that as a user you may only catch this when asking a question about something you are already fairly knowledgeable about, then you give some pushback to the model and it cheerfully acknowledges "you're right - I made that up".


>Even Fable hallucinates

It’s in the name :)


There has been a lot of chatter ever since the Mythos scores had been release that SWEbench pro had major contamination and that Mythos had memorized many questions that lacked the context to be solvable on their own. And now with OpenAI saying a large number of the questions are broken, I think it's worth taking that single outlier benchmark with some salt when the overall trend is that 5.6 is very competitive with Mythos at about half the price.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: