Interesting approach. I would like to see a benchmark like this without web or mobile app code. In fact just: systems code (operating systems, drivers, low level applications), libraries, native apps, firmware, robotics software, signal processing, embedded code, bare-metal, RTOS, hardware design languages, EDA tools, etc.
I can call a coworker right now and have a real time conversation with them without feeling like I’m talking to a frustrating machine. Most of them, anyway. But I guess that’s “moving the goalposts”.
> I can call a coworker right now and have a real time conversation with them without feeling like I’m talking to a frustrating machine.
These machines were designed with the express purpose of being condescendingly sycophant to the point they hallucinate just to state "you are absolutely right".
No wonder some people even find these chatbots to be wife material.
And it disagrees with me a lot, granted this is in the webchat which I don't use for programming but for general usage seems to work, it isn't sycophantic
I don't have those system prompts, but gemini pushes back on its own quite a bit. It'll tell me when I wrong, or when there are better options to consider. They won't be exhaustive, but good enough for 95% of my queries, so it's also my go-to web chat.
Random humans don't have a significant cross section of human knowledge available in real-time, although many like to pretend they do, especially in internet comments :P Being able to compete with the capabilities of a median human would be an absolutely world changing achievement.
What do you mean? Virtually all humans have access to the internet. That's literally "a significant cross section of human knowledge available in real-time".
Oh, the median human can't process that in realtime, you say? Looks like they can't compete with the capabilities of the frontier AI then.
Yep - I like to phrase it as "AI is better at most tasks than most people". AI will still be beat at experts at specific tasks, but in general I find it to be better than me at the areas where I have no expertise. It's the ultimate generalist.
Have you? It’s a dumb model that gets a lot wrong. Codex voice is IMO only good for when I cannot dictate into the good models. The delay introduced by letting it work with a good model kills it for me.
Do labs come back from disasters like GDM’s 3.5 pretrain? I am thinking of Meta’s Llama 4. Meta is just now starting to be taken seriously again but they are definitely not at the frontier. And when I say “come back” I mean have an Opus 4.5 moment, which was really mind blowing for me at the time. Fable was a similar leap, just not as big.
Unless the company is going under, why not? Let's say Google releases Gemini Pro 4 tomorrow, and it's better than Fable and Sol; lots of people would switch over to it.
AI models are almost completely interchangeable, so the best/cheapest/fastest whatever will always have a market.
I agree we’d switch to it. I guess what I’m doubting is if a company can recover from that sort of stumble in the first place.
And they might not want to either. They might think there’s more value somewhere else besides trying to get back to the absolute performance and capability frontier. Smaller models targeted to specific domains that large models would be too inefficient at no matter how large they get or how clever you are at distillation, for example.
At some point it won’t be realistic to ask a human to solve a technical problem any more. I’m not suggesting it will happen soon but I can at least imagine that day coming. Whereas before modern “AI”, that possibility wasn’t even on my radar.
We didn’t replicate the human brain. We built systems that can statistically approximate some of what the human brain might output in certain limited situations.
> We didn’t replicate the human brain. We built systems that can statistically approximate some of what the human brain might output in certain limited situations.
Which is equivalent to
"We didn't replicate the human brain. We partially replicated its functionality."
One of the reasons this analogy is unconvincing is that humans have compared themselves and their inner workings to "the current technology of the time" for millenia:
- ~3rd century BCE : The invention of hydraulic engineering (eg aqueducs) in the 3rd century BCE led to the popularity of a hydraulic model of human intelligence, the idea that the flow of different fluids in the body accounted for both physical and mental functioning.
- Pre-Socratic Greece : The ancient Greeks saw the mind as a chariot pulled by horses of reason and emotion.
- 1500s-1600s: Automata powered by springs and gears had been devised, Descartes suggested that cerebral hydraulic automata produced behavior by powering "animal spirits" through the nerves.
- 1700s–1800s: The mind worked like clockwork.
- Industrial revolution (1800s) : In the industrial revolution, the mind was understood as a steam engine.
- Late 1800s : Hermann von Helmholtz compared the mind's workings to telegraphy and hydraulics; the brain was likened to a telegraph network or a complex switchboard
- 1895 : Sigmund Freud in 1895 described a Project for a Scientific Psychology using a crude neural network model and borrowed concepts from thermodynamics, speaking of psychic energy, pressure, and discharge, essentially a hydraulic model of the psyche.
- 1930s–present : "our brain is a computer"
- and 2022-now: "We are LLMs!"
I can't wait to become a quantum chip, an NFT, and so on, as new things arrive. It doesn't make any of those models accurate. They're just the metaphor of the day.
I suppose you could argue that none of those had the actual engineers behind them attempting to replicate "intelligence" though. Just because random people made metaphors doesn't mean the engineers designing them had any illusion that it was nothing like the brain.
My personal opinion is that eventually we may just get advanced enough genetic engineering combined with brain / computer networks that the real A.I. will just be a brain like thing grown in a lab but more specialised. Or who knows - our own brains might have some quantum inner workings as well we are unaware of !
Counterpoint: Our engineers are currently trying to replicate "intelligence" and we think it's like the brain because today we think the brain is where "intelligence" reside, and today we strongly believe that "we" are our brains.
I would argue that the inventors of those past technologies were trying to replicate other things (movement, energy, pressure, pneuma, psyche, physis, whatever) that felt deeply human to them, and that in the zeitgeist of his time, Hero of Alexandria (1 BC) could also have said:
> I suppose you could argue that none of the stories of the past had the actual philosophers behind them attempting to replicate "pneuma" though.
Paracelse was trying to create a homonculus by using semen, manure, and blood in the 16th century or so.
A few decades or centuries from now, maybe engineers will try to replicate "consciousness" (though it's pretty clear that Anthropic is already trying) when it has become clear to 22nd century people that giving an object "intelligence" gets you no closer to recreating humanity than "movement through pneumatics" does.
If you do not think there is a difference between "your reflection in a mirror" and "you", it opens so many fascinating questions. I'm curious:
- Do you think a live video, shown on a phone screen, of you, is "you"?
- Do you think a still photograph of you is "you"?
- Do you think a set of bytes representing that photograph (or video) digitally is "you"?
- Do you think a compressed version of that photograph is "you"? Is there a limit to how much I can size down the image or compress it until it's no longer "you"?
- Do you think the base-10 number equivalent to that digitized picture is also "you"? Can I memorize "you" if I learn all the digits of that number? Can I write "you" on a piece of paper from memory? Is Pi a person?
- There is a very large number of reflecting surfaces in the world. How many of you are there?
- Does the "you" in the mirror persist if you walk off the frame and can no longer see yourself in the mirror? What happened to him? Does he live in a left-handed world? What happens if I shatter or paint over the mirror?
- If I draw you, is my drawing "you"? Does the accuracy of the drawing influence whether it is really "you" or not? If so, then does the accuracy/quality of the mirror influence whether it is "you" or not in the reflection? Are "you" fatter or slimmer, depending if the mirror is warped?
- If you're standing far from the mirror, but I'm close to it and I can see "you", why can I talk or signal to you and you don't respond?
Not only that! Does the decimal representation of π (which is infinite in length) contain all persons who ever existed, and will ever exist? Since π itself is a known reason, but its decimal representation is infinite, it means π cannot contain itself. So if it can contain every person that ever existed, but can't contain itself (which could conceivably contain everyone), then what does that even mean?
I do love the idea that Pi contains all of us. It means everytime you put on a wedding ring, time a pendulum, look at a rainbow, or land on a spherical planet, you can whip out your ruler and get access to every single human that ever lived or will live. What a concept!
Spot me after the next rain. I'll be in the color indigo, right above the pot of gold, waving back.
I think "reflect top to bottom" is intended to mean "swap top and button". A mirror reflects left, right, top and bottom perfectly.
It's front and back that it swaps.
Someone saying that a mirror swaps left and right is comparing it to a photograph, and only because we, as bipedal creatures, really prefer to orient images of other humans with heads up.
Someone saying that a mirror is swapped left and right is because they rotated themselves 180 degrees about the vertical axis to face the vertically aligned mirror. If they used a horizontal axis instead, they would have swapped top and bottom. And if the mirror is horizontally mounted on the floor, anything goes. You'd probably say it swaps up and down, which is front and back from the mirror's point of view.
reply