Hacker Newsnew | past | comments | ask | show | jobs | submit | tracyhenry's commentslogin

Generative UI is unsolved because current models do not have taste and end up generating the same kind of slop.

I'm amazed that this blog doesn't even have a single screenshot/photo of the kind of UI they can generate.

Focusing on benchmarks in this domain feels very wrong.


Gen UI is meant to be design agnostic, the output is just the content and the form. It is on the implementation, agentic or human to make it look good.

design is the hard problem. i don't know what hard problem this is trying to solve here

There's an image in the article and a full website with more media is just 1 click away. Instead you resorted to typing 272 characters not including ENTER, and I doubt that was easier than clicking the logo to visit the homepage.

Typing this comment also did not solve your problem, because that would require the author to read your comment, add more screenshots and it would require that you revisit it.


When the blog title is "world's first model for Generative UI", you are supposed to show something that GPT/Claude couldn't do. Instead it's all benchmark.

I don't know what media you are talking about. It's all slop worse than current slop.


I usually don't reply to people who comment this obnoxiously, but the article is literally:

Headline

Introduction

Animation <------ LOOK HERE

Rest of article


No one on HN seems to care about its crazy ability in 3D modeling?

3D might have just been solved like coding


I care. I've been feeling bad for all the people who have dedicated their lives to that craft. It took years of daily practise to be good at 3D modelling...

When Astra becomes wildly available it seems likely a single 3d modeller (or even someone with no 3d modelling knowledge at all) could do the work of 10 today. Those guys are going to find it very hard to find work.


The only times I found the memory feature useful are in "projects" I created myself.

In a project my questions are usually revolved around the same topic. Having context carried across threads actually make a lot of sense.

In the general mode where I'm expecting models to be *stateless*, having memory is very annoying.


Not regularly, but I do watch movies on it once in a while (in those beautiful environments), especially when I'm on a flight.

WWDC also just rolled out some quite exciting features to RealityKit: https://developer.apple.com/videos/play/wwdc2026/279/ as well as visionOS itself (https://developer.apple.com/videos/play/wwdc2026/287/)

It makes me wonder if Apple is really giving it up as news have claimed.


> It makes me wonder if Apple is really giving it up as news have claimed.

I don't think anyone serious has claimed that.


Vision Pro's top talent all moved with their boss, Mike Rockwell, to Siri where they have at least a prayer of promotions, not to mention the satisfaction of working on a product that's not a miserable failure in the market.

Apple shelved the follow up v2 over a year ago and signaled very clearly it was refocusing its XR efforts toward AR glasses.

There hasn't been a new Vision Pro component order since the late 2023, early 2024 original orders that were capped (by SSS's display capacity) at 500k units worth, so Apple hasn't sold even that tiny amount yet.

It was Tim Cook's baby, his last shot at a product legacy, and the guy that just inherited Cooks role put even the lesser Vision Air on ice, the last goggles product that was still on Apple's roadmap.

So, it's not dead, but it's clearly on life support with no goggles of any kind remaining on the roadmap and Apple's attention fully moved to glasses and AI.


I don’t think Vision Pro was intended as a mass consumer product. There is so much cost reduction they could have done on the device that they didn’t do to hit a lower price point, and it’s obvious to everyone that $3500 wasn’t going to sell a lot of units (and it must have been obvious to Apple execs as well).

My guess is it is basically a real world survey of how people would use such a device, and what developers would do with it. Then they could later focus on a cheaper product that is only a display, or a product that is only a media viewer, or a product that focuses on VR/AR applications while deemphasizing other use cases.

This is the version of reality that doesn’t require anyone to be stupid. In my experience if a reality candidate requires that someone is just really really stupid to arrive at the outcome (in this case Apple execs), probably you are missing something about someone’s perspective.


This is why I haven't bought one again. I had an M2 and sold it. Then I was thinking I missed it and bought a used Galaxy XR. It is not bad for watching movies. It is close to the same experience for video consumption but significantly worse for everything else. I am always itching to get an M5 AVP, but the org changes are red flags. I will end up with a $4k paper weight.



Your ego must be insane to wear a Vision Pro on a flight.


What are the logistics of wearing this on a flight?

It seems like it would take up a majority of your personal item space. As someone who only packs a personal item for most flights, that’s a hard sell.

Can you reliably plug it in? It seems like the battery doesn’t last long enough for a long haul flight, and for shorter flights where the battery would hold up, the bulk doesn’t seem worth it. Of course, on the long flights, I’d hope to sleep, and having a cumbersome VR headset seems like it would be more trouble than it’s worth vs just watching the screen in the seat back.

When you get to your destination, do you use it there, or does it just sit around like a neck pillow?


At the end of the day, it only makes sense if watching a movie on AVP gives you immense joy, which is the case for me.

If you can charge your phone, you can charge your vision pro. The battery itself can last any movie with no problem.


After all these, I still feel their voice AI interrupts quite a lot, especially when I pause just for 0.5 sec. Interestingly, when I tell it to interrupt less, it seems to be better.


off topic but I just wonder if this page is AI-designed. It looks quite good to my eyes. I feel like prior to coding agents this would instead be a blog post with some charts.


No doubt, these sites are now a dime a dozen. Flashy but really low signal to noise ratio, you can't unsee it.


I'm building Eima (https://eima.app) which combines your Todo List and Calendar, allowing you to schedule your todos by simply dragging them onto the calendar.

Similar apps have existed before (like Amie), but they were nearly all VC-backed and had pretty much all pivoted to AI (e.g. being an AI note taker). Their approaches to a Todo-focused calendar has been largely unsatisfying due to the focus on Enterprise users and whatever is trendy.

Eima, in contrast, focuses on personal use and does one thing very well: scheduling your todos. In particular, I spent a lot of time making sure multi-occurrence todos work smoothly (e.g. todos that need multiple attempts or simply recurring todos). These were not addressed by prior tools at all and had been my biggest motivation to build Eima.

Would love some test users! If you end up wanting to give Eima a try please use the code EARLYEIMA to get it for free.


> they break at large enterprise repos.

I don't know where you get this. you should ask folks at Meta. They are probably the biggest and happiest users of CC


You mean the company where engineers ask chat bots to write chess games in their spare time in order to hit their AI usage requirements? That Meta?


idk why you bring this up. this is irrelevant to whether CC actually works at big corps


I missed that, source?


I'm honestly surprised by so much hate. IMHO it's more important to look at 1) the progress we've made + what this can potentially do in 5 years and 2) how much it's already helping people write code than dismissing it based on its current state.


This looks great. I would love to know more what makes Confident AI/DeepEval special compared to tons of other LLM Eval tools out there.


Thanks and great question! There's a ton of eval tools out there but there are only a few that actually focuses on evals. The quality of LLM evaluation depends on the quality of dataset and the quality of metrics, and so tools that are more focused on the platform side of things (observability/tracing) tend to fall short on the ability to do accurate and reliable benchmarking. What tends to happen for those tools are users use them for one-off debugging, but when errors only happen 1% of the time, there is no capability for regression testing.

Since we own the metrics and the algorithms that we've spent the last year iterating on with our users, we balance between giving engineers the ability to customize our metric algorithms and evaluation techniques, while offering the ability for them to bring it to the cloud for their organization when they're ready.

This brings me to the tools that does have their own metrics and evals. Including us, there's only 3 companies out there that does this to a good extent (excuse me for this one), and we're the only one with a self-served platform such that any open-source user can get the benefit of Confident AI as well.

That's not all the difference, because if you were to compare DeepEval's metrics on more nuance details (which I think is very important), we provide the most customizable metrics out there. This includes researched-backed SOTA LLM-as-a-judge G-Eval for any criteria, and the recently released DAG metric that is a decision-based that is virtually deterministic despite being LLM-evaluated. This means as user's use cases get more and more specific, they can stick with our metrics and benefit from DeepEval's ecosystem as well (metric caching, cost tracking, parallelization, integrated with Pytest for CI/CD, Confident AI, etc)

There's so much more, such as generating synthetic data to get started with testing even if you don't have a prepared test set, red-teaming for safety testing (so not just testing for functionality), but I'm going to stop here for now.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: