Hacker Newsnew | past | comments | ask | show | jobs | submit | thisgoodlife's commentslogin

I’m curious how you guys keep track of each model’s coding capabilities. The landscape keeps changing. I don’t suppose you benchmark all frontier models every other month, right?

I use them. Daily. Gemini hasn’t been a contender by comparison for a long time.

I wonder if it's a harness thing or a model thing at this point. I feel all coding models are quite capable for most tasks I want them to do.

Most of the time I don't need what the bench tests and I'm not really giving them completely ambiguous tasks without any refinement.

I only find marginal differences between models at this point and it almost feels like personality quirks in each model than anything.


When comparing OpenAI and Claude thats pretty much true, but not Gemini... And have you tried Antigravity? Yikes

The CLI version of agy is great. Have you tried it?

Do you dangerously allow permissions? I absolutely cannot use it until they ship an auto approver. As it is now I have it write one bash/python script to do everything it wants to, then I review that. Otherwise it is COMPLETELY unusable and it shocks me when I hear people are using it.

Sounds like they shipped some changes today that might reduce approvals: https://x.com/antigravity/status/2100001904969297980

  alias agy="agy --dangerously-skip-permissions"

Why not allow everything? It's not like it can do much inside the container.

Yes, but I dangerously allow claude and codex, too...

is anti gravity open sourced just like codex or grok code?

Compared to gemini-cli that they took out behind the woodshed, I hate it.

I've used Antigravity as my main coding agent on one of my biggest projects for about a year. It's been great for me. (and I use Claude, Codex, Grok and Muse for all the other projects)

True but it has a niche in SQL reviews for me. Looks like Google has a lot of good sql in their corpus and in their RL digital lobotomy factory.

I did a test involving implementing cobol control flow in Java for a source to source translation project. Gemini was the only model to get the edge cases. Cobol is very peculiar in this regard.

It's very good at Elixir in my experience too. And it just does what I ask and doesn't wind me up like Opus. I don't think I've had to insult it more than once per day.

I had a typical $20 Gemini plan that I just downgraded to their $5 plan (to keep access to some of the models). It had been so long since I let Gemini work on (or review) any code / design / html (anything) that I couldn't justify bothering to keep wasting money on it. It fell behind badly over the past year. Astra might as well be an alien super intelligence at code compared to Gemini. I enjoy talking to Gemini, it is very good at conversation, I get solid answers to everyday questions. I intend to keep the $5 plan indefinitely for basic use. I don't expect they'll ever resurface as a competitor in coding with Astra & Fable et al.

Even for battle-tested software like Postgresql, always read the changelog and test it thoroughly before upgrading.


I like the size of a regular iPhone. What I really want is a lighter phone. Unfortunately, compared to the iPhone 17, the Air is about 30% thinner, with worse battery life, camera, etc, but only around 7% lighter. I was expecting at least 20% lighter if it's called "Air".


Also combine fuchsia while you are at it. You don’t need so many operating systems


Can it draw basic stuff like circles, arrows, and lines now?


Circle: Ellipse select -> Bucket fill

Line: Rectangle select -> Bucket fill -> Transform

There should be some basic arrows but there doesn't seem to be.


You also need to check "Fixed aspect ratio" in the Ellipse select options, otherwise you will get an ellipse instead of a circle.


Or you can use the right modifier key.

Haven't used Gimp in a long time, but try holding down ctrl, shift or alt while doing the circle (or rectangular for that matter) select.

One is for aspect ratio, one is for center from starting point and the last I cannot remember.

IIRC the same goes for combining selections: modifier keys can be used to add, subtract or make intersections.


I just use shift, but that's photoshop muscle memory for ya.


I tired that originally and it did not work. The issue was that I pressed shift before clicking. After testing, MS paint and Photopea both support clicking shift before starting the selection to create a circle. Krita does nothing if you hold shift before selecting. I don't have PS right now so I don't know what that does. Seems unintuitive to me.


I guess it’s the latter. If you can afford, give your users a generous offer, but never unlimited. Otherwise, some people will find very creative ways to abuse it.


Yep, full time streamers run up a lot of hours.

https://bsky.app/profile/authorblu.es/post/3likxmdytys2l

Assuming ~6000kbit/sec that's about 17TB of archived video for that guy alone.


That's assuming none of that video is something that Twitch is storing for any other reason (i.e., other users have highlights of the same thing, or they would store the videos internally for some reason).

It's possible the actual additional storage requirements for that specific user are minuscule, since we don't know what data they are/aren't archiving themselves, if they're doing any deduplicating, etc.


yeah, kinda, but VODs (the automatic recordings) are not covered by this change. This is about edits & uploads, so stuff you would usually put on youtube. If you're a full time streamer and stream every day, Twitch will still provide your past streams for 2 (or 3? not sure) months (or less if you're not popular) and this will not change anything for you.


That's just $500 worth of storage though (and your 6Mbps is likely bit high IMHO).

Billing the owner a few bucks each month the each thousand hours of extra storage would make much more sense than removing everything.


> just $500

That's the cost for just buying disks, but storing data in the cloud costs more than that and it's an ongoing cost.

S3 charges 1.25c/GB/month for this sort of data. So that's $200/month for just this guy. There may be 100s or thousands of these people. Easily adds up.


> That's the cost for just buying disks, but storing data in the cloud costs more than that and it's an ongoing cost.

> S3 charges 1.25c/GB/month for this sort of data.

It doesn't cost them anywhere close that. Their competitors charge twice as less or more an still make money.

Twitch belongs to Amazon, they are the cloud.

Setting up your own infra to handle this is of course going to cost you a lot more than that, but when you have the infra set up then the marginal price is hardware (+ a monthly electricity bill, which is not as high as for other kind of workload).

And even if they had to charge $200 a month, they should probably offer the option instead of just removing the content: we're talking about professionals who make money out of the platform (and earn Twitch their income), they can make the choice whether or not they can afford it.


> And even if they had to charge $200 a month, they should probably offer the option instead of just removing the content: we're talking about professionals who make money out of the platform

There's no way these professionals have 6000 hours of interesting content and there's no way they would pay $200/month to store it. They're just saving everything they ever record because it's free.

Implementing that feature would cost more money than it would ever make.


> There's no way these professionals have 6000 hours of interesting content and there's no way they would pay $200/month to store it. They're just saving everything they ever record because it's free.

Some of these people have been streaming for 15 years, it's far from “everything they ever record” (and some content creators in twitter/bluesky links elsewhere in this discussion explicitly said they did select content).

Likely they would be more picky in their selection if they had to pay, but that doesn't mean they would be ready to pay something for a thousand hours instead of 100. 100 hours is a ridiculous amount!

> Implementing that feature would cost more money than it would ever make.

It's no more work than implementing a hard threshold. They did change the system in the first place, they could have made this change much better had they cared…


I don't think you've thought this through. You can't _just_ bill the owner a couple of bucks each month. You need a whole infrastructure to do that. You need to plan, design, build, test, deploy, maintain, and provide customer service for an entire new feature of your site. You need to research, test, revise and communicate what the price for storage is going to be (and handle the immediate and ongoing backlash). You need to catrgorize and plan for this new income stream AS WELL AS the costs to get it started and the ongoing costs to maintain it.

That's all just off the top of my head, and all of that is going to be fighting against all the other projects that people want to get done, projects that are likely way more profitable and way closer to the primary goal of the company -- being an intentional streaming service, not an accidental video hosting service.


That’s just looking at theoretical costs but completely ignores the actual revenue side.

If they annoy the most active streamers to the point they leave to another site, why should a viewer stay at Twitch versus just using another site?

I’m assuming some of these accounts bring in far more than the $500-1000 it costs to host old video.

Going from an unknown limit down to 100 hours with little notice shows how shortsighted Twitch was here.


Doesn't the infrastructure already exist?


Pardon, not the storage infrastructure, but the tracking, billing, taxation, customer support, etc. infrastructure.

It's a whole new income stream, which becomes a whole new line of business, and that business requires a variety of infrastructure to support it, especially at a large company.


Of course it does, these dudes already have this amount of video stored on Twitch's server.


It’s Twitch not some indie startup.

You’re not entirely wrong but you’re exaggerating the difficulty.


> and your 6Mbps is likely bit high IMHO

6Mbps is Twitches recommended ingest bitrate, and their highest quality just serves the ingested stream back to viewers without transcoding. In reality the storage would actually be a little higher still because they have to store all the transcoded lower resolution versions as well.


> 6Mbps is Twitches recommended ingest bitrate, and their highest quality just serves the ingested stream back to viewers without transcoding.

Interesting, I wouldn't have guessed that it would make sense for them with regard to bandwidth cost. TIL, thanks.


It's a trade-off between bandwidth and encoding capacity. Twitch actually only guarentees transcoding for "partnered" streamers above a certain viewership threshold, so when watching a smaller streamer you might only be able to view the "source" quality if there isn't enough encoding capacity to go around.


That makes sense.


That’s kind of the main consideration with production LLM apps right now. Really looking for a startup that solves this out of the box (llm credit payment system that manages the reality that remote LLM usage can never be unlimited).

Twitch will offer a premium sub for heavy users most likely.


I create a .venv directory for each project(even for those test projects named pytest, djangotest). And each project has its own requirements file. Personally, Python packaging has never been a problem.


What do you do when you accidentally run pip install -r requirements.txt with the wrong .venv activated?

If your answer is "delete the venv and recreate it", what do you do when your code now has a bunch of errors it didn't have before?

If your answer is "ignore it", what do you do when you try to run the project on a new system and find half the imports are missing?

None of these problems are insurmountable of course. But they're niggling irritations. And of course they become a lot harder when you try to work with someone else's project, or come back to a project from a couple of years ago and find it doesn't work.


>What do you do when you accidentally run pip install -r requirements.txt with the wrong .venv activated?

As someone with a similar approach (not using requirements.txt, but using all the basic tools and not using any kind of workflow tool or sophisticated package manager), I don't understand the question. I just have a workflow where this isn't feasible.

Why would the wrong venv be activated?

I activate a venv according to the project I'm currently working on. If the venv for my current code isn't active, it's because nothing is active. And I use my one global Pip through a wrapper, which (politely and tersely) bonks me if I don't have a virtual environment active. (Other users could rely on the distro bonking them, assuming Python>=3.11. But my global Pip is actually the Pipx-vendored one, so I protect myself from installing into its environment.)

You might as well be asking Poetry or uv users: "what do you do when you 'accidentally' manually copy another project's pyproject.toml over the current one and then try to update?" I'm pretty sure they won't be able to protect you from that.

>If your answer is "delete the venv and recreate it", what do you do when your code now has a bunch of errors it didn't have before?

If it did somehow happen, that would be the approach - but the code simply wouldn't have those errors. Because that venv has its own up-to-date listing of requirements; so when I recreated the venv, it would naturally just contain what it needs to. If the listing were somehow out of date, I would have to fix that anyway, and this would be a prompt to do so. Do tools like Poetry and uv scan my source code and somehow figure out what dependencies (and versions) I need? If not, I'm not any further behind here.

>And of course they become a lot harder when you try to work with someone else's project, or come back to a project from a couple of years ago and find it doesn't work.

I spent this morning exploring ways to install Pip 0.2 in a Python 2.7 virtual environment, "cleanly" (i.e. without directly editing/moving/copying stuff) starting from scratch with system Python 3.12. (It can't be done directly, for a variety of reasons; the simplest approach is to let a specific version of `virtualenv` make the environment with an "up-to-date" 20.3.4 Pip bootstrap, and then have that Pip downgrade itself.)

I can deal with someone else's (or past me's) requirements.txt being a little wonky.


> Why would the wrong venv be activated?

Because when you activate a venv in a given terminal window it stays active until you deliberately deactivate it, and one terminal and one venv looks much like another.

> I activate a venv according to the project I'm currently working on.

So just manual discipline? It works (most of the time), but in my experience there's a "discipline budget"; every little niggle you have to worry about manually saps your ability to think about the actual business problem.

> "what do you do when you 'accidentally' manually copy another project's pyproject.toml over the current one and then try to update?" I'm pretty sure they won't be able to protect you from that.

Copying pyproject.toml is a lot less routine than changing directories in a terminal window. But if I did that I'd just git checkout/revert to the original version.

> the code simply wouldn't have those errors. Because that venv has its own up-to-date listing of requirements; so when I recreated the venv, it would naturally just contain what it needs to.

So how do you ensure that? pip dependency resolution is nondeterministic, dependency versions aren't locked by default and even if you lock the versions of your immediate dependencies, the versions of your transitive dependencies are still unlocked.

> If the listing were somehow out of date, I would have to fix that anyway, and this would be a prompt to do so.

Flagging up outdated dependencies can be helpful, but getting forced to update while you're in the middle of working on a feature (or maybe even working on a different project) is rather less so. Especially since you don't know what you're updating - the old versions were in the venv you just clobbered and then deleted, so you don't know which dependency is causing the error and you've got no way to bisect versions to find out when a change happened.

> Do tools like Poetry and uv scan my source code and somehow figure out what dependencies (and versions) I need? If not, I'm not any further behind here.

uv has deterministic dependency resolution with a lock file that, crucially, it uses by default without you needing to do anything. So if you wiped out your cache or something (or even switched to a new computer) you get the same dependency versions you had before. There's no venv to clobber in the first place because you're not activating environments and installing dependencies - when you "uv run myproject" the dependencies you listed in pyproject.toml, there's no intermediate non-version-controlled thing to get out of sync and cause confusion. (I mean, maybe there is a virtualenv somewhere, but if so it's transparent to me as a user)

> I spent this morning exploring ways to install Pip 0.2 in a Python 2.7 virtual environment, "cleanly" (i.e. without directly editing/moving/copying stuff) starting from scratch with system Python 3.12. (It can't be done directly, for a variety of reasons; the simplest approach is to let a specific version of `virtualenv` make the environment with an "up-to-date" 20.3.4 Pip bootstrap, and then have that Pip downgrade itself.)

Putting pip inside Python was dumb and is another pitfall uv avoids/fixes.


>So just manual discipline? It works (most of the time), but in my experience there's a "discipline budget"; every little niggle you have to worry about manually saps your ability to think about the actual business problem.

>...but getting forced to update while you're in the middle of working on a feature...

I feel like trying to work on more than one project in the same session would require more such discipline.

>So how do you ensure that? pip dependency resolution is nondeterministic, dependency versions aren't locked by default and even if you lock the versions of your immediate dependencies, the versions of your transitive dependencies are still unlocked.

Ah, so this is really about lock files. I primarily develop libraries; if something breaks this way, I want to find out about it as soon as possible, so that I can advertise correct dependency ranges to my downstream.

The requirements.txt approach does, of course, allow you to list transitive dependencies explicitly, and pin everything. It's not a proper lock file (in the sense that it says nothing about supply chains, hashes etc.) but it does mean you get predictable versions of everything from PyPI (assuming your platform doesn't somehow change).

If I needed proper lock files, then I would take an approach that involves them, yes. Fortunately, it looks like I'd be able to take advantage of the PEP 751 standard if and when I need that.

>Putting pip inside Python was dumb and is another pitfall uv avoids/fixes.

Agreed completely! (Of course I was only using a venv so that I could have a separate, parallel version of Pip for testing.) Rather, the Pip bootstrapping system (which you can completely skip now, thanks to the `--python` hack) is dumb, along with all the other nonsense it's enabled (such as other programs trying to use Pip programmatically without a proper API, and without declaring it as a dependency; and such as empowering the Pip team to go so long without even as functional of a solution as `--python`; and such as making lots of people think that Python venv creation has to be much slower than it really does).

I'll be fixing this with Paper, too, of course.


> I feel like trying to work on more than one project in the same session would require more such discipline.

We all know that multitasking reduces productivity. But business often demands it (hopefully while being conscious of what it's costing).

You also don't have to be working in the "same session" to trip yourself up this way - "this terminal tab still has the venv from what I was working on yesterday/last week" is a way I've had it happen.

> I primarily develop libraries; if something breaks this way, I want to find out about it as soon as possible, so that I can advertise correct dependency ranges to my downstream.

If you want to find out as soon as possible, better to have a systematic way of finding out (e.g. a daily "edge build") than pick up new dependencies essentially at random.

> The requirements.txt approach does, of course, allow you to list transitive dependencies explicitly, and pin everything.

It allows you to, but it doesn't make it easy or natural. Especially if you're making a library, you probably don't want to list all your transitive dependencies or pin exact versions in your requirements.txt (at least not the one you're publishing). So you end up with something like two different requirements.txt where you use a frozen one for development and then switch to an unfrozen one for release or when you need to add or change dependencies, and regenerate the frozen one every so often. None of which is impossible, but it's all tedious and error-prone and there's no real standardisation (so e.g. even if you come up with a good workflow for your project, will your IDE understand it?).

> Fortunately, it looks like I'd be able to take advantage of the PEP 751 standard if and when I need that.

That's a standard written in response to the rise of uv, that still hasn't been agreed to, much less implemented, much less turned on by default (and unfortunately most of the time when you realise you need a lock file, you need the lock file that the first run of your tool would have generated when it was run, not the lock file it would generate now - so an optional lock file is of limited effectiveness). I don't think it justifies a "python packaging has never been a problem" stance - quite the opposite, it's an acknowledgement that pre-uv python packaging really was as broken as many of us were saying.


>None of which is impossible, but it's all tedious and error-prone and there's no real standardisation (so e.g. even if you come up with a good workflow for your project, will your IDE understand it?).

I mean, my "IDE" is Vim, and I'm not even a Vim power-user or anything.

People gravitate towards tools according to their needs and preferences. My own needs are simple, and my aesthetic sense is such that I strongly prefer to use many small tools instead of an opinionated, over-arching workflow tool. Getting into the details probably isn't productive any further from here.

>That's a standard written in response to the rise of uv

I know it looks this way given the timing, but I really don't think that's accurate. Python packaging discussion moves slowly and people have been talking about lock files for a long time. PEP 751 has seen multiple iterations, and it's not the first attempt, either. When uv first appeared, a lot of important people were taken completely by surprise; they hadn't heard of the project at all. My impression is that the Astral team liked it just fine that way, too. But it's not as if someone like Brett Cannon had an epiphany from seeing uv's approach. Poetry has been doing its own lock files for years.

>so an optional lock file is of limited effectiveness

The problem is that you aren't going to just get everyone to do everything "professionally". Python is where it is because of the low barrier to entry. A quite large fraction of Python programmers likely still don't even know what pyproject.toml is.

>I don't think it justifies a "python packaging has never been a problem" stance

That's certainly not my stance and I don't think it's the other guy's stance. I just shy away from heavyweight solutions on principle. Simple is better than complex, and all that. And I end up noticing problems that others don't, this way.


> People gravitate towards tools according to their needs and preferences.

Up to a point, but people are also nudged, not always consciously, by the reality of what tools exist in their ecosystem. The fact that Python makes "heavy" tools difficult to write and use is a significant factor in what many Python developers think is just a personal preference, IME. (I'd also argue that if you want to use lots of small tools you actually have more need for a standard format for your dependencies and your lockfile, since all the tools need to understand it).

> The problem is that you aren't going to just get everyone to do everything "professionally". Python is where it is because of the low barrier to entry. A quite large fraction of Python programmers likely still don't even know what pyproject.toml is.

Yes and no. I agree that many Python programmers aren't going to change the defaults and may not even know where their tool config file is. Any approach that requires extra effort from the user is not going to succeed. That's exactly why I think lockfiles need to be on by default, which is not something that has to make things harder for users (e.g. npm is a similarly beginner-first ecosystem but they have lockfiles and I've never seen it cited as something that makes it harder to get started or anything like that).


uv basically does that + python version handling + conveniences like auto-activating venv and installing dependencies


it was a massive problem at our company's hackathon. just so many hours wasted


Yeah, this is where I've been for a while. Maybe it helps that I don't do any ML work with lots of C or Fortran libraries that depend on exact versions of Python or whatever. But for just writing an application in Python, venv and pip are fine. I'll probably still try uv eventually if everyone really decides they're adopting it, but I won't rush.


I don’t care much about package managers. I left Ubuntu because it put a snap directory in my home folder, put ads in mtod, and made mount output unreadable. It felt like I lost control of my computer.


So dumb, it’s brilliant!


Cool project! If the input is an array of objects, the output graph is just a single block. Is this the expected output?


Thanks for letting me know this, I will make some tests and deploy the fix.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: