Hacker Newsnew | past | comments | ask | show | jobs | submit | benswerd's commentslogin

So dope

I agree.

My first time playing StarCraft was at summer camp around a decade after it came out.

All the smartest people played it so I wanted to too. Great decision, I have been continually impressed with the people who StarCraft introduced me to.


I might open this up to a tournament if enough people want. Any interest?

Hah, yeah that's definitely interesting, though maybe a general platform for this stuff would be even more interesting.

Although, maybe benchmarking an agent on "how well can you command a swarm to annihilate the Terrans" is how it all starts going downhill...


I predict LLMs will reach superhuman level and beat even that model in the next 12 months

Starcraft is APM-dependent. Unless the latency will improve greatly in frontier reasoning LLMs (which is unlikely), it will remain a bit like knitting with an excavator.

I predict latency will improve greatly in the next 12 months to more than 4x speed on current frontier tasks

The best sc2 bots these days play in the ~50,000 APM range (they could mostly go higher as well, but the game client breaks somewhere around 100k APM). I don't see an LLM-based bot getting up to that sort of speed anytime soon.

On the other hand, I do think LLM-based bots will quickly outperform the decision-making of many of the hand-coded bots, so maybe they won't need so much APM to be competitive.


Yeah but Starcraft needs, like, 10-20x the APM these agents are doing.

I’m not convinced a lot of it can’t be solved with code mode.

Marine staggering for example seems like an ideal code mode task.


Do the humans get to use this auto-stagger too?

Yeah. Some of it may just be "thinking" less rather than faster token generation.

Should be possible to play this with Jev.

Not hard to build. I was shocked at how fast/easy this was to pull together.

Oh sorry I should be more clear on that. Will add to report.

For agent harness I did Claude Code, Codex, Grok Build. This was primarily a cost driven decision — I have a lot of free tokens and I didn't want to pay API prices for this.

For game harness I used minimal BW-API issue command and get observation apis as tools. I felt this was the most fair way to do it on my small scale.

In the future I would like to integrate code mode and multiple games/I think if it was a best of 5 where each agent could learn from its past games and build its own automations over time that would be much more interesting.


Given that a lot of their failures are from fairly basic mistakes related to the unique setup (eg, thinking rather than defending immediately) I'd love to know how much they improve with basic tips.

Or possibly even whether they can learn from a game themselves. "Analyse your game for your failures" -> Then give a fresh agent of the same model that "learnings" doc for the next match. Do the rankings change over time, if models can write instructions for future selves?


Try it, you can run your own games on bw.swerdlow.dev

+ Playable Agent driven Starcraft

Dibs on the fly brain

"THIS FLY CAN BEAT YOU AT STAR CRAFT: HAS SCIENCE GONE TOO FAR????"

What is the hardest part of building this for you?


I have a payments background. So maintaining a very high bar on security, reliability and performance as usage scales is super important.


What made you choose digital ocean?


YC gives DO credits for startups.


Their machines are good & they have a partnership program that's fast and compatible with the model.


Its not a zero day. I found this while zero day hunting.

It is a surprise and I cannot reveal what it is 4 the sake of the memes.


Nobody is going to clone a repo just to look at it. If you want people to look at it, you must provide some info or it will just be ignored.


I mean, I can guess, given one of the URLs. I look forward to subsequent reporting about it.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: