Hacker Newsnew | past | comments | ask | show | jobs | submit | armanified's commentslogin

I'm a professional web scraper and have pretty much exhausted most open-source stealth browsers. The paid ones aren't a very economical solution when you want to scrape at scale and have to deal with heavy bot mitigation.

The issue I repeatedly faced with most tools is that almost all of them run on Linux and try to pretend they aren't, which might work for a while, but given enough samples, it's pretty obvious to large bot-mitigation infrastructures, and they start blocking or at least push a bit harder.

So I tried to build a browser myself, which does a few things:

* Randomize profiling based on the given seed * Match the timezone based on the proxy exit IP * Can outsource Canvas to a different machine ( from the target OS )

This canvas outsourcing is what really makes it different, because pretending isn't enough. Some of the toughest bot-mitigation infrastructures probe deep into the canvas, a depth that a simple patch can't bypass on a different OS. But this is optional and should only be done if everything else fails.

Right now, this is an early release. Any feedback is highly appreciated


True, but most would ignore LM if it weren't LLM.


This title might have triggered something in those bots; most of them have sneaky AI SaaS links in their bio.

Honestly, I never expected this post to become so popular. It was just the outcome of a weekend practice session.


OMG! Why didn't I thought fo this first :P


OMG! You just gave me the next idea..


Pretty neat! I'll definitely take a deeper look into this.


Uppercase letters were intentionally ignored.


My initial idea was to train a navigation decision model with 25M parameters for a Raspberry Pi, which, in testing, was getting about 60% of tool calls correct. IMO, it seems like around 20M parameters would be a good size for following some narrow & basic language instructions.


Ok. This makes me wonder about a broader question. Is there a scientific approach showing a pyramid of cognitive functions, and how many parameters are (minimally) required for each layer in this pyramid?


It mostly doesn't, at 9M it has very limited capacity. The whole idea of this project is to demonstrate how Language Models work.


I haven't compared it with anything yet. Thanks for the suggestion; I'll look into these.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: