There are a lot of non-white british people (just like there are non-white americans), and he basically ignores that entire population. He assumes that you are only british if you are white and that london is worse now because there are less whites there. That has nothing to do with meritocracy, and arguable little to do with cultural non-fit.
There's a lot more to unpack in that article aside from just that, but that is the most blatantly obvious one.. And for what it's worth, I don't disagree with DHH on mass immigration. But he seems to have a lot more severe beliefs than just "mass immigration is bad".
> There are a lot of non-white british people (just like there are non-white americans), and he basically ignores that entire population.
There are many people who are mostly or wholly culturally British. I am not white and I live in the north of England and the biggest cultural non-fit is that my speech is obviously southern.
London does not, on the whole, feel very different culturally, from when I grew up there in the 80s. The biggest change in terms of the cultural impact of immigrants has been that of Eastern Europeans: things such as the availability if Polish food.
How is that the same thing at all? You didn't understand the comment you replied to and were confidently wrong. The person replying to you understood what you were saying very clearly and didn't make any ridiculous "you may not last long" comments like you did...
The truth is people just don't seem care about code quality anymore (if they ever did?). I don't think it generates great code (yes even using the state of the art, frontier, super max pro 9000 turbo boost ++ models), but it usually generates code that works. Sadly, that's all that most of the people in charge care about. I mean shit, half the time it doesn't even have to work great, look at the state of a lot of modern software. People complain about it all the time. Quite sad, but that seems to just be the state of things these days.
I don't think corporate ever gave a shit about code quality. Honestly if there weren't building codes skyscrapers would be made with the cheapest shit that didn't collapse immediately and your house would be made of cardboard. Similarly if users were OK with buggy crap or if they had no choice (e.g. windows) then you can get away with lots of crap.
You could only argue about it when an engineer was in charge and even so when they're high up enough the pressures and incentives are to just ship fast and break stuff. Not every company can afford to be a NASA making uncrashable code that runs for 100 years and goes to pluto and all.
It’s quite an uncanny alignment with Nvidia and other semiconductor companies that benefit from their customers to bloating software and buying as many of their chips with no real thought given to efficiency of the runtime.
This is pretty typical for a new blizzard game. For example Diablo 3 and 4 were both announced 4 years prior to release. Not saying it’s good or bad, but it’s not unusual.
I came to say the same - I remember it being an excruciatingly long wait for SC2, which was also apparently over 3 years between announcement and release [0]
The [lack of] integrity of OpenAI (and any other frontier lab) should already be pretty solidified. Among other horrible things, these companies stole millions of IPs and no one seems to care anymore. Regardless of what you think of the product they are making and the success of ai/its impact on humanity, these companies objectively do not have much integrity.
How do you feel about the integrity of the machine learning researchers over the past twenty years who trained models on scraped internet data that weren't particularly powerful and didn't attract any attention?
If they scraped internet data in the same way as current day frontier labs do, then I feel the same exact way about them. Why would I feel any different if that is the case?
My point is that researchers and academics really have been doing this for decades - it's the reason projects like Common Crawl and LAION exist.
I think it's notable that nobody was calling out those researchers for their lack of integrity, because the systems they were building did not seem like a threat to anyone.
OpenAI etc get accused of a lack of integrity on this precisely because the systems they are building work, and are profitable.
My personal opinion here is that integrity is more about what you build with the data. I think saying "scraping means you lack integrity" is a simplification.
You're right, it was an over simplification. I think public exchange of data is great for innovation and research (Common Crawl/LAION). But I still think scraping proprietary data without consent or attribution is generally bad (also Common Crawl/LAION).
Then you have OpenAI etc.. who build these multi-billion (trillion??) dollar machines and sell them back to people, using everyone's proprietary data, and (among other things) tell everyone it's going to take their jobs. That combination of things doesn't scream integrity to me.
Still, it's undeniable that these machines could be beneficial for humanity (cancer research and such). So, I'm sure many people would say the good out-ways the bad. I don't know. Seems that would set a risky precedent for future companies, but maybe not.
you massively collapsed what AI companies have been doing by comparing it to old internet-scraping. Facebook flat-out admitted that they scanned copyrighted books for their AI. The image generators most definitely trained on copyrighted images.
LAION and Common Crawl both scraped copyrighted images. From what I can tell (I'm not an expert in this domain at all), the main difference between those two and frontier labs is in how they stored and used the data. CC and LAION seem to be actually open (unlike "Open"AI) and are more centered around publicly sharing the data they scrape to support research and innovation.
OpenAI et al also stole everything from everyone. But then they raised billions of dollars from that data and sell back their LLM to people (again, among other things). They are also very much NOT open in any way, aside from sharing their benchmarks of new models.
What I meant more is that the dataset they scrape is openly available for download by anyone, unlike any of the frontier labs. Not that they don’t scrape copyrighted content. Still sketch, but at least they don’t call themselves “OpenLAION”.
Also my understanding was they’re not storing the actual music, but the metadata and a link to the YouTube video.
The thread is not really about what's legal; the topic is integrity. It sounds like, based on the fact that you wouldn't do it yourself, you agree that it's not a good thing to do.
i think that this case, if they did train on buckmaster and alpöge, amounts to an attempt to steal the millenium prize, bypassing all attribution.
legally speaking, the default privacy notice gives them an irrevocable license to your content. they may read and use the prompts for research. so it is very possible they simply stole the navier-stokes solution.
that is the same principle as any other prompt but this would be a concrete example.
there would be some difference between simply giving the model some prompts to read, which they are entitled to do on the default policy, and putting it into aggregate training data.
reply