Hacker Newsnew | past | comments | ask | show | jobs | submit | lbriner's commentslogin

I suspect there is a name for thinking that, "if I do something, it will get resolved quicker". Lots of us won't just sit there and wait for someone to come and tell us when it is fixed because we don't trust that other people also want to resolve the problem.

To be fair, this sort of thinking is not entirely wrong (although it may be in this case). Your priorities are not necessarily the same as whatever person/worker you are dealing with. And the priorities of the worker you are dealing with are not necessarily the same whatever organization/company/government they are working for.

For example, I was recently at the hospital to pick up a friend who was due to be discharged that day. I arrived in their room, spoke briefly with the attending nurse who said it would be a long while before the patient was discharged (as it always is at hospitals) and sat down to wait. After a couple of hours I checked in with the nurse, who said he was waiting on the social worker. My friend noted that he had already talked to the social worker, and despite her insistence, said that he didn't want or need any sort of nursing or in-home care after discharge. When the nurse heard this, he went back, printed up the discharge papers and we were soon on our way.

Why did this happen? Because job of the social worker is to push the hospital/network's post-discharge nursing and rehab system. Once this was turned down, she had no incentive to stay involved or update the discharge status. The busy nurse had no incentive to pro-actively scrutinize the forms and decide the patient was ready for discharge. The hospital itself (as an organization) had no incentive to speed up the discharge, as they were making tens or hundreds of thousands of dollars from his excellent insurance coverage.

If we had just sat around, and said nothing, we almost certainly would have been there for several more hours, until the next shift of nurses came in and reviewed the patient files. Sometimes you are just screaming into the void but many times it helps to speak up.


Very little you can do will have any effect on how an airplane airplanes, particularly in a systematic failure.

Advocating for your health care can have very large effects, healthcare suffers from a strong lack of someone being in charge of and following your care, it's more like a collection of poorly communicating agents giving you brief spurts of attention somehow leading to a resolution. In other words, folks are barely paying attention to you.


Sounds like pedestrian push-buttons on crossings - satisfies people but it doesnt actually do anything.

This is, at best, misleading. It is common that a junction running at capacity with a pedestrian phase will always run that phase with the same timing regardless of buttons, but that same junction (never mind other junctions) may often not be running at capacity, whereupon that pedestrian phase is inserted when demanded.

Also of course outside of a junction the control system has no reason to stop the vehicular traffic if nobody demands to cross. So even in the middle of the day the only reason somebody who refuses to press the button can ever cross is that other people will press the button.


It wont turn green unless you push it here. If you dont push it, you will stand there like an idiot until someone else comes and pushes it.

https://www.bbc.co.uk/future/article/20150415-the-buttons-th...

There's also placebo buttons in many offices for aircon/heating.


I KNOW our work, because when I dont notice them, I stand there like an idiot until someome goes around and press.

Maybe there are the ones that dont work, somewhere, but there are also the one ones that do.


Alternatively you can play a round of hyper realistic frogger at any time

I am on a pretty strong winning streak.

I think a large part is being close enough to existing utilities to save the installation costs of 40 miles of cable, telecomms etc (all duplicated of course). However, some data centres are built in the middle of nowhere but that is potentially then a nature problem because why spoil this nice countryside with a massive non-descript building?

I like Copenhill (https://www.visitcopenhagen.com/copenhagen/planning/copenhil...), where they decided that waste handling is not nice so why not make it nice by building a ski slope on it. I suspect that a bit more creativity could make people want data centres (free heating and hot water anybody? Hot spring baths? Free hosting for local schools and charities?)


Clearly, it's both. Surely tech has a bad image and Flock does whatever they've done, but just because "one data center caused sludge" nowhere near equals "all data centres cause sludge", which is the implication from a lot of the anti-tech influencers. It's like saying, "data centers consume vast amounts of water" which, of course, is sometimes true but also sometimes not true at all. Data centres don't create many ongoing jobs, true but so what? Don't build anything that doesn't provide jobs?

People like my dad with no technical knowledge at all read an article in some rage-bait newspaper and it must be true (like foreigners took all your jobs, and the left-wing government wants to tax all your money and give it to scroungers).

If people are not taught how to ask "how often is this true" or "how much is this true" or "will this always be true", then they will fear whatever they are told to fear.


They had an article about Australia embracing AI data centers to be suppliers of the future and not just customers and one big pull is "the vast amounts of land in Australia" then said that one of them is being built within 100m of the nearest suburban house and 200m of a school.

It seems with all the other battles to fight, that moving it 2km down the road and surrounding it with trees and parkland would knock the noise argument out of the window. At least in the US and Australia that really do have vast amounts of open space.


That is an unfair conclusion. These people run complex networks like the rest of us, they probably have a range of detection systems and, also like the rest of us, an almost impossibly large attack surface to consider internally and on their supply chain.

The problem is that it is really, really hard to make something secure even if you try and follow all the best-practices you know.

I guess the awkward bit is marketing everything as certificate this, accreditation that and overselling how secure it is although I don't really know how else you would word it, "as secure as we know how"?


As Ops person, massive doubt. I've been at companies that have gotten hacked twice now, neither my department though. Both times, security vulnerabilities that hackers got into were well known, the tickets were in the backlog and deprioritized over feature requests.

I've also seen cases where it's like, maybe we shouldn't share S3 Root Creds or put it on the VPC so we can monitor outgoing traffic but too many applications would need to be redeployed for that so skip it. Those 2 year old tickets were still sitting in the backlog when I left.

If we ever get report, it's extremely likely going to be massive failure and only way to change this is fines for company that are bankrupting.

EDIT: Oh yea, SOC2 needs to go away. It's security theater that's just giving cover to companies.


I'd say it's not so much "as secure as we know how" and a lot more "as secure as we're willing to pay for". I'm sure this company has competent sysadmins and devs who'd be happy to lock things down. Usually management doesn't want the expense or the hassle.

This is kinda my point though - just don’t do it in the first place is the answer. Nothing is unhackable. So don’t create a massive honeypot in the first place.

Those aren’t the people in charge. People like 93 year old Senator Chuck Grassley are calling the shots.

It is not so much the eco-system, it is the need for a full food-chain of animals to keep populations in-check. We like deer but without wolves or large cats, the deer become a pest but we don't like wolves and large cats because we are all taught as children that they are scary and dangerous - although more children die from car accidents and stabbings than wild animals so whatever!

So natural predators for small rodents keep their numbers in check but that will be a long time because "won't somebody think of the children"..."and the farmers"


> it is the need for a full food-chain of animals to keep populations in-check

The error you make is that there is only one, global ecosystem, and it sorts itself out. We won't be matching the power of the dinosaur comet anytime soon, and even that gave rise to mammals and eventually space travel so just chill... We dont need to do a thing


> We won't be matching the power of the dinosaur comet anytime soon, and even that gave rise to mammals

Great, biodiversity will be restored in 100 million years, just in time for my great-great-great-great......grandchildren to enjoy.


Lol you wouldn't be here without it though.

> biodiversity will be restored

Restored? What are the appropriate levels of "diversity"?

Some facts humans cannot handle: - lifeforms suffer unimaginable agony every single day. - life is finite - humans are not beyond the natural order, human made global warming is natural, it won't end the planet or even us. - humans think because suffering makes them sad, suffering must be the bad. (Apart from when we don't like the other person or they've transgressed us in some way, were ethically clear to mete out violence)


> Lol you wouldn't be here without it though.

And?


Another disaster would be great the for the universe the same way we think we were great for the universe

Judging by the seats, perhaps making it as light as possible. Existing airframes are relatively light but let's be honest, they aren't trying to save 500g here and there when they have massive jet engines to get everything airborne.

I discussed with people working on aircraft components, and my understanding is that weight is a constant obsession for them. Propose a new system to Airbus/Boeing, and the first question they'll ask is "how much does it weight?"

Every kg saved is a kg more of freight that can be transported (or a little less fuel used to keep the plane in the air); save 500g for each seat of a 200-seat plane, and you can put one more paying passenger in the cabin.


Agreed. But for related reasons, aircraft aren't really designed for either the structural robustness or balance to deal with big heavy batteries, which is what motivated the clean sheet design.


After having my first child (a girl who is now nearly 3 years old) these stats impact me so much more. My first roll was a girl in the Iraniam plateau in the 18 century who had a 33% chance of dying before they reached 5. She managed to live to 70 but experience so much death in her life.

I can't imagine my daughter dying before she reaches 5, it feels like it would be the worst feeling I could possibly imagine and then some.


Visit any older cemetery and look for the tiny headstones and the matching very brief lifespans, or read an older obituary for someone from your great-grandparent's generation or further back, or watch a couple of seasons of _Call the Midwife_.

A text editor is very personal to your use-case and experience, which is why not everyone likes any particular editor.

Before I knew about multi-line cursors, I wouldn't have cared a less about whether an editor did that. Now that I know it and use it and love it, I would never choose to use an editor that doesn't have it.

But since I not really spent much time identifying what should be quicker for me in a text editor, I am probably much happier still with simple/quick/responsive compared to some of the vi Gods who require at least 8000 macros to be productive and would never live with a mortal text editor (or emacs :-)


It is often not worth optimising in the early days. You don't know how popular it will become, you might not know how many DNS records you will hold, it was possibly written in an earlier language and ported as-is.

At the point someone queries the 100TB of RAM, then maybe it is worth revisiting but even that has risks. You have to design the migration path, have fallback mechanisms etc.


It's also often that you can avoid all those future migration/fallback risks and pains if you invest a little bit of design thinking upfront.

So how would you decide which path to take in situations like this?


It only looks super obvious in hindsight and the well explained blog post. when a team of 5 is tasked with getting a completely new DNS up at the scale and integrate well with cloudflare.

if you spend cycles on nitty gritty opinions like this time to market goes out further and further out. some napkin math, 130 gen13 servers cost "only" ~$2.6M. relative to the importance of the 1.1.1.1 and the market at the time. that is nothing to cloudflare.

this is not to say good system design does not matter. it very much does, but making that call at that time would've butchered the prodcut very much similar to google+, youtube etc.


This one also looks pretty obvious "in foresight" (using the same tools that existed back then. Maybe owner dedupe might be less obvious and require a bit of knowledge and probing into actual data, but for rw vs ro you are fine knowing nothing?) and you forgot the napkin math re. how much your precious "time to market" would have been delayed by.

It's also not nothing, otherwise it would never be optimized away now, but left as is. After all, wasting time on optimization delays "time to market" for other useful features.

I also don't get the reference to YouTube, it's a very successful product, how was it butchered by good system design???


Imagine you're an engineer at cloudflare, an 8 year old (at the time of launch of 1.1.1.1) company. The company is wildly popular and any service launched is going to have a lot of traffic and a lot of attacks right away. Any problems with it are going to embarass the company a lot.

You're tasked with making a DNS caching recursive resolver that can operate at a large scale and will be run on thousands of servers each of which has a lot of GBs of ram.

You are given some period of time to build this and make it production ready. How do you spend your time:

* Focusing on making sure that the resolver works correctly?

* Focusing on make sure that it actually provides improved DNS performance for internet users?

* Handles an very large number of record requests/s?

* Saves a few GB of ram per server?

There are tradeoffs to consider. RAM is cheap, even at today's prices RAM is not the most expensive thing that can go wrong in such a scenario. Having the responses be slow or incorrect is a far more expensive problem. A good engineer would pick a simple data structure that has the right shape but might not be optimal in footprint to focus on correctness and response time. The few extra GBs of RAM per server can be dealt with later.

When building things at scale you want to make sure it works correctly, fails correctly, and does the thing quickly before worrying about reducing resource consumption. I've never seen a project fail on Vec<T> vs Box<[T]> memory differeneces, or even on a few GBs of RAM usage per instance. I have seen them fail on "one wierd corner case of correctness" though, and on poorly thought through failure modes.


> The company is wildly popular and any service launched is going to have a lot of traffic and a lot of attacks right away.

Doesn't this also inform you that your cache will be very large, so you shouldn't use growable structures with slack space when cache entries won't grow; slop space reduces the size of your cache. And also that the query volume will be high so the cached data should require as little work as possible before returning data; spending time marshalling response data on every cache hit increases response time and decreases capacity.


RAM is cheap. I'd find myself far far more concerned with:

* unbounded growth of the cache and properly invalidating after TTL expires (a few GBs of slop is nothing on a server with 64 or more GBs of ram, unbounded growth is a problem).

* making sure the DNS implementation works correctly on both the serving side and recursive resolution side.

* What strategy is best for deduping recursive requests across machines (if something a few miliseconds away has a live result, why do a full lookup taking hundreds or thousands of milliseconds?). This potentially improves RAM usage across the datacenter too from not having a given record on dozens (or more) machines' local cache. I don't know exactly how they do it, but naively I'd look at some sort of DHT shaped solution to look for records in peers within the datacenter. Or maybe some sort of tiered caching with the upper tier being sharded on domain name or the like.

* The biggest performance gains cloudflare can provide in Web and DNS cache come from a cache hit. This is on the order of 10s or 100s of ms due to having a big cache and short distance to the requesting machine. A suboptimal lookup algorithm that is a few microseconds slower in local compute and ram access is just not as important as the other concerns for dedup and cache sharing. That's not to say it's unimportant, just that it's not the top priority when you're trying to deliver this much larger performance gains from other aspects of the system. Thats why they are getting to it several years after release.

Cloudflare writes a lot about distributed systems solutions to various problems. They likely don't think as hard about single machine performance as much as whole datacenter performance when approaching problems.

Keep in mind that the per-server cost of the whole program pre-optimization seems to be about 10GB (from the graph in the post). IME that's not bad for a big busy caching service.


> The biggest performance gains cloudflare can provide in Web and DNS cache come from a cache hit.

Using twice as much ram per cache entry makes the cache half as large, assuming your cache is bounded by ram, unless the queried, unexpired result set is less than the ram budget (which I would tend to doubt... lots of randomized queries out there; maybe I'm wrong if the cache size dropped).

When you're storing billions of records, it makes sense to spend a few minutes to consider how they're used and make a good choice about how to store them.

When you're getting a cache hit tons of times per second, it makes sense to consider every step and which ones don't need to happen every time. You have to consider every step while you're pursing correctness anyway, so might as well have the performance lens active too.

I'm not asking for heroic optimization: I didn't ask for vectorized stuff or kernel/nic offloading or kernel bypass networking... Just you have to use some data structures, you might as well not use ones that are expensive for features you don't need; and you have to store something in your cache, you may as well store something that requires less munging on the way out.

If this were a small local cache, that didn't want to use something already existing like unbound for some reason then yeah, data structures don't make a huge difference, extra marshalling doesn't make a huge difference, just don't reimplement all the CVEs that BIND had in the 90s. But if you're going to allocate 100 TB of ram, make it count. Even if you do use twice the ram but you get value from it, maybe that's fine... I've run wacky systems with bloated storage when there was a benefit. Vec doesn't give any value over a Box<[]> in this case; convenience or lazyness would be fine except that the sheer number of objects makes it worth the few minutes it takes to do something better.


> Using twice as much ram per cache entry makes the cache half as large, assuming your cache is bounded by ram, unless the queried, unexpired result set is less than the ram budget (which I would tend to doubt... lots of randomized queries out there; maybe I'm wrong if the cache size dropped).

This is true. I'm arguing that its unlikely this was ever bound by available RAM. Cloudflare is a DDoS protection company that absorbs attacks. They have a lot of available capacity at any moment. When you're building a service in a sitaution where you have more capacity than you'll likely need.

The savings were 100 TB across >300 data centers. The savings were on the order of 50%. So prior to this reduction the service was using something less than 2/3 of TB per datacenter. The service ram usage was about 10GB per instance according to the graph in post. IDK how cloudflare divides thier stuff between machines, but assuming they don't run less than 64 GB per server that's less than 12 servers per datacenter of ram for a flagship product, and they likely run it spread across 65 of the machines in the datacenter that are also doing other stuff. The per-instance RAM likely isn't the the concerning limit.

Overall RAM usage is proabably a bigger concern. Thats why I would think about dedup between instances and distributed caching strategy first. I could focus on redudcing the ram needed per service instance and get a 50% reduction per machine. Or I could focus on deduping 1/n (where n > 2) reduction in total memory usage across all instances. Personally if I was worried about reducing RAM I'd put more energy into growing N.

However all this is a red herring. The assumption people are making is that the cache was always read-only, and it's obvious that Box<[T]> was the best decision because in a RO cache smaller entries hold more things.

The 1.1.1.1 service advertises improved DNS performace. That's its value add. The biggest performance gain you can have from a cache is not having a cache miss, and in DNS a cache miss means a very expensive recursive lookup. So there's concerns about how to minimize those lookups. If one instance has does a lookup, it makes sense to share that result to the other instances that may need to do a lookup [1]. I don't know off the top of my head if it makes sense to get those updates and modify the existing record or just replace it in the local cache. That comes down to locking strategies and reading patterns in the specific code and service traffic patterns. Until i have hard evidence one way or another I'd like my cache to be able to do both and keep the data structs modifiable until that's nailed down. If per-isntance ram ever becomes the issue, there's easy wins there to buy me time to find better large scale solutions to the problem.

No one is disagreeing that the larger datastructures are larger. No one is disagreeing that they take more RAM, and or even if RAM was the the problem reducing it would be good.

The thing people are pointing out is that this isn't a homework problem about an optimal cache structure in a vacuum. We're pointing our that engineering real large scale solutions has a lot more to consider than a homework problem, and that the thing you're harping about likely didn't have any real budgetary or noticable performance impact on bulding that system. The reduction in ram is just a smallish improvement in operating costs after all the more expensive stuff was figured out.

Put another way 100TB of RAM is ~$350K. Thats one engineer year for a mid-level engineer.[2] Would you rather spend that money to save an equivalent amount of money somewhere, or... would you spend that money putting the engineer on something that saved $700K elsewhere (alternately that generated $700K)?

[1] I talked a lot about dedup and the simple gotcha is "hahah then its not deduped so you need smaller objects". But on a service that is running on a few dozen instances having a few redundant copies to deal with loss of a machine and/or load can still result in 1/(n>2) savings in total ram.

[2] I'm not saying someone worked on this for a year btw, a couple people likely spent a couple months on the code, validation and testing of it. A manager spent time overseeing it. Operations people spent time understaning any effects it had on running systems. Costs add up and it wouldn't suprise me if this didn't end up being roughly break-even for the year.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: