saved
Some very non-obvious lessons about how the cloud works
by Slava Akhmechet · X · published
Hraness republishes this public post from a saved copy. The post is the author’s own words.
Some very non-obvious lessons about how the cloud works:
- There are ~500 services (give or take), each more or less its own independent business. So you've got 500 startups, many of them hypergrowth, operating under one roof.
- Like all startups, services are governed by the power law. Some are zombies that barely make a few million after many years, some are in the billions and growing very fast.
- Pace and urgency are very similar to a hypergrowth startup. Every Monday it's hard to remember the previous week because so much has happened.
- Each service recursively depends on other services. E.g. queues depend on compute, storage, networking, telemetry, and a dozen other services, each of which has its own dependencies.
- Which means if anything breaks, everything up the stack breaks.
- All these services are in dozens of regions, so something is always broken somewhere (and thus everything above it is broken)
- There is a lot of infra to help with reliability-- internal services to help detect failures, trace dependencies, mitigate common problems, etc.
- Depending on how you look at it, this infra is mature and very helpful, but is also evolving and improving at crazy speed.
- Everyone is always capacity constrained, despite constant hardware buildouts. Sometimes important customers can't get capacity at any price. (Now that I think about it, I don't understand why we don't auction off capacity. I'll ask around.)
- Customers buy into an ecosystem, not into an individual service. So once they bought into a cloud, each service gets free tailwind.
- But the flipside is also true-- when a service falters it puts the whole ecosystem at risk (which you never want to find yourself doing, obviously).
- Fleets are huge. Just crazy scale.
- Operational load is extremely spiky. Maybe you normally have to manage a few failovers here and there, but then an availability zone fails and you get swarmed by a zillion failovers you have to handle _now_.
... And lots and lots of other stuff. I'm running out of steam, maybe will pick it up again in another post.