Platform & architecture
Black Friday doesn't fail for lack of servers
Most e-commerce operations that fail on Black Friday don’t fail for lack of servers.
It’s counterintuitive, because servers are where the preparation goes. Horizontal scaling, CDN, caching, load tests on the core infrastructure. All of it is necessary, and almost everyone does it.
What almost nobody prepares with the same care are the parts that are not under your control.
Where the peak actually breaks
The bottlenecks that cause the most trouble under peak demand rarely sit in the core of the platform. They sit at the edges:
- A partner API with no documented rate limit, and no defined behavior for when the limit is hit. Payments, fraud prevention, shipping quotes, address lookup.
- A database query that works under normal load and locks up under real volume, because nobody tested it with production data and the real access pattern.
- A synchronous inventory integration, where every product lookup blocks until it hears back from a system that was never sized for that volume.
- User sessions spilling into a database that is already saturated, instead of living in their own layer.
The pattern repeats: your own infrastructure holds. The bottleneck is some integration that was treated as an implementation detail rather than an operational risk.
Why it goes unnoticed
Because load tests usually test what is ours.
Simulating the peak on your platform is fairly simple. Simulating it on a partner is hard: staging environments that don’t reproduce production, contracts that don’t allow volume testing, SLAs written for the average and not for the peak.
The result is preparation that measures the easiest part of the system and leaves untested exactly the part that is going to fail.
And there’s an aggravating factor. On a DTC platform that absorbs several times its baseline volume at peak, a slow dependency doesn’t only take itself down. It holds connections, fills queues and contaminates the rest of the flow. The symptom shows up at checkout; the cause is three integrations upstream.
What real preparation takes
1. Map every external dependency and how it behaves under volume. Not the list of integrations: how each one behaves when demand multiplies. What’s the limit? What happens when it’s reached? Who gets paged?
2. Load test the full flow. From search to confirmed order, through every integration on the way, not just your own infrastructure.
3. Define a circuit breaker and a fallback for every critical integration. If shipping doesn’t answer within two seconds, what’s the default response? If fraud prevention goes down, does the order wait, go through or stop? These decisions need to be made calmly in October, not at ten o’clock on Friday night.
4. Separate the essential from the accessory. Personalized recommendations, product reviews and the free-shipping badge are great. None of them is worth taking down checkout. Whatever can be switched off under pressure should have a switch.
The question that changes the preparation
The wrong question is “can our infrastructure handle the peak?”. It almost always can.
The right question is “which integration fails first, and what happens to the order when it does?”.
Servers are the easiest thing to scale. They’re rarely where the problem is.