The systems that taught me the most were not the ones that crashed spectacularly. They were the familiar endpoints that slowly became expensive to change: a small shortcut here, a hand-run release there, one query nobody measured. Then traffic arrived, latency climbed, and every innocent line had an opinion.
I do not have five tricks for handling millions of requests. I have five habits that keep a system legible while the requests arrive. They make incident response calmer because they make the code less surprising.
1. Simplicity over cleverness
The clever version of a decision usually saves lines, not time. Under pressure, the next person needs to see the business state without mentally executing a puzzle.
Before
const status = paid
? "paid"
: failed
? "failed"
: retrying
? "retry"
: "pending";
After
function paymentStatus(payment: Payment): PaymentStatus {
if (payment.paid) return "paid";
if (payment.failed) return "failed";
if (payment.retrying) return "retry";
return "pending";
}
The second version makes precedence explicit and gives the decision a name. That name becomes useful when the rules grow, when we log a transition, or when someone asks why a failed payment still appears pending.
Trade-off: named branches can feel verbose for a two-state decision. Keep a ternary when it is truly binary and obvious. The line is crossed when a reader must remember precedence or domain rules to understand it.
2. Scale by design
Scaling is not adding a queue after the outage. It is deciding which work belongs to the request path before the endpoint becomes popular. A user should not wait for every email provider round-trip just because they changed a setting.
Before
for (const user of users) {
await sendEmail(user);
}
return { delivered: users.length };
After
await notificationQueue.addBulk(
users.map((user) => ({
name: "email",
data: { userId: user.id },
})),
);
return { accepted: users.length };
The API now owns acceptance and validation; workers own retries, provider limits, and delivery. The response contract is honest about that boundary.
Trade-off: queues add monitoring, idempotency, and eventual consistency. Do not queue work whose result the caller must receive synchronously. Design the boundary early; do not use a queue as decorative infrastructure.
3. Measure before optimize
The slowest thing in an incident is often the argument about what is slow. I have seen teams rewrite a cache layer while one unindexed query was doing all the damage.
Before
const result = await expensiveQuery();
return compressResult(result);
After
const startedAt = performance.now();
const result = await expensiveQuery();
metrics.histogram(
"report.query_ms",
performance.now() - startedAt,
);
return compressResult(result);
Measure a boundary that can drive a decision: query duration, queue age, dependency latency, error rate. A metric without an owner or a threshold is just telemetry confetti.
Trade-off: instrumentation costs time, storage, and attention. Start with the request path and the user-visible objective. Measuring every function can be as distracting as measuring nothing.
4. Automate everything repeatable
If a release requires a senior engineer to remember six terminal commands, it is not a process. It is a ritual with a high bus factor.
Before
npm run test && npm run build
scp -r dist/* server:/var/www
ssh server "systemctl restart portfolio"
After
deploy:
needs: [test, build]
steps:
- uses: actions/download-artifact@v4
- run: pnpm deploy
Automation does more than save clicks. It records the exact path to production, makes it reviewable, and fails consistently when a prerequisite is missing.
Trade-off: automate a stable process first. An unreliable manual process becomes an unreliable automated process faster. Write the checklist once, then encode it.
5. Clean code survives longer
Traffic does not kill code as often as change does. A checkout function that validates input, calculates prices, charges a card, writes data, and sends a receipt can work for months. It becomes dangerous the first time one of those responsibilities changes alone.
Before
async function checkout(req: Request) {
const input = await req.json();
const total = price(input.items);
const charge = await chargeCard(input.card, total);
await database.orders.insert({ input, total, charge });
await sendReceipt(input.email, total);
}
After
const order = await orderService.create(input);
await paymentService.capture(order);
await receiptService.send(order);
The orchestration still exists, but each collaborator has one reason to change and one surface to test. That makes failure handling explicit instead of accidental.
Trade-off: do not split code merely to create more files. Extract a boundary when it has distinct policies, failure modes, or tests. Clean code is not maximal abstraction; it is code whose next change has an obvious home.
Before I call it ready
- Can a new owner explain the decision path without decoding a trick?
- Which part of the work must finish before the request can return?
- What metric proves the change helped the user?
- Can the release and recovery path run without memory-based instructions?
- Does each change have one obvious place to live and one focused test?
These questions do not guarantee a quiet pager. They make the answer to a noisy pager easier to find.
Comments (0)