본문으로 건너뛰기
L14.12

Capacity Planning and Failure Recovery

Goal

Estimate safe service capacity, leave room for bursts and failures, and define what the service should do when a replica or dependency fails.

Plan with headroom and failure reserve​

Capacity planning is easier to understand with numbers. Suppose one replica can safely complete about 10 requests per second under the tested workload, but you choose a target utilization of 80% so normal planning counts only 8 requests per second. Three replicas then provide about 24 requests per second of planned safe capacity.

Now test a failure. If one replica disappears, only two remain, so planned safe capacity falls to about 16 requests per second. If normal demand is 20, the three-replica service handles normal traffic but does not survive one replica failure at the chosen safety target. That tells you the plan needs another replica, a lower workload, or a defined degraded mode.

The unused part is headroom: capacity intentionally kept available for bursts, retries, and failures. It is similar to leaving empty seats on a bus when you know a group may join at the next stop. Running at the absolute maximum leaves no room for surprises.

safe capacity per replica at target utilization = 8 req/s
3 replicas → 24 req/s
lose 1 replica → 16 req/s
normal demand = 20 req/s → failure reserve is not enough

Capacity planning therefore asks whether the healthy service can handle normal traffic, bursts, and failures. Counting GPUs is not enough. Queue rules, token length, model size, batching, and reserve capacity all change how much work one replica can safely handle.

Measure workload shape, not one average​

Measure the shape of the workload. Requests per second matters, but a long request can use much more compute and KV-cache memory than a short one. Track input length, output length, and concurrency instead of planning from one "average request."

Replica count should include a failure reserve. If normal traffic needs every replica, losing one creates overload. The exact reserve depends on cost and service importance, but it should be a deliberate choice.

Use queues and degraded modes intentionally​

A growing queue is an early warning. If requests arrive faster than the service finishes them, waiting work keeps building. Backpressure slows new work; load shedding rejects some work. Either can be safer than letting everything wait until clients time out.

Recovery rules should separate a temporary process failure from a persistent model or device failure. Repeatedly restarting a process that always runs out of memory creates a crash loop. Stop routing to it and escalate after a bounded number of attempts.

A degraded mode keeps the most important service working with fewer resources. It might shorten maximum output, turn off an optional feature, use a smaller fallback model, or reject low-priority batch work. Define and test these modes before an incident.

Re-measure after capacity changes​

Capacity changes can change behavior too. Adding replicas changes routing and cache warmth. Removing replicas can increase queue pressure. Measure the effect instead of assuming scaling is invisible.

A useful plan states expected demand, safe capacity per replica, headroom, minimum healthy replicas, queue and load-shed rules, fallback modes, and the telemetry that triggers action.

Predict

A service needs all four replicas at 95% utilization to meet normal traffic. What is the operational concern?

Run the Docker-environment Lab preflight​

This activity is registered for the Docker-oriented production environment used in the second half of Level 14. Start with the deterministic Python preflight so service-logic failures remain distinguishable from container, GPU, cluster, or credential setup problems.

Run:

python3 labs/notebooks/level-14/l14-12-capacity-recovery.py

The Lab sizes capacity for normal traffic and then removes one replica.

  1. Run it unchanged with target_utilization = 0.8. Record required/planned replicas and the utilization after one failure.
  2. Before editing, predict what happens to reserved headroom if the planner is allowed to target 98% utilization instead. The traffic rate and per-replica safe capacity stay fixed.
  3. Change only target_utilization = 0.8 to target_utilization = 0.98, then rerun.
  4. Confirm the plan uses less reserve and the one-replica-failure scenario crosses into high queue risk. Higher nominal utilization trades away failure/burst headroom.

Loading lab…

Write the core logic yourself​

Open:

labs/notebooks/level-14/l14-12-capacity-recovery-exercise.py

Implement required-replica calculation with target utilization, then check whether one replica can fail without dropping below the requirement.

Run:

python3 labs/notebooks/level-14/l14-12-capacity-recovery-exercise.py

The starter intentionally stops at TODO until you implement the missing logic. A correct solution reaches the final PASS: marker. Use the solved deterministic Lab as a comparison only after your own attempt.

Quick Check

1. Why plan from token/work distributions instead of request count alone?
2. What happens when sustained arrival exceeds completion capacity?
3. What is a useful degraded-mode property?

0 of 3 questions answered.

Explain it back​

Create a capacity plan for a three-replica service. Include target utilization, one-replica failure reserve, queue limit, load-shed rule, and one degraded mode.

Key Takeaways

  • Capacity planning combines workload shape and per-replica capability.
  • Keep headroom for bursts and failures.
  • Persistent arrival above completion grows queues.
  • Bound restart/recovery attempts and remove unhealthy replicas from routing.
  • Predefine degraded modes and load-shed rules.

Next Lesson

Next, L14.13 — Production AI Operations Workshop integrates serving, deployment, telemetry, rollout, and recovery into one release decision.

References

Lesson actions

Completion is stored locally on this device.

View progress