Skip to main content
L14.3

Batching and Streaming

Goal

Explain why a serving system groups requests, why waiting for a batch can hurt interactive users, how streaming changes what the user sees, and how scheduling choices trade throughput against delay.

Start with three requests arriving together​

Suppose three people ask the same service for different amounts of text:

A: 40 output tokens
B: 200 output tokens
C: 25 output tokens

One simple server could finish A, then B, then C. That is easy to understand, but a GPU can often do more useful work when it processes several sequences together.

Batching groups active requests so one model step can serve more than one sequence. The benefit is better hardware use. The cost is waiting: a request that arrives now may sit in a queue while the server decides what else should run with it.

That trade-off becomes obvious with a fixed batch of 32. During heavy traffic, 32 requests may arrive quickly. Late at night, the first request might wait a long time for 31 more requests that are not coming. A throughput optimization has turned into user-visible delay.

So do not ask only, “What batch size makes the GPU busiest?” Ask two questions together: “How much work does the hardware finish?” and “How long does an individual request wait before useful output begins?”

Continuous batching changes admission​

Continuous batching changes the scheduling model. Modern serving runtimes can admit new work as other sequences finish rather than keeping one fixed batch until every sequence completes. vLLM documents continuous batching as one of its throughput-oriented serving features. The exact scheduler changes over time, but the stable idea is dynamic admission of active requests.

Stream for responsiveness​

Streaming solves a different problem. Instead of waiting until every output token is generated, a server can send partial output as tokens become available. This can improve time to first token, even though total generation time may remain similar.

Batching and streaming can coexist. The scheduler may batch several active sequences on the device while each client receives its own stream. Request identity is essential so partial outputs are routed to the correct connection.

GPU step 1: [A, B, C]
GPU step 2: [A, B, C]
B finishes → D can enter
client A receives tokens as they are produced

Protect fairness and cancellation​

Fairness matters when requests have very different lengths. One extremely long generation should not cause a queue of short interactive requests to wait indefinitely. Schedulers can use policies such as admission limits, age, request class, or token budget to balance throughput and responsiveness.

Cancellation also interacts with streaming. A client may disconnect after receiving enough output. The service should stop scheduling useless future work when safe to do so, release request-specific KV cache, and record a cancellation terminal reason.

Measure queue time separately​

Measure queue time separately from compute time. A faster model does not help much if requests spend most of their latency waiting for admission. Likewise, a larger batch that improves tokens per second may hurt interactive time to first token.

The practical goal is not maximum batch size. It is a scheduling policy that matches workload goals: interactive chat may prioritize first-token latency, while offline batch generation may prioritize total throughput and cost.

Predict

Traffic is low and the service waits for a full batch of 32 before running. What is the likely problem for interactive users?

Run the local Lab​

Run:

python3 labs/notebooks/level-14/l14-03-batching-streaming.py

The Lab simulates fixed arrivals under a maximum batch size and wait deadline.

  1. Run it unchanged with wait_deadline = 2. Record the batch count and longest queue wait.
  2. Before editing, predict the tradeoff if the scheduler may wait up to 10 ms: should requests form fewer/larger batches, and should the oldest request be allowed to wait longer?
  3. Change only wait_deadline = 2 to wait_deadline = 10, then rerun.
  4. Confirm the same arrivals now form fewer batches while the longest permitted queue wait increases. The workload did not change; only the batching deadline did.

Loading lab…

Write the core logic yourself​

Open:

labs/notebooks/level-14/l14-03-batching-streaming-exercise.py

Implement bounded batch formation using both maximum batch size and wait deadline. The supplied arrivals must split into two batches for the correct reason, not because of a hard-coded answer.

Run:

python3 labs/notebooks/level-14/l14-03-batching-streaming-exercise.py

The starter intentionally stops at TODO until you implement the missing logic. A correct solution reaches the final PASS: marker. Use the solved deterministic Lab as a comparison only after your own attempt.

Quick Check

1. What is continuous batching trying to improve?
2. What does output streaming most directly improve?
3. Why measure queue time separately?

0 of 3 questions answered.

Explain it back​

Compare an interactive chat workload with an offline document-generation workload. Explain how you would choose batch-wait time, maximum concurrency, and streaming behavior for each.

Key Takeaways

  • Batching trades waiting time for hardware efficiency.
  • Continuous batching dynamically mixes active requests.
  • Streaming targets time to first output rather than total compute alone.
  • Fairness and cancellation are scheduler responsibilities.
  • Measure queue and compute latency separately.

Next Lesson

Next, L14.4 — Model Loading and Memory accounts for where serving memory is actually spent.

References

Lesson actions

Completion is stored locally on this device.

View progress