From Notebook to Service
Goal
Turn notebook inference code into a small shared service by separating model loading, request validation, request-local state, health checks, configuration, and deployment identity.
Start with the notebook that worked for one person
You run a notebook, load a model, type one prompt, and get the expected answer. Then a teammate asks to use the same model from another application.
The first version might simply run the notebook code every time a request arrives. That works for a demo, but two problems appear quickly:
- loading the model again for every request wastes time and memory;
- a remote caller can send values the notebook author never would, such as a missing prompt or an enormous
max_tokens.
A service keeps the useful inference function but adds the surrounding jobs needed for many callers.
service starts
→ load model once
→ verify required files/device resources
→ mark service ready
request A → validate → inference → response
request B → validate → inference → response
The model lifecycle belongs to service startup rather than to each request. Request handling then works with an already loaded runtime.
Validation belongs before inference. If prompt must be text and max_tokens must be a positive number below a known limit, reject a bad request there. A clear client error is easier to fix than an exception deep inside the model runtime.
This is the first difference between “my notebook works” and “other software can depend on this.” The model computation may be unchanged, but the application now has to manage shared resources and unpredictable callers.
Separate liveness, readiness, and request state
Health has more than one meaning. Liveness asks whether the process is running. Readiness asks whether it is actually prepared to serve traffic. A process can be alive while the model is still loading or while an essential dependency is unavailable. Routing traffic before readiness is true creates avoidable failures.
Concurrency changes state ownership. Two requests may reach the same process at nearly the same time. Request-specific generation state should not leak between them. Shared read-only model weights are useful; mutable request buffers need their own identity and lifecycle.
Carry deployment identity into operations
A service should expose deployment identity with operational evidence. Useful fields include application version, model identifier, model revision or digest where available, runtime version, and configuration hash. When an output changes after a release, those fields make it possible to compare the exact serving environment.
Containers help package the runtime environment, libraries, and service code. Docker documentation distinguishes image tags from content-addressed digests; a digest can identify exact image content rather than only a movable human-readable label. That makes deployment evidence stronger when reproducing an incident.
Keep secrets and policy outside model code
Do not put secrets or policy inside the model code simply because the notebook had convenient global variables. Credentials, request authentication, rate limits, and environment configuration belong in explicit service or deployment boundaries. The inference function should receive already validated inputs.
The first production design goal is therefore not maximum speed. It is a clean boundary: one model lifecycle, explicit requests, explicit response/error shapes, safe request-local state, health checks, and deployment identity. Once those exist, later performance work has something reliable to optimize.
Predict
Run the local Lab
Run:
python3 labs/notebooks/level-14/l14-01-service-boundary.py
The Lab separates process liveness from traffic readiness.
- Run it unchanged. Confirm
liveness: True,readiness: True, andsend user traffic? yes. - In
state, find"model_loaded": True. Before editing, predict which health signal should change if the process stays alive but the model is not loaded. - Change only
"model_loaded": Trueto"model_loaded": False, then rerun. - Confirm liveness stays
True, readiness becomesFalse, and the service stays out of the load balancer.
Loading lab…
Quick Check
Explain it back
Turn a notebook with load_model() and generate(prompt) into a service design. State what belongs in startup, request validation, request-local state, health, and deployment metadata.
Key Takeaways
- A service separates model lifecycle from request handling.
- Validate untrusted client inputs at the boundary.
- Liveness and readiness should answer different questions.
- Concurrent requests need isolated request state.
- Record enough deployment identity to reproduce behavior.
Next Lesson
Next, L14.2 — Inference APIs turns the service boundary into a predictable client contract.
References
- Docker, Documentation.
- vLLM, Documentation.
Completion is stored locally on this device.