Project 10
An LLM inference server built from scratch to understand why vLLM is fast. The goal was not a vLLM clone but a runtime where every optimization is switchable and measured against a baseline on a real GPU, including the ones that turn out not to help.
The server owns its decode loop and layers features one phase at a time: continuous batching, batched prefill, a block-based paged KV cache, chunked prefill, a prefix cache, int8 weights, speculative decoding and CUDA-graph decode. One engine-loop thread owns the batcher and bridges to async HTTP through a queue per request, so the API streams over SSE with cancellation and admission control.
Measured on a rented RTX 4090 with Qwen2.5-0.5B in bf16. At 8 req/s, continuous batching cut median time-to-first-token from 0.63s to 0.03s and more than doubled SLO goodput, while being roughly equal to static batching at saturation.
Profiling then showed decode is CPU launch-bound, so CUDA-graph replay made each step 2.2-6.3x faster with identical logits. The paged cache served 25-40% more tokens per second and held 44-60% more concurrent sequences in the same KV memory. A prefix cache gave 2.2x lower TTFT at a 6K shared prefix.
Several things did not pay. int8 weights saved 36% of weight memory but ran 11-17% slower than bf16. Chunked prefill halved the worst inter-token gap at 8K prompts (113ms to 51ms) but cost 3-15% throughput and long-prompt latency. A 0.5B draft model for speculative decoding did not pay (0.6-1.1x), though prompt-lookup hit 3.1-4.6x on copy-heavy answers. vLLM on the same GPU was 8-17x faster over HTTP, and the measured reason is the same launch overhead; that comparison predates the CUDA-graph work and has not been rerun. Stress testing also caught a real bug where one over-long request with CUDA graphs on failed every in-flight request. It is fixed and covered by tests.
What I Learned
Tech Stack
Links