Send a request
t 0.00sCluster
Requests
newest firstWhat is actually running: the real scheduling logic — 16-token paged blocks, chain-hashed prefix matching, refcounts and LRU eviction, continuous batching, and the migrate-vs-recompute cost model priced against a token-bucket link. Step costs are timings measured on real hardware (T4 prefill 0.59 ms/token, decode 52 ms/step). What is not here is GPT-2 — a 124M-parameter model will not run in a browser tab, so generated text is a stand-in. The Python system this mirrors runs the real model on real CUDA; the kernels link above compiles and benchmarks it on a free GPU.
The kernels underneath
Profiling the decode step showed a third of it was the host gathering KV blocks into contiguous tensors — overhead invented by paging the cache. Three kernel generations later, each written in response to a measurement rather than a hunch:
| path | per call | vs gather | GB/s | % of peak |
|---|
batch 32 · context 2048 · 201 MB of KV per call. PyTorch's SDPA — cuDNN and FlashAttention underneath — barely beats a naive einsum, because both still pay the gather. That is the case for a paged kernel in one line: a fast dense kernel does not help once the cache is paged, because it cannot read a block table. Nsight then showed v2 was occupancy-starved rather than bandwidth-starved — 0.1 waves across 40 SMs — which is what produced v3's context split.