hetero-servepaged attention
served
0
ttft p50
—
cache hit
0%
migrations
0
kv moved
0 MB
Run the kernels GitHub

Send a prompt twice.
Watch the second one skip the work.

A KV-cache-aware LLM serving scheduler — paged cache, prefix sharing, continuous batching, and a cost model that decides whether to move a cache across the network or recompute it. Running here, in your browser.

peak bandwidth used
55.4%
vs PyTorch SDPA
10–22×
CUDA kernels
3
tests
47 + 22

Send a request

t 0.00s
Try in order

Cluster

Interconnect · KV migrations
idle

Requests

newest first
Send a prompt — then send another that starts the same way.

What is actually running: the real scheduling logic — 16-token paged blocks, chain-hashed prefix matching, refcounts and LRU eviction, continuous batching, and the migrate-vs-recompute cost model priced against a token-bucket link. Step costs are timings measured on real hardware (T4 prefill 0.59 ms/token, decode 52 ms/step). What is not here is GPT-2 — a 124M-parameter model will not run in a browser tab, so generated text is a stand-in. The Python system this mirrors runs the real model on real CUDA; the kernels link above compiles and benchmarks it on a free GPU.

measured · tesla t4 · 320 gb/s peak

The kernels underneath

Profiling the decode step showed a third of it was the host gathering KV blocks into contiguous tensors — overhead invented by paging the cache. Three kernel generations later, each written in response to a measurement rather than a hunch:

pathper callvs gatherGB/s% of peak

batch 32 · context 2048 · 201 MB of KV per call. PyTorch's SDPA — cuDNN and FlashAttention underneath — barely beats a naive einsum, because both still pay the gather. That is the case for a paged kernel in one line: a fast dense kernel does not help once the cache is paged, because it cannot read a block table. Nsight then showed v2 was occupancy-starved rather than bandwidth-starved — 0.1 waves across 40 SMs — which is what produced v3's context split.