Python DSL → optimized IR → fused CUDA

Attention variants,
compiled—not hand-written.

Compose causal masks, sliding windows, GQA, softcaps, and ALiBi. attnc lowers the composition into one kernel and proves its semantics against a NumPy oracle.

160differential cases
2Python frontends
3tile states
1fused forward kernel
01 / LIVE EXPLORER

Forge a kernel.

This browser demo mirrors attnc's lowering rules. Flip a feature and inspect the compiler artifacts change together.


          

          

          
        
STATIC TILE PLAN
query tiles ↓KV tiles →
fast pathpredicatedskipped

Causal + window bounds become a closed-form key interval. Fully masked tiles issue no memory traffic.

02 / COMPILER PIPELINE

Small compiler. Real compiler.

Both frontends converge on one IR, then pass through conservative optimizations before runtime dispatch.

01Capture

Immutable combinators or proxy-object tracing record score and mask semantics.

score - slopes[h] * Δ
02Optimize

Constant folding, boolean simplification, bound inference, and dead-branch removal.

causal ∧ window(4096)
03Classify

Every key tile becomes fully masked, fully unmasked, or partial.

skip · fast · predicate
04Emit + JIT

Splice expressions into an online-softmax skeleton, compile with NVRTC, and cache.

hash(IR, shape, arch)
✓
Correctness is a first-class backend

The same IR runs through an independent NumPy interpreter, enabling differential fuzzing across variants, shapes, and seeds.

160 / 160CPU differential cases passing