Expert kernels,
for any stack.
The fastest kernel for your model probably exists already. It is written for someone else's accelerator, in someone else's kernel language. KerneLab builds agentic machinery that ports, tunes and verifies kernels across both, with numerical correctness as a gate and measured runtime as the only score.
The kernel layer does not travel.
A kernel is the routine that actually runs on the accelerator: a matrix multiply, an attention block, a mixture-of-experts layer. The best ones are written by a small number of experts, for one chip, in one kernel language. Everything around them is moving.
Kernels are tied to a language
CuTe DSL, Pallas, Triton, hand-written CUDA. A kernel written in one does not drop into a stack built on another, however good it is.
Every chip changes the rules
Each accelerator generation brings new instructions, new memory, new ways to schedule work. Last generation's tuning is this generation's starting point at best.
Models arrive faster than experts
New models and new numeric formats land every few months. The people who can write a tensor-core pipeline by hand do not multiply at that rate.
Compilers leave a gap
General-purpose compilers do a great deal, and the last stretch to an expert kernel is still done by hand, one workload at a time.
Treat kernel engineering as a search that has to prove itself.
Our agents do what an expert kernel engineer does: read the reference, write a candidate, compare the machine code, change one thing, measure again. They do it inside a harness that will not let them fool themselves, or us.
- Runtime is the objective. Measured on the device, not predicted.
- Numerics are a gate. A candidate that fails the numerics check is never timed.
- Noise is measured. An improvement smaller than the run-to-run noise is not an improvement.
- Every gain has a cause. One change per edit, so each improvement is attributable.
- CaptureLower the reference kernel and record everything about how it runs: its code, its launch configuration, its device timings, its outputs.
- ProposeAgents write a candidate in the target kernel language, then edit it one discrepancy at a time. Parameters first, rewrites after.
- VerifyCheck the candidate against a high-precision oracle. Only if it passes is it timed on the real device.
- ConvergeKeep what is faster by more than the noise. Stop when the candidate is inside the reference's own noise floor.
Two kernels, two directions of travel.
Both were carried out almost entirely by agents working inside our convergence harness, and both ran on real Blackwell hardware. We describe what was done and what we learned; we will add figures as we are ready to publish them.
A mixture-of-experts kernel, from CuTe DSL to Pallas
FlashInfer's masked MoE path, rebuilt in JAX's Pallas Mosaic GPU for Blackwell with its 4-bit number format intact. Its core multiply is bit-exact against a high-precision oracle.
Splash attention, from TPU to NVIDIA GPU
JAX's block-sparse flash attention existed only for TPU. Now it has a Blackwell implementation with forward and backward kernels, every mask type, and three tiers of verification behind it.
One workload, one kernel.
Once kernels can be rewritten and re-verified cheaply, the boundaries between them are no longer fixed. A model's forward pass is a long chain of library kernels, and every seam costs a launch and often a round trip through memory.
A megakernel fuses that chain into a single kernel shaped around one application. We have already seen what this is worth in miniature: fusing three stages of the MoE path removed a large intermediate tensor, and gave one of the biggest end-to-end gains of that project. Timed in isolation, the same change looked marginal.
A new effort, working in the open as it can.
KerneLab is just getting started. The machinery exists and has been run end to end on real kernels and real hardware. What comes next is more kernels, more kernel languages and more accelerators.
We publish what we can defend. This site describes the method, the work and the open questions in that spirit.
Kernels on Blackwell
Ports between kernel languages and from TPU to GPU, converged and verified by agents.
Working upstream
Contributing what the ports teach us back to the kernel languages we target.
Want the detail?
The technology page walks through the convergence loop, the fingerprint layers and the acceptance rule. The research page lists what we do not know yet.