KerneLab
Agentic kernel engineering

Expert kernels,
for any stack.

The fastest kernel for your model probably exists already. It is written for someone else's accelerator, in someone else's kernel language. KerneLab builds agentic machinery that ports, tunes and verifies kernels across both, with numerical correctness as a gate and measured runtime as the only score.

candidate vs referencecapturing
LayerDistance from referenceState
L0Skeleton
L1Inventory
L2Structure
L3Order
L4FP flavor
L5Machine
Schematic of a convergence run. Not measured data.
The problem

The kernel layer does not travel.

A kernel is the routine that actually runs on the accelerator: a matrix multiply, an attention block, a mixture-of-experts layer. The best ones are written by a small number of experts, for one chip, in one kernel language. Everything around them is moving.

01

Kernels are tied to a language

CuTe DSL, Pallas, Triton, hand-written CUDA. A kernel written in one does not drop into a stack built on another, however good it is.

02

Every chip changes the rules

Each accelerator generation brings new instructions, new memory, new ways to schedule work. Last generation's tuning is this generation's starting point at best.

03

Models arrive faster than experts

New models and new numeric formats land every few months. The people who can write a tensor-core pipeline by hand do not multiply at that rate.

04

Compilers leave a gap

General-purpose compilers do a great deal, and the last stretch to an expert kernel is still done by hand, one workload at a time.

The approach

Treat kernel engineering as a search that has to prove itself.

Our agents do what an expert kernel engineer does: read the reference, write a candidate, compare the machine code, change one thing, measure again. They do it inside a harness that will not let them fool themselves, or us.

  • Runtime is the objective. Measured on the device, not predicted.
  • Numerics are a gate. A candidate that fails the numerics check is never timed.
  • Noise is measured. An improvement smaller than the run-to-run noise is not an improvement.
  • Every gain has a cause. One change per edit, so each improvement is attributable.

Read the technical deep-dive →

  1. CaptureLower the reference kernel and record everything about how it runs: its code, its launch configuration, its device timings, its outputs.
  2. ProposeAgents write a candidate in the target kernel language, then edit it one discrepancy at a time. Parameters first, rewrites after.
  3. VerifyCheck the candidate against a high-precision oracle. Only if it passes is it timed on the real device.
  4. ConvergeKeep what is faster by more than the noise. Stop when the candidate is inside the reference's own noise floor.
Where this leads

One workload, one kernel.

Once kernels can be rewritten and re-verified cheaply, the boundaries between them are no longer fixed. A model's forward pass is a long chain of library kernels, and every seam costs a launch and often a round trip through memory.

A megakernel fuses that chain into a single kernel shaped around one application. We have already seen what this is worth in miniature: fusing three stages of the MoE path removed a large intermediate tensor, and gave one of the biggest end-to-end gains of that project. Timed in isolation, the same change looked marginal.

Why boundaries cost what they cost →

forward passkernel library
kernel boundaries 21 stalls so far 0
Where we are

A new effort, working in the open as it can.

KerneLab is just getting started. The machinery exists and has been run end to end on real kernels and real hardware. What comes next is more kernels, more kernel languages and more accelerators.

We publish what we can defend. This site describes the method, the work and the open questions in that spirit.

Now

Kernels on Blackwell

Ports between kernel languages and from TPU to GPU, converged and verified by agents.

Next

Working upstream

Contributing what the ports teach us back to the kernel languages we target.

Then

Opening up

Repositories will appear on GitHub as we open up our work.

Want the detail?

The technology page walks through the convergence loop, the fingerprint layers and the acceptance rule. The research page lists what we do not know yet.