KerneLab
Research

What we know, and what we do not.

Agent-written kernels are a young field with a few striking results and at least one public retraction. This page sets out the questions we think decide whether the approach works, the notes we have taken from real ports, and the work we are building on.

FIG. 07A space of candidates
Exploring a space of candidate kernelsA conceptual grid contains untested candidates, cyan verified candidates, crossed-out rejected candidates and an amber incumbent. It illustrates search and verification, not benchmark results. CANDIDATE STRATEGIES○ UNTESTED■ VERIFIED× REJECTED07 / EXPLORE. CHECK. KEEP WHAT SURVIVES.
Conceptual search space; not measured data. Amber marks an incumbent.
Open questions

Six problems worth being wrong about in public

Q1

Can the verifier resist the optimizer?

Search finds whatever the checker fails to forbid. The field has seen reported speedups shrink or vanish once contaminated tasks and gamed checks were removed. Which checks stay sound when the thing being checked is trying to pass?

Q2

What does "correct" mean at low precision?

Some kernels can be bit-exact. Others pass through a re-quantization step where a one-unit difference upstream becomes a whole step downstream. Tolerances have to be derived from the format, then defended.

Q3

When does fusion pay?

Published megakernel results are strongest for small models at low batch and smaller at high throughput. We want a predictive account of where boundaries dominate, before spending a search on it.

Q4

What can a kernel language not say?

Kernel languages differ in what they can express. Is there a systematic way to map the differences between two of them before a port begins?

Q5

Is machine-code distance a useful signal?

Fingerprint distance fell alongside runtime in one convergence and rose slightly after a real improvement in another. When does it guide a search and when does it mislead one?

Q6

How far does autonomy go?

Others report that agent-assisted kernels still beat fully autonomous ones. Where exactly does human direction still earn its place in the loop, and can the harness absorb it?

A toy version

Propose, verify, converge, in forty generations

A synthetic landscape, not our data. Each cell is a candidate strategy; brighter is faster. Three explorers sample at random each round and seven refiners work near the best verified candidate. Some candidates are fast and wrong. Watch what the gate does with them.

toy strategy space · 36 × 14 candidates
Exploring a space of candidate kernelsA conceptual grid contains untested candidates, cyan verified candidates, crossed-out rejected candidates and an amber incumbent. It illustrates search and verification, not benchmark results. CANDIDATE STRATEGIES○ UNTESTED■ VERIFIED× REJECTED07 / EXPLORE. CHECK. KEEP WHAT SURVIVES.

A conceptual strategy space. Verified candidates survive the gate; incorrect candidates are discarded.

Generation0
Evaluated0
Rejected by gate0
Best verified–

Filled cells passed verification; crosses failed it and are discarded however fast they looked. The outlined cell is the best verified candidate. The landscape has a second, shallower basin, so a run can settle in the wrong place: reset and run again to see it happen.

Field notes

Things real kernels taught us

Qualitative findings from the ports described in the case studies. Each one changed how the loop works.

N1

The launch record carries what the code does not

Whether a kernel is persistent, and how many threads a block runs, are not in its instruction listing. Reading disassembly alone missed both for one reference. Capturing a real launch found them.

Separate kernels and a fused regionThree separate operations have intermediate memory boundaries. A fused region removes those depicted boundaries; real fusion may introduce resource and scheduling costs. SEPARATE KERNELSFUSED REGION MEMORYMEMORY Operation AOperation BOperation CABC
Illustrative fusion schematic
N2

Isolated benchmarks mislead

A fused kernel timed against the two it replaced looked barely worthwhile. End to end it was one of the largest gains of the project, because the intermediate tensor between the two never shows up when they are timed back to back.

Six layers of a kernel fingerprintAn exploded conceptual diagram of L0 skeleton, L1 inventory, L2 structure, L3 order, L4 floating-point flavor and L5 machine code. These are comparison layers, not physical chip layers. L0 / SKELETONL1 / INVENTORYL2 / STRUCTUREL3 / ORDERL4 / FP FLAVORL5 / MACHINE03 / COARSE STRUCTURE → FINE DETAIL
Conceptual configuration layers
N3

A configuration is a package

We reconstructed a reference's exact tile shape and pipeline depth and ran it. It was slower. A reference's settings work as a set; copied piecemeal, they do not carry the benefit across.

N4

The obvious lead can be wrong

The candidate issued several times as many block-wide barriers as the reference. It looked like the answer. Removing them changed nothing measurable. We keep the record because the next reader will suspect the same thing.

N5

Idioms do not survive translation

A standard activation function did not reach the hardware's exponential unit. Rewriting it around an explicit base-two exponential was bit-identical and markedly cheaper.

N6

Hardware finds what interpreters cannot

Allocating exactly every register on a processor left a kernel waiting forever at full power. Leaving slack fixed it. Nothing short of the real device showed the problem.

Prior art

Work we build on, and work that keeps us honest

None of this is ours. It is the context any claim in this area should be read against.

Megakernels

Agents and search for kernels

Verification, and what happens without it

The substrate

  • Pallas Mosaic GPUJAXJAX's kernel-authoring layer for NVIDIA GPUs. The target of both ports so far.
  • Splash attention for TPUJAXThe block-sparse flash attention kernel our GPU port takes as its reference.
  • FlashInferOpen sourceThe kernel library whose CuTe DSL mixture-of-experts path was the reference for our first port.
  • Getting Started with CUDA GraphsNVIDIA · 2019Measured per-kernel launch costs, and what batching launches recovers.
Publications

Nothing yet.

We have internal reports on each port. We intend to publish them, with reproductions, as the work opens up.

Standing rule

Every claim carries its conditions.

A result is published with its reference, its hardware and software versions, its correctness check and a script that reproduces it.