What we know, and what we do not.
Agent-written kernels are a young field with a few striking results and at least one public retraction. This page sets out the questions we think decide whether the approach works, the notes we have taken from real ports, and the work we are building on.
Six problems worth being wrong about in public
Can the verifier resist the optimizer?
Search finds whatever the checker fails to forbid. The field has seen reported speedups shrink or vanish once contaminated tasks and gamed checks were removed. Which checks stay sound when the thing being checked is trying to pass?
What does "correct" mean at low precision?
Some kernels can be bit-exact. Others pass through a re-quantization step where a one-unit difference upstream becomes a whole step downstream. Tolerances have to be derived from the format, then defended.
When does fusion pay?
Published megakernel results are strongest for small models at low batch and smaller at high throughput. We want a predictive account of where boundaries dominate, before spending a search on it.
What can a kernel language not say?
Kernel languages differ in what they can express. Is there a systematic way to map the differences between two of them before a port begins?
Is machine-code distance a useful signal?
Fingerprint distance fell alongside runtime in one convergence and rose slightly after a real improvement in another. When does it guide a search and when does it mislead one?
How far does autonomy go?
Others report that agent-assisted kernels still beat fully autonomous ones. Where exactly does human direction still earn its place in the loop, and can the harness absorb it?
Propose, verify, converge, in forty generations
A synthetic landscape, not our data. Each cell is a candidate strategy; brighter is faster. Three explorers sample at random each round and seven refiners work near the best verified candidate. Some candidates are fast and wrong. Watch what the gate does with them.
A conceptual strategy space. Verified candidates survive the gate; incorrect candidates are discarded.
Filled cells passed verification; crosses failed it and are discarded however fast they looked. The outlined cell is the best verified candidate. The landscape has a second, shallower basin, so a run can settle in the wrong place: reset and run again to see it happen.
Things real kernels taught us
Qualitative findings from the ports described in the case studies. Each one changed how the loop works.
The launch record carries what the code does not
Whether a kernel is persistent, and how many threads a block runs, are not in its instruction listing. Reading disassembly alone missed both for one reference. Capturing a real launch found them.
Isolated benchmarks mislead
A fused kernel timed against the two it replaced looked barely worthwhile. End to end it was one of the largest gains of the project, because the intermediate tensor between the two never shows up when they are timed back to back.
A configuration is a package
We reconstructed a reference's exact tile shape and pipeline depth and ran it. It was slower. A reference's settings work as a set; copied piecemeal, they do not carry the benefit across.
The obvious lead can be wrong
The candidate issued several times as many block-wide barriers as the reference. It looked like the answer. Removing them changed nothing measurable. We keep the record because the next reader will suspect the same thing.
Idioms do not survive translation
A standard activation function did not reach the hardware's exponential unit. Rewriting it around an explicit base-two exponential was bit-identical and markedly cheaper.
Hardware finds what interpreters cannot
Allocating exactly every register on a processor left a kernel waiting forever at full power. Leaving slack fixed it. Nothing short of the real device showed the problem.
Work we build on, and work that keeps us honest
None of this is ours. It is the context any claim in this area should be read against.
Megakernels
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1BHazy Research · 2025A hand-built single kernel for a whole forward pass. The existence proof.
- A Compiler and Runtime for Mega-Kernelizing Tensor Programs (MPK)OSDI · 2026A compiler that generates megakernels automatically, with no language model in the loop.
- Compiling Models to MegakernelsLuminal · 2026Search-based compilation of a forward pass into one kernel.
- Megakernels Considered Harmful: Wavefront Path Tracing on GPUsHPG · 2013Where the term comes from, and the standing reminder that fusion is not free.
Agents and search for kernels
- KernelBench: Can LLMs Write Efficient GPU Kernels?ICML · 2025The benchmark that framed the problem as propose, check and time.
- Learning to Optimize Halide with Tree Search and Random ProgramsSIGGRAPH · 2019Search plus a learned cost model, beating expert schedules on average.
- Ansor: Generating High-Performance Tensor Programs for Deep LearningOSDI · 2020Hierarchical search spaces and evolutionary search for tensor programs.
Verification, and what happens without it
- Towards Robust Agentic CUDA Kernel Benchmarking, Verification, and OptimizationSakana AI · 2025Written after an earlier system exploited its own evaluation. Documents how models game kernel benchmarks.
- Measuring Automated Kernel EngineeringMETR · 2025Benchmark hygiene as half the work, including a result removed as implausibly fast.
- Kernel ContractsarXiv · 2026A specification language for what a kernel must compute, with tolerances and a reference oracle.
The substrate
- Pallas Mosaic GPUJAXJAX's kernel-authoring layer for NVIDIA GPUs. The target of both ports so far.
- Splash attention for TPUJAXThe block-sparse flash attention kernel our GPU port takes as its reference.
- FlashInferOpen sourceThe kernel library whose CuTe DSL mixture-of-experts path was the reference for our first port.
- Getting Started with CUDA GraphsNVIDIA · 2019Measured per-kernel launch costs, and what batching launches recovers.
Nothing yet.
We have internal reports on each port. We intend to publish them, with reproductions, as the work opens up.
Every claim carries its conditions.
A result is published with its reference, its hardware and software versions, its correctness check and a script that reproduces it.