Introduction
Most inference optimization is focused on improving the performance of individual kernels in models, such as attention, mixture of experts, matmuls and others. That work is extremely important and is a big part of how frameworks like vLLM, SGLang and TensorRT-LLM get impressive performance for LLM inference.
However, in order to squeeze the most out of your GPUs, it is not enough to simply optimize standalone kernels. In memory-bound regimes, most notably in LLM decode, a big chunk of performance is lost between kernel launches.
The inter-kernel overhead consists of three main parts:
- Kernel launch overhead, which is several microseconds before every kernel launch. This overhead can be hidden in compute-intensive workloads, but LLM decode has many short kernels, which can’t be overlapped with the launch.
- Wave quantization, which occurs when the number of tasks is not divisible by the number of SMs. Then, at the epilogue of that operation, some SMs have to stay idle due to lack of work.
- Lack of communication and computation overlap. That kind of overlap is, in general, one of the main ways to hide latency in kernel optimization. Overlapping one kernel’s computation with another’s loads and stores is very beneficial.
These three failure modes of the standard programming model motivate a new, different paradigm for kernel design, called megakernels. The next section will focus on explaining this mysterious and often misunderstood term.
What is a megakernel?
Many people in the AI performance engineering community have heard the term megakernel in the last year or so. Hazy Research Lab coined the term, and it has become one of those overloaded computer science terms (like the term “kernel” itself) where you need to explicitly state what you mean when you say “megakernel”. It has even been memed for how vague the term actually is:
In this blog post, we will explain parts of this iceberg, and which part of it Hoid takes.
Megakernel iceberg
Tip of the iceberg says that a large fused kernel is a megakernel. Kernel fusion is the process of joining two or more kernels together to avoid having to store and load the intermediate results back to global memory. This is an important AI compiler optimization, but it doesn’t encapsulate the core of the megakernel. A megakernel’s biggest advantage is removing the scheduling overhead of standard GPU kernels.
Layer 2 of the iceberg says that GPU-side scheduling is a megakernel. This states that the GPU has scheduling logic that decides which operation executes next, and not the host (CPU). While this often ends up being the case in efficient megakernels, it is only a single step toward making megakernels actually efficient.
Layer 3 says that a CUDA graph is a megakernel. CUDA graphs record the exact sequence of kernels that you want to execute and then record all the parameters for launching them. The next time the same graph should be executed, CUDA graphs just replay the sequence of kernels that was prerecorded. That way, they remove the kernel launch overhead before every kernel. They are very useful and widely used in the ecosystem, especially in LLM decode, when launch overhead accounts for a larger portion of the whole execution time.
Unfortunately, there are two main limitations of CUDA graphs:
- They reduce the kernel launch overhead, but still ~1 microsecond per kernel is needed for launching the kernel.
- They still enforce synchronization between two separate kernels, which means that all blocks of kernel 1 have to finish and before moving on to kernel 2. This results in wave quantization and doesn’t allow overlapping two consecutive kernel executions.
Layer 4 states that device-launched CUDA graphs are megakernels. As the name suggests, this allows the GPU to decide whether to launch a CUDA graph without host intervention. While this adds more flexibility to the CUDA graph execution, its main problems remain.
Layer 5 says that PDL (Programmatic Dependent Launch) is a megakernel. PDL enables overlapping part of the end of one kernel with the loading of static inputs of the next kernel. It is one step beyond CUDA graphs and helps reduce the inter-kernel overhead. It cannot, however, remove the boundaries between two kernels, and the blocks of the first kernel still have to finish before any block of the second kernel starts executing.
Layer 6 states that fine-grained overlap is a megakernel. This is the layer where Hoid’s megakernels stand! This means that you merge multiple kernels into a single megakernel in such a way that you schedule, dispatch, and manage all dependencies on device. There is no global synchronization between two different operations. Each tile can be executed at the exact moment its dependencies finish executing.
Additionally, managing dependencies manually gives you the ability to prefetch any inputs.
Layer 7. This layer is just for the meme!
Hoid’s approach to megakernels
Hoid is adopting the on-GPU interpreter megakernel model to manage fine-grained dependencies between operations. Unlike the traditional kernel programming model, where each kernel poses a synchronization boundary, the on-GPU interpreter treats one block as a single unit of work. That means that each block will only wait for its direct producer blocks to finish before launching. The standard kernel model forces every block of the latter kernel to wait for every block of the former kernel to finish. This is a simple way to avoid race conditions, but it comes at the cost of performance.
b) When we adopt the On-GPU interpreter model, one unit of work is one block. Therefore, a single block only has to wait for its producers to finish, which removes the idle time on SMs.
With this said, it is obvious that such a megakernel design is a big leap in terms of complexity compared to the standard kernel design. That leads to one of the key takeaways about megakernels: it is not hard to make a megakernel, but it is very hard to make a performant megakernel!
Two failure modes that are common in megakernels are scheduling overhead and inefficient block-level implementations of kernels.
The craft of making performant megakernels boils down to avoiding these two failure modes. Hoid takes an agent–compiler co-design approach for this purpose. AI compiler infrastructure provides scheduling primitives, and agents have the ability to write optimized block-level kernels. We believe this is the only viable way to generate performant megakernels at scale.
Results
The following figure shows a comparison of Hoid’s megakernel with vLLM, SGLang, and TensorRT-LLM. We are beating all the popular inference engines on a variety of batch sizes. We are also making our work open source! The code to run the vLLM server with Hoid’s megakernel and to reproduce the benchmarks is available at: https://github.com/hoid-ai/hoid-megakernel-qwen/.
This is just the beginning. We are aiming to build the first AI compiler that can consistently generate performant megakernels.