Plain language

What this result means

This is the most directly practical result in the ledger. It connects research output to actual runtime: when a model is bottlenecked by small batched matrices, fine-grained MoE, depthwise convolution, or skinny int8 decode, a kernel that fills the GPU or removes a memory pass can matter immediately.

  • The same report also records losses. Large square GEMM remains cuBLAS territory, and a generic W4A16 kernel is much slower than tinygemm.
  • The kernel speedups are measured as ratios under paired, interleaved timing, because the GPU was shared and absolute timings were noisy.
  • The op-count records and the kernel records are separate. Fewer arithmetic operations did not automatically make a faster GPU kernel.

Visual notes

How to read the result

Dataflow comparison of separate convolution and activation kernels versus one fused kernel, showing the intermediate global-memory write and read removed by fusion.
Remove the intermediate memory tripThe stock depthwise path writes a temporary activation tensor to global memory and reads it back for ReLU6. The fused kernel keeps that intermediate on chip. This schematic shows activation-tensor traffic; weights, bias and cache effects are omitted.Open full-size figure ↗
Two heatmaps of measured FP16 and FP32 depthwise speedups, arranged by batch size, channel count and spatial resolution.
Every measured shape in the locked gridFP16 and FP32 speedups for fused depthwise 3×3 + bias + ReLU6 against cuDNN convolution followed by activation. All 12 locked cases per precision are shown. Empty cells were not in this grid; both precision variants matched the reference bit for bit in the reported tests.Open full-size figure ↗

From kernel to workload

How much of the work can you accelerate?

A kernel’s speedup matters most when it covers a large share of the original runtime.

Explore the Amdahl model

20%
4.0×
1.18×modeled whole-workload speedup
Other workOptimized workTime saved
S=1(1−f)+f/sS=\frac{1}{(1-f)+f/s}

f is the original time share; s is the kernel speedup. This idealized model holds other work fixed and adds no new overhead.

Selected measured results · L40S

2.85×Fine-grained MoE layerE128 · K1024 · grouped GEMM
1.23×MobileNetV3 blockC72 · 56×56 · fused depthwise
1.07×Transformer blockSwin · S64 · batched matmul

These are measured examples from the report, each with its own workload and kernel. They are separate from the model above.

Result table

Measured L40S speedups for small batched GEMM, fused depthwise conv, int8 decode, MoE, and GQA decode.

CellBaselineNumaroDeltaNote
Batched matmultorch.bmm2-4.5xfastersmall/medium matrices, batch >= 512
Fused depthwise 3x3cuDNN + activation1.55x fp16 / 1.31x fp32fasterbit-exact
int8 W8A8torch._int_mmup to 3.7xfasterskinny-M decode
MoE grouped GEMMper-expert loop2.5-3.6x end-to-endfasterfine-grained MoE
W4A16tinygemm~12x slowerlosskept as a boundary condition

Method

How it was found

Each kernel targets a specific production gap: not enough occupancy, an avoidable memory pass, a missing primitive, or a launch-heavy loop.

  • Mapped where cuBLAS/cuDNN/PyTorch under-filled the GPU or forced extra memory traffic.
  • Wrote narrow Triton kernels for those exact regimes instead of trying to beat vendor libraries everywhere.
  • Timed stock and custom kernels back-to-back under a GPU timing lock.
  • Kept losses in the report so the boundary of the method is visible.

Verification

How it was checked

Each subdirectory has a verifier that checks correctness before speed. Some kernels are bit-identical; others are compared against a higher-precision or stock reference with an explicit tolerance where reduction order differs.

Scope

What is not being claimed

All numbers are on one L40S under the stated environment. They are not claims for every GPU, every shape, or cross-library SOTA. Large square GEMM and 4-bit weight-only decode are explicitly not beaten.

References

Baseline sources

Citation

How to cite

Numaro AI Autoresearch Team. "Bit-exact GPU kernels in regimes vendor libraries leave open." Numaro Research Report NUMARO-2026-003, 2026.

@techreport{numaro2026FasterMlKernels,
  title = {Bit-exact GPU kernels in regimes vendor libraries leave open},
  author = {Numaro AI Autoresearch Team},
  institution = {Numaro},
  number = {NUMARO-2026-003},
  year = {2026},
  url = {https://numaro.tech/research/faster-ml-kernels-2026/}
}