AI workloads & kernels
Custom GPU kernels, fused operations, and better memory use. Built for your model.
AI Performance / Numaro
Kernels, compilers, and specialized languages. We optimize the software around your workload and hardware.
Your performance target
Across the stack
From one operation
to a whole new toolchain.
Custom GPU kernels, fused operations, and better memory use. Built for your model.
Scheduling, code generation, and workload mapping for GPUs and custom accelerators. Target stacks include MLIR/IREE, TVM, and LLVM, with hardware or simulator evaluation.
A language and compiler built around your domain. A deeper collaboration, from design to execution.
Our team guides the autoresearch engine and reviews the results.
How it worksFrom our optimization work
2–4.5×
Selected small / medium matrices · batch ≥ 512
Hatched area: variation across tested shapes.
Selected small and medium matrices, batch ≥ 512, versus torch.bmm / cuBLAS. Outputs matched the baseline bit for bit in the reported tests.
Measurements on one NVIDIA L40S in the report’s stated environment, using paired, interleaved timing on a shared GPU. Results depend on workload, hardware, and baseline; a pilot measures your stack and the end-to-end impact.
Read the GPU benchmarks ↗1.55×
Fused 3×3 + bias + ReLU6 · FP16 geometric mean
FP16 geometric mean for depthwise 3×3 convolution, bias, and ReLU6 versus cuDNN convolution followed by activation. Bit-exact in the reported tests.
Measurements on one NVIDIA L40S in the report’s stated environment, using paired, interleaved timing on a shared GPU. Results depend on workload, hardware, and baseline; a pilot measures your stack and the end-to-end impact.
Read the GPU benchmarks ↗2.93×
80 C-reference benchmarks · geometric mean
41 wins · 7 ties · 32 losses. 1.17× when excluding Numaro runs under 1 ms.
Geometric-mean execution speedup across 80 C-reference benchmarks versus Clang 21 -O2, using Numaro’s experimental compiler. The programs express equivalent computations in different source languages and match the benchmark checksums. This measures execution time, not compilation time or Rust performance.
Recorded August 8, 2026 on an x86-64 CPU. Timings include process startup, with a warmup followed by the minimum of five runs pinned to one core. This is a historical benchmark snapshot, not a new timing run.
Against Clang 21 -O2: 41 wins, 7 ties within 2%, and 32 losses. The median runtime is 0.9733× the baseline. Large algorithmic gains and short-running cases have a substantial effect on the 2.93× geometric mean.
Restricting the comparison to the 66 cases where Numaro’s execution takes at least 1 ms gives a 1.17× geometric-mean speedup. Across all 80 cases, the geometric mean versus Clang -O3 -march=native is 2.48×.
The custom compiler is not a drop-in C or Rust compiler. These results do not establish faster compilation, a universal runtime advantage, or performance on a custom accelerator.
View the benchmark data ↗The people behind the work
Work directly with us,
from the first benchmark to handover.
ML & autonomous research
Vadim guides workload selection, ML evaluation, and the research campaign. His work spans generative models and efficient inference.
Programming languages & compilers
Philipp guides language and compiler design, optimization strategy, and technical review. His research connects programming language theory to native code generation.
Work with Numaro
Your engineers and our team choose a workload, establish the baseline, and agree on performance and correctness targets.
Our team directs the campaign on your hardware or simulator, reviews candidates, and checks the results with you.
We deliver evaluated code and reproducible benchmarks, walk your engineers through the changes, and guide integration.
Targets follow baseline review. Simulated results are labeled.
contact@numaro.tech