Projects
Inference research, benchmarks, and kernels.

Zero-Free Conditioning for Ideogram 4
Jul 2026, OptimizationIdeogram 4 rebuilds a 53,248-wide conditioning tensor that is mostly zeros at every denoising step. Keeping only the real text rows saves 9.46 GiB at 2048² and is bitwise identical.
Memory-Efficient LM-Head Inference for LLaDA
Jul 2026, KernelsFused CUDA and Triton kernels combine the LM-head projection with a streaming logsumexp reduction, cutting peak memory 4.5× without ever materializing the logits tensor.