MachGen has published the source code for Open-VC Attention — open FP8 attention kernels for NVIDIA Blackwell accelerators (B200 and B300), built on the FlashAttention-4 CuTe DSL core and implementing the ExpCast technique from the recent VC-Attention paper, under the BSD-3-Clause license. On a real capture of a denoising step from the MiniMax-H3 video model, the kernel runs nearly twice as fast as BF16 FlashAttention-4, while the numerical error remains controlled. For teams generating long video on Blackwell, this is a ready-made way to reduce the cost of the most expensive stage of generation, and for the industry — a clear example of how a fresh paper turns into open production code in a matter of weeks.


What happened
MachGen has open-sourced Open-VC Attention: these are FP8 attention kernels for NVIDIA Blackwell (B200 and B300), built on the FlashAttention-4 CuTe DSL core and implementing the ExpCast technique from the VC-Attention paper (arXiv 2609.15810, published 14.09.2026). The code is distributed under the BSD-3-Clause license. Benchmarks were run on a real BF16 capture of a single denoising step of the MiniMax-H3 video model — sequence length 73397, 56 attention heads, head dim 128. Compared to BF16 FlashAttention-4 (flash-attn-4 4.0.0b33), the kernel delivers a 2.00–2.03× speedup, and outpaces the paper's author implementation by roughly 1.6×; the relative L2 error is 2.98–3.19%. Instead of the paper's proposed V-Smooth technique — online k-means clustering over V values — the open version uses "V residual repair": quantization residuals are applied only to the worst-energy V tokens; in a 15-second MiniMax-H3 clip, repair was needed for roughly 0.5% of tokens, without which the character's face behind the visor collapsed. The repository includes benchmark methodology and a smoke check: open-vc-attn-check --smoke.
Context
To understand the significance of the release, it is worth recalling how attention is computed on Blackwell. Attention is the most expensive part of video generation by diffusion transformers, and FP8 acceleration of tensor cores on these chips is undermined by the slow MUFU special-function unit: its throughput is about 2.53 PFLOP/s, roughly half the FP8 peak of the B200, and it is precisely through this unit that standard kernels run softmax exponentials. The ExpCast technique removes this bottleneck: FP8 softmax probabilities are computed with a single FMA instruction without calling the exponential function. However, it is not just the paper's idea: the open code's advantage is explained by CuTe DSL engineering — a warp-specialized pipeline, mid-window maximum search that reduces the number of rescales, packed transposed V for tensor cores, and input preparation fused into a CUDA graph (0.86 ms, about 1.5% overhead). For comparison, the original paper claimed a 1.47–1.6× method win. The targeted repair of the worst-energy V tokens, replacing the paper's mechanism, turns quantization error from a fixed cost into a tunable budget — an approach closer to production practice than online clustering.
Why this matters for the industry
The industry significance is that this is the first open production path to such video-diffusion attention acceleration on Blackwell: on the most expensive stage of generation — long video denoising — computations become roughly twice as cheap. More telling is the format itself: a low-bit technique from a fresh paper turned into open, reproducible BSD-3-Clause code with documented SGLang integration in a matter of weeks, rather than remaining a private advantage of one team — compute moats around custom attention kernels are quickly eroding. Over the coming months, an absorption trajectory is likely: the ExpCast and error-budget management techniques may be picked up by serving stacks and kernel competitors like FlashAttention-4 and cuDNN; expect comparisons on a broader set of models and the first ports beyond video attention. If independent reproductions confirm the claimed numbers, FP8 attention with a tunable error budget could become a default option in video serving, and within a couple of years, low-bit attention with adjustable precision risks becoming the inference default for video and long context — at which point competition will shift from the kernels themselves to error-budget management and product quality.
Why this matters for users
Teams with a B200 or B300 profile and long video denoising can apply the release immediately: MHA with head dim 128 and sequence length from ~32768. The kernel is installed via pip — Linux, Python 3.10+, PyTorch with CUDA 13, B200 requires SM100 architecture, B300 requires SM103 — and is called as open_vc_attn.attention(q, k, v) with BF16 inputs. The acceptable error level is set by the V residual repair budget: the repository example uses budget=0.005. A practical pilot sequence: install the package, run the attached smoke check, then benchmark your own model using the published methodology. The benefit is real only when the profile matches; teams outside this corridor are not directly suited to the kernel today, but the open code with methodology allows dissecting the ExpCast and targeted V-value repair techniques for their own experiments.
What is still unknown / limitations
The "error as a knob" formulation at the product level is premature: the budget=0.005 value is a V residual repair parameter tuned on a single 15-second clip from one model, and the relationship between repair budget and visual quality has not been validated on other models, prompts, and scenes. The claimed relative L2 error is a numerical metric that does not replace quality checks on your own data until perceptual validation. The claimed speedup numbers have not yet been confirmed by independent reproductions. Finally, applicability is limited to the narrow hardware and workload profile described above, so measured benefit is not guaranteed outside this profile.
Sources
- Open-VC Attention: Open-Source, Error-Tunable VC-Attention Variant — MachGen blog
- MachGen/open-vc-attention — open FP8 implementation of VC-Attention for Blackwell (GitHub, BSD-3-Clause)
- VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention (arXiv 2609.15810)
Author
Look at AI, editorial team
