The IST-DASLab team (Andrei Panferov, Maximilian Kleinegger, Sweta Priyadarshi, Tijmen Blankevoort, Dan Alistarh) in the paper “Disaggregated Quantization: Specializing LLM Prefill and Decode” (arXiv 2609.26333, September 22, 2026) proposes processing LLM prefill and decode with different quantizations of the same model and jointly fine-tuning both branches via the QADD method. This phase-specific format split recovers a significant portion of the accuracy lost during aggressive weight compression and does not require changing an existing decode checkpoint. The approach is already reproducible on top of popular inference stacks, and ready-made checkpoints are available for download.



What happened
The paper describes the QADD method: an SFT-token mask determines which phase-specific linear path processes each token, so prefill and decode of the same model live in different quantizations and are fine-tuned jointly. The dedicated NVFP4 prefiller — a compute-native format for Blackwell GPUs — is trained on top of a frozen decoder: on Qwen3.8-27B with an Unsloth IQ1_S GGUF decoder, accuracy rises on MMLU-Pro from 29.04% to 61.54% and on MMMU-Pro from 24.39% to 59.65%, while the decode checkpoint remains unchanged. The offloaded disaggregated prefill (ODP) scheme streams an additional 12.8 GiB prefill checkpoint from SSD through a ring buffer carved out of the output head memory, so no extra VRAM is needed at all. In a custom llama.cpp integration on DGX Spark, TTFT on an 8K context drops from 12.27 to 6.90 seconds; the method is enabled with a single flag in llama.cpp and validated in disaggregated vLLM serving.
Context
LLM inference is structurally asymmetric, and the method targets exactly this asymmetry. Prefill is a compute-bound phase that benefits from compute-native formats like NVFP4, designed for Blackwell GPU tensors; decode is a memory-bound phase where ultra-low-bit GGUF weights win by saving memory. The default practice was different: one quantization for the entire run, so aggressive weight compression hurt quality both on long prompts and on generation. Fine-tuning via QADD recovers roughly half of the accuracy lost in such compression, without rewriting existing decoders — essentially a continuation of the industry-normalized disaggregation of roles in serving, pushed down to the level of weight formats.
Why this matters for the industry
Inference is no longer “one quantization for the entire run”: compute-bound prefill stays in hardware-native NVFP4, and memory-bound decode in ultra-low-bit GGUF, and this is a reproducible layer on top of existing stacks, not theory. The numbers are also industrial: at 2-bit decode, full disaggregation exceeds weight-only baselines by 7.1 and 4.5 points (Qwen 3 and Gemma 3) on decode-heavy tasks and by 12.6 and 8.9 points on prefill-heavy RULER. The economics of local inference change: quality close to a full-size model on long prompts is achieved without increased VRAM and without modifying existing decoders, and PTQ disaggregation has already been confirmed on models up to 2.8 trillion parameters. If the line holds, “quantization” as a single deployment parameter will break down into phase-specific profiles: compute-native formats for compute-bound phases, ultra-low-bit weights for memory-bound.
Why this matters for users
On a single consumer Blackwell GPU (RTX 50-series or DGX Spark GB10), you can already today download ready-made 1–2-bit GGUF Qwen3.8-27B and connect an NVFP4 prefiller from Hugging Face, enabling the --odp-blocks flag in llama.cpp. GPU memory stays at the level of the decode checkpoint, accuracy on long prompts noticeably improves relative to pure ultra-low-bit decode, and TTFT on an 8K context drops by roughly half. Nuances honestly disclosed in the model card: release llama.cpp kernels give about 1.3× speedup, and short prompts remain slower due to loading the prefill checkpoint from SSD. The minimal day-one scenario is a local assistant working with long documents on desktop hardware.
What is still unknown / limitations
There is no comparison with full-size Qwen3.8-27B in the provided data, so the phrasing “almost like the full-size model” is an extrapolation; the measured gain is relative to the Unsloth IQ1_S decoder. The claimed speedup range of 1.38–1.78× applies to 4K–32K contexts, with 1.78× achieved only on a not-yet-released FlashInfer integration, while release llama.cpp kernels give about 1.3×; the current llama.cpp integration is custom, not mainstream. The scheme depends on the speed and stability of the SSD subsystem, and short prompts remain slower due to streaming the prefill checkpoint. Full disaggregation with QADD has so far been shown on Qwen 3 and Gemma 3, and ready-made prefillers cover a limited set of GGUF decoders (including IQ1_S and IQ1_M), so extension to other model families is a future question.
Sources
- Disaggregated Quantization: Specializing LLM Prefill and Decode (arXiv 2609.26333)
- IST-DASLab/disaggregated-quantization — official codebase (GitHub)
- ISTA-DASLab/Qwen3.8-27B-NVFP4-prefiller — prefillers for GGUF decoders (Hugging Face)
Author
Look at AI, editorial team
