Google researchers — Chih-Wei Hsu, Moonkyung Ryu, and co-authors — published a post on the Google Research Blog on September 29, 2026, introducing Diffusion Controller (DiffCon): a framework that describes the control of image diffusion generators using the language of control theory. From a single formalism, practical fine-tuning algorithms are derived, and a small "side" network allows the model to be controlled with frozen weights, including closed models. On Stable Diffusion v1.4, the approach has already outperformed LoRA in two fine-tuning modes, but it remains a research result rather than a ready-made product tool.

image
image

What happened

The post on the Google Research Blog is accompanied by the paper "Diffusion Controller: Framework, Algorithms and Parameterization" on arXiv under number 2603.06981. The key idea is that reverse diffusion sampling is formulated as stochastic control in linearly solvable Markov processes (LS-MDPs): control reweights the transition kernels of the pre-trained model, balancing the target reward and an f-divergence penalty that prevents generation quality from degrading. From the optimality conditions, the authors derive two practical fine-tuning methods — a regularized policy gradient with a PPO rule and reward-weighted regression with a guarantee of preserving the KL-divergence minimizer. Controlling the model with a frozen backbone is enabled by a "side" auxiliary network that introduces corrections at intermediate denoising results, including gray-box access to closed models. In experiments on Stable Diffusion v1.4, the gray-box version of DiffCon outperformed LoRA in SFT and RWL modes on the HPS-v2 metric, affecting fewer internal layers, while the white-box version showed a 90% win rate against the base model.

Context

Until recently, methods for controlling diffusion generation have developed independently of each other: classifier-free guidance sets the strength of adherence to the prompt, LoRA adapts weights to a style or task, and the family of reward methods fine-tunes the model to a measurable reward. Each approach has its own mechanics, its own trade-offs between controllability and quality degradation, and its own tooling stack, making the comparison and combination of methods cumbersome. DiffCon proposes a common formalism in which these techniques become special cases of reweighting transition kernels, and transfers to diffusion generators the alignment methodology familiar from the RLHF stack of language models. A separate architectural rearrangement is important: controllability is transformed from a property of the weights into an external module that is placed on top of a frozen backbone and adjusts the denoising process at intermediate steps.

Why this matters for the industry

For the industry, DiffCon is a template for a product control layer: a single framework sets the reward, style constraints, and safety requirements, while the layer itself remains a lightweight add-on to the base model. For teams where LoRA "pulls" the model and breaks the original quality, this is a candidate for replacing part of the adaptation pipeline, and it can be tested right now: the paper and code framework are available, and the experiments are reproducible on Stable Diffusion v1.4. The gray-box route is particularly significant for services working with closed models: controllability and safety filters can be added without access to the weights, prototyping a sidecar on open backbones. In the next six months, the key test will be independent re-checks on modern architectures, such as Flux; if the superiority over LoRA is maintained, the first open implementations of controllers in inference frameworks and requests to API providers for access to intermediate denoising steps are likely. In the longer term, the same pattern, according to the Google Research Blog description, is applicable to video models, where reweighting transition kernels may become a standard formalism for controlling generation.

Why this matters for users

For practitioners working with Stable Diffusion, Flux, or Nano Banana, DiffCon describes an alternative to LoRA: a lightweight controller add-on on top of a frozen model with a single guidance strength parameter at inference, which, as stated in the Google post, can be smoothly adjusted without quality distortions. Instead of selecting and storing a set of adapters for each task, there is an external module that adjusts generation and can be removed and replaced without retraining the backbone. It is already possible to study the method independently: the paper on arXiv and the code framework are open, and the confirmed experiments are reproducible on Stable Diffusion v1.4, so comparing with LoRA on your own dataset and your own eval harness is a real first step. For those who only use closed APIs, it should be remembered that the gray-box route to them is not confirmed in the sources: today this is a direction for developers of open models and research teams.

What is still unknown / limitations

All confirmed results so far relate to Stable Diffusion v1.4 — a 2022 model, so transferability to modern and future backbones requires verification. The 90% win rate is measured against a simple base model, which is a low bar, and the comparison by HPS-v2 remains a proxy metric requiring manual vertical quality verification. The sources have no data on latency, pricing, and managed API, so it is impossible to judge production applicability. The statement about a single guidance strength parameter that is smoothly adjusted without distortions is taken from the description in the Google Research Blog post and is not a verified evaluation result. The gray-box route to closed services like Nano Banana is not confirmed in the sources, and the question of whether providers will open access to intermediate denoising steps remains open.

Sources

Author

Look at AI, editorial team