🖼️ Qwen Suggests the Next Image Edits — and Verifies They Are Feasible
Alibaba's Qwen Business Unit, together with Southeast University and ShanghaiTech University, published What to Edit Next (arXiv:2608.07565): a three-stage framework that, in an image-generation chat, suggests the next edits verified against the current image. The Qwen3-VL-8B model was trained on 100,000 Qwen App dialogues and 173,071 click pairs.
🌍 Recipe for multimodal assistants: an ordered source-target verifier yielded 0.6% false rejections versus 22.2% for a single-pass VLM. RL based on clicks increased CTR but did not eliminate visually infeasible edits.
👤 In a 14-day A/B test with millions of users, the CTR of recommendations increased by 32.70%, and the inconsistency of suggestions dropped from 3.7% to 0.9%. The paper is open on arXiv; there is no code or weights — this is a methodology reference.
Source 1: https://what-to-edit-next.github.io/ Source 2: https://arxiv.org/abs/2608.07565
