CAST: Causal Advantage-Structured Training with Spatially Grounded Compositional Rewards for Diffusion Models

Shu Yu1,2,3, Chaochao Lu1,†
1Shanghai Artificial Intelligence Laboratory  2Shanghai Innovation Institute  3Fudan University
†Corresponding author.

Visual Results

Additional qualitative comparisons on FLUX.2-dev (part I)
Additional qualitative comparisons on FLUX.2-dev (part I). GenEval 2 prompts with multi-object counts, attributes, and chained spatial relations. GenEval 2 score reported below each image.
Additional qualitative comparisons on FLUX.2-dev (part II)
Additional qualitative comparisons on FLUX.2-dev (part II). GenEval 2 score reported below each image.
Compositional failures in baselines and CAST's improvement
Compositional failures in baselines and CAST's improvement. GenEval 2 prompts: (a) "five striped monkeys in front of three pink cows to the left of four clocks." (b) "a plastic horse under a glass raccoon." GenEval 2 score reported below each image. Only CAST satisfies all verifiable-atoms in both prompts.

Abstract

Online reinforcement learning has been successfully extended to flow matching for diffusion model (DM) image generation. However, this paradigm faces three limitations: (1) Window selection. Existing methods typically manually set the stochastic differential equation (SDE) sampling window, i.e., the denoising steps where exploration noise is injected. We instead determine it from each model's denoising trajectory. (2) Reward saturation. Current methods rely heavily on scoring models trained on human annotations; we find that such scores are extremely high and nearly indistinguishable on the latest SOTA open-source DMs, making advantage estimation largely ineffective. (3) Sample inefficiency. A single scalar reward collapses different failure modes into almost identical scores, leaving minimal gradient guidance for targeted improvement.

To address these issues, we propose CAST (Causal Advantage-Structured Training), an RL fine-tuning method for pretrained DMs, which (1) identifies the denoising step at which each model fixes the objects and their spatial arrangement in the image and uses that timing to set the SDE window, (2) decomposes each prompt via Causal Scene Graphs (CSG) into verifiable-atoms, i.e., minimal semantic units such as an object, count, attribute, or spatial relation that can each be checked independently, and rewards each atom separately, and (3) projects the signed atom-level advantages into pixel space through teacher-forced attention and uses them to spatially weight the SDE policy objective. We fine-tune two of the strongest open-source DMs, FLUX.2-dev and Qwen-Image-2512, with CAST, and evaluate them on GenEval 2, a compositional benchmark on which these models still fail frequently, and on Qwen-Image-Bench for overall quality. Within almost the same training budget, CAST's improvement over the base model on the most challenging GenEval 2 prompts is up to 3.07× that of Flow-GRPO, while overall generation quality also improves.

Motivation

We identify two properties of modern diffusion models that motivate CAST's design:

Image structure forms in early denoising steps
Image structure forms in early denoising steps. Left: intermediate decoded images at five timesteps for three diffusion models. Dashed borders mark the knee point where CLIP similarity plateaus. Right: CLIP-ViT-L/14 similarity between Tweedie estimates and the text prompt. The knee point (triangle) occurs before τ = 0.25 across all models, indicating that the main layout and composition are established early, so the SDE exploration window should be set from each model's own trajectory.
PickScore saturates and weakly tracks structured correctness
Preference scores saturate and weakly track structured correctness. (a) Score distributions pooled across the four DMs: PickScore concentrates near its ceiling, whereas GenEval 2 retains a broad dynamic range. (b) The contrast persists within every model's generations. (c) Preference agreement between PickScore and GenEval 2 is only slightly above the 50% random baseline (dashed line).

Method: CAST Overview

Overview of CAST
Overview of CAST. A Causal Scene Graph (CSG) decomposes the prompt into K verifiable-atoms. Three representative types of facts are shown: attributes (red), counts (blue), and spatial relations (green). A frozen VLM uses separate passes to score each atom and extract its teacher-forced attention map. The resulting atom advantages and maps are combined into a signed spatial map that weights the SDE policy objective.

Main Results

Within the same 2,048-image and 128-update budget and under the same reward signal, the Hard Case gain of CAST over the base model is 1.93× that of scalar-reward Flow-GRPO on FLUX.2-dev and 3.07× on Qwen-Image-2512. CAST also surpasses the high-selectivity distillation upper bound SFT (3/64) on both backbones, whose Hard Case gain amounts to only 37% and 58% of CAST's despite its 32× larger image budget, while attaining the highest Qwen-Image-Bench Overall on both backbones.

(a) Final performance

Final held-out performance on GenEval 2 and Qwen-Image-Bench. G2 Hard-50 is the GenEval 2 Overall score on the 50-prompt Hard Case subset of unsaturated evaluation prompts. Values are mean ± sample standard deviation over four matched generation-seed replicates.

Backbone Method G2 Hard-50 ↑ G2 Overall ↑ G2 Count ↑ G2 Position ↑ QIB Overall ↑
FLUX.2-devBase65.72 ± 1.8383.16 ± 0.2066.15 ± 0.6073.29 ± 0.4252.84 ± 0.27
SFT (4/8)67.28 ± 2.0683.31 ± 0.2466.37 ± 0.6073.46 ± 0.5152.76 ± 0.25
SFT (3/64)69.76 ± 2.4883.93 ± 0.4067.27 ± 0.6074.15 ± 0.8852.42 ± 0.16
Flow-GRPO (Orig.)64.26 ± 1.7681.85 ± 0.2264.18 ± 0.4472.71 ± 1.1353.34 ± 0.28
Flow-GRPO (G2)71.37 ± 1.9084.54 ± 0.1168.39 ± 0.3274.98 ± 0.6252.73 ± 0.06
CAST76.64 ± 2.3985.95 ± 0.1070.85 ± 0.2976.90 ± 0.5353.44 ± 0.29
Qwen-Image-2512Base67.79 ± 1.8178.02 ± 0.5667.14 ± 1.5959.60 ± 0.8450.68 ± 0.23
SFT (4/8)68.55 ± 1.7978.27 ± 0.5467.64 ± 1.5559.62 ± 0.7850.45 ± 0.24
SFT (3/64)70.12 ± 1.8479.29 ± 0.4569.64 ± 1.3759.72 ± 0.5349.55 ± 0.29
Flow-GRPO (Orig.)68.60 ± 1.7577.90 ± 0.4467.15 ± 1.0759.92 ± 0.9150.11 ± 0.26
Flow-GRPO (G2)69.11 ± 1.5679.59 ± 0.2270.48 ± 0.6258.59 ± 0.2550.70 ± 0.15
CAST71.84 ± 1.8580.81 ± 0.2372.35 ± 0.5560.45 ± 0.4351.32 ± 0.15

(b) Training cost

Raw training cost within each backbone. All generated and scored candidates, including filtered or discarded candidates, are counted.

Method Images ↓ Updates ↓ FLUX H200-h ↓ Qwen H200-h ↓
SFT (4/8)8,19238495.0090.24
SFT (3/64)65,536288749.51723.70
Flow-GRPO (Orig.)2,04812828.9128.75
Flow-GRPO (G2)2,04812832.4332.21
CAST2,04812832.9032.51

Ablation

Ablation of atom-level credit assignment and spatial weighting on FLUX.2-dev. All variants use the same training prompts, SDE window, generated-image budget, and optimization settings. Values are mean ± sample standard deviation over four matched generation-seed replicates.

Variant Advantage Spatial map G2 Overall ↑ QIB Overall ↑
Sum-reward baselineper-imagenone84.54 ± 0.1152.73 ± 0.06
Per-atom advantageper-atomnone85.12 ± 0.1452.98 ± 0.18
Spatial weightingper-imageaggregated attention84.91 ± 0.1653.06 ± 0.20
Shuffled correspondenceper-atomshuffled atom maps85.00 ± 0.1552.89 ± 0.22
CASTper-atommatched atom maps85.95 ± 0.1053.44 ± 0.29

Attention-Head Dynamics

Attention-head sensitivity across denoising steps
Attention-head sensitivity across denoising steps. Each heatmap cell averages a 4×4 group of transformer layers and attention heads. The side panels show the layer-wise mean sensitivity at each step, and the dashed line separates the dual-stream layers (0–7) from the single-stream layers (8–55). Spatial relation is most sensitive at the early step and then declines; attribute becomes broadly sensitive from step 7 onward; count is most clearly separated at step 7. The most sensitive heads overlap substantially across the three aspects (count, attribute, and spatial relation), indicating temporal differences rather than a one-to-one assignment between aspects and attention heads.

Training Dynamics

Training dynamics on FLUX.2-dev
Training dynamics on FLUX.2-dev. Left: SFT self-distillation loss under three learning rates; the reported setting is emphasized, and all runs plateau within the budget. Right: mean Soft-TIFA reward on the shared RL manifest; the CAST curve is an atom-averaged diagnostic rather than its optimized objective, and the aesthetic reward (right axis) is on a separate scale.

BibTeX

If you find our work useful in your research, please cite:

@article{yu2026cast,
      title={CAST: Causal Advantage-Structured Training with Spatially Grounded Compositional Rewards for Diffusion Models},
      author={Shu Yu and Chaochao Lu},
      year={2026},
      journal={arXiv preprint arXiv:2609.39441},
}