GRPO-Guard — Mitigating Implicit Over-Optimization via Regulated Clipping¶
| Field | Value |
|---|---|
| arXiv | 2510.22319 |
| Submitted | 2025-10-25 (revised 2025-10-30) |
| Venue | — (preprint) |
| Authors | Jing Wang, Jiajun Liang, Jie Liu, Henglin Liu, Gongye Liu, Jun Zheng, Wanyuan Pang, Ao Ma, Zhenyu Xie, Xintao Wang, Meng Wang, Pengfei Wan, Xiaodan Liang |
| GitHub | https://github.com/yifan123/flow_grpo (same repo as Flow-GRPO; scripts/multi_node/sd3_grpo_guard.sh) · project page |
| Paradigm | Policy Gradient — modifier applied inside the GRPO objective; no change to sampling |
| Cites | FlowGRPO (2505.05470), DanceGRPO (2505.07818), PPO, GRPO (DeepSeek-R1) |
| Cited by | FlowDPPO, UniGRPO |
Context¶
GRPO-Guard is a modifier applied to the training objective of any policy-gradient GRPO method (FlowGRPO, DanceGRPO, MixGRPO). It changes nothing about how images are sampled; it only changes how the importance ratio is computed and weighted inside the loss. The core issue it addresses is that PPO's clip mechanism — designed around the assumption \(\mathbb{E}[\rho_t] \approx 1\) — breaks down in flow-matching GRPO, causing reward hacking. This file builds on the GRPO objective from flow_grpo.md.
Problem 1 — The importance ratio is systematically left-shifted; PPO clipping does not engage¶
Issue: In flow-matching GRPO, empirical measurement shows that the mean of \(\rho_t^{(i)}\) — the raw per-step importance ratio for sample \(i\) at denoising step \(t\), i.e. the current-policy transition density over the old-policy density, \(\rho_t^{(i)} = \pi_\theta(x_{t-\Delta t}^{(i)} \mid x_t^{(i)}, c) / \pi_{\theta_\text{old}}(x_{t-\Delta t}^{(i)} \mid x_t^{(i)}, c)\) — is systematically below 1 across all timesteps. PPO clipping is designed to catch ratios outside \([1-\epsilon, 1+\epsilon]\) (here \(\epsilon\) is the small clip half-width, typically \(\sim 10^{-4}\)), assuming the ratio is centred near 1. When the mean is left-shifted (e.g., mean \(\approx 0.7\)), most positive-advantage samples already satisfy \(\rho_t^{(i)} < 1-\epsilon\), so the clip is never activated — the update is unconstrained. This allows the policy to take unbounded steps on high-reward samples, driving reward hacking: proxy reward keeps climbing but image quality degrades.
Why it happens: For a Gaussian policy at step \(t\), the log importance ratio is:
where \(\mu_\theta = \mu_\theta(x_t, t, c)\) and \(\mu_{\theta_\text{old}}\) are the Gaussian transition means under the current and old policy, \(\sigma_t\) is the per-step diffusion (noise) scale, and \(\Delta t\) is the discretisation step size of the reverse SDE. When \(\theta\) has been updated in a direction that increases reward, \(\mu_\theta\) moves away from \(\mu_{\theta_\text{old}}\), and the numerator tends to be negative for randomly-drawn \(x_{t-\Delta t}^{(i)}\) (the sample is "left behind" the shifting mean). This structurally left-shifts \(\rho_t\).
Idea — RatioNorm: Standardise the log importance ratio within the group at each timestep, removing the systematic bias and scale. Define the normalised log ratio:
Here \(\hat\rho_t^{(i)}\) is the RatioNorm-normalised ratio: the raw \(\rho_t^{(i)}\) after subtracting the within-group mean and dividing by the within-group standard deviation of \(\log\rho_t\) at step \(t\), where \(\mathbb{E}_i\) and \(\mathrm{std}_i\) average over the \(N_g\) group members and \(\delta > 0\) is a small numerical-stability constant. Using the Gaussian structure of the flow policy, this simplifies to a closed form involving only the dot product of the per-step mean shift with the injected noise:
where \(\Delta\mu_t = \mu_{\theta_\text{old}}(x_t, t, c) - \mu_\theta(x_t, t, c)\) is the per-step shift between old and current transition means and \(\epsilon_t^{(i)} \sim \mathcal{N}(0, I)\) is the noise injected at step \(t\) for sample \(i\).
Why this works: After RatioNorm, \(\mathbb{E}[\log\hat\rho_t] \approx 0\) and \(\mathrm{std}(\log\hat\rho_t) \approx 1\) for all \(t\). This matches PPO's design assumption — the clip band \([1-\epsilon, 1+\epsilon]\) now activates as intended, constraining large updates when the policy has drifted far from the old policy.
Result: RatioNorm restores a balanced, step-consistent ratio (Fig. 2) so the clip engages — the payoff shows in the "Gold score" (true quality, vs proxy reward): on SD3.5-M it rises while the proxy holds — GenEval Gold 0.84 → 0.89, TextRender Gold 0.88 → 0.99, PickScore Gold 1.16 → 1.20; on FLUX.1-dev GenEval Gold 0.88 → 1.02 (Tab. 1) — i.e. much less over-optimization at equal-or-better proxy reward.
Problem 2 — Gradient magnitude varies ~20× across timesteps; late steps dominate¶
Issue: Even after RatioNorm, the gradient magnitude contributed by each timestep is not uniform. For a Gaussian policy, the gradient of \(\log\rho_t\) with respect to \(\theta\) scales as \(1/(\sigma_t^2\Delta t)\). Steps with small \(\sigma_t\) (late denoising, low noise) produce large gradients; steps with large \(\sigma_t\) (early denoising, high noise) produce small ones. The ~20× variation means late timesteps dominate training, causing the model to over-optimise the final appearance while under-optimising the coarse structure.
Idea — Gradient reweighting: Multiply each timestep's loss contribution by a per-timestep scalar \(\delta_t\) (the per-step gradient reweight, a fixed function of \(t\)) that equalises gradient magnitudes:
where \(\Delta t\) is the reverse-SDE step size and, in the DDPM case, \(\beta_t = 1 + \eta^2(1-t)/(2t)\) is the noise-schedule-dependent variance factor (\(\eta\) being the SDE stochasticity level). Why this works: \(\delta_t\) scales each step's contribution inversely proportional to its variance, restoring approximately uniform gradient magnitudes across timesteps.
Result: Gradient-magnitude variation across timesteps drops from ~20× to ~2.5× (Fig. 3), so coarse-structure (high-noise) steps are no longer drowned out by fine-detail steps — this is what makes the Gold-score gains in Problem 1 stable rather than transient.
Training Objective¶
where \(N_g\) is the group size (number of images sampled per prompt, written \(G\) in the paper), \(T\) is the number of denoising steps, \(\hat A^{(i)}\) is the group-relative advantage of sample \(i\) — its reward standardised within the group, \(\hat A^{(i)} = (r^{(i)} - \mathrm{mean}(\lbrace r^{(j)}\rbrace)) / \mathrm{std}(\lbrace r^{(j)}\rbrace)\) — and \(\epsilon\) is the PPO clip half-width. Compared to plain FlowGRPO: \(\rho_t \to \hat\rho_t\) (RatioNorm) and an extra \(\delta_t\) factor. Sampling, advantage estimation, and KL penalty are unchanged.
Reference Implementation (VeRL-Omni)¶
Condensed from GRPOGuardLoss in diffusion_algos.py (@register_diffusion_loss("grpo_guard")). It is identical to FlowGRPOLoss except the four # <<< lines — it adds a reverse-SDE mean-drift bias projected onto the per-step scale sqrt_dt·σ_t (RatioNorm), then rescales the loss by 1/sqrt_dt² so gradient magnitude is consistent across timesteps (\(\delta_t = 1/\Delta t\)). The extra inputs prev_sample_mean, old_prev_sample_mean, std_dev_t, sqrt_dt come from the rollout:
@register_diffusion_loss("grpo_guard")
def loss_grpo_guard(old_lp, lp, adv, mean_θ, mean_old, std_t, sqrt_dt, cfg):
c = cfg.diffusion_loss
adv = clamp(adv, -c.adv_clip_max, c.adv_clip_max) # same as FlowGRPO
scale = sqrt_dt.mean() * std_t.mean() # <<< shared per-step scalar
mean_diff_sq = ((mean_θ - mean_old) ** 2).mean(non_batch_dims) # <<< reverse-SDE mean drift
ratio_bias = mean_diff_sq / (2 * scale ** 2) # <<< RatioNorm bias
ratio = exp((lp - old_lp + ratio_bias) * scale) # <<< FlowGRPO: exp(lp - old_lp)
unclipped = -adv * ratio # ┐
clipped = -adv * clamp(ratio, 1 - c.clip_ratio, 1 + c.clip_ratio) # ├ identical PPO-clip body
return mean(max(unclipped, clipped)) / sqrt_dt.mean() ** 2 # <<< FlowGRPO: mean(...); here ÷Δt
Effect on Training Stability¶
| Metric | FlowGRPO baseline | With GRPO-Guard |
|---|---|---|
| Mean of \(\rho_t\) | \(<1\) (left-shifted) | \(\approx 1\) (centred) |
| Variance spread across timesteps | \(\sim 20\times\) | \(\sim 2.5\times\) |
| PPO clip activation | Rarely | As designed |
| Reward hacking onset | Early (reward climbs, quality drops) | Substantially delayed |
Compatibility¶
GRPO-Guard is a pure objective modifier — it is composable with all policy-gradient methods:
- FlowGRPO + Guard ✓
- DanceGRPO + Guard ✓
- MixGRPO + Guard ✓ (apply RatioNorm and \(\delta_t\) to window steps only)
- CPS + Guard ✓ (CPS handles sampling quality; Guard handles objective stability)
- UniGRPO adopts RatioNorm from this work ✓
Limitations¶
- RatioNorm uses per-group statistics (group size \(N_g\)) rather than population statistics — the normalisation is an approximation that degrades for small groups (\(N_g < 4\)).
- Gradient reweighting \(\delta_t\) is a fixed schedule, not adaptive to the current model's gradient landscape.
- Does not address SDE noise artifacts (→ CPS) or computational cost (→ MixGRPO).