SGF+ splits gradient roles for autoregressive video generation
An arXiv paper (arXiv:2610.10429) introduces SGF+ (Self Gradient Forcing Plus), which assigns separate parameters to the two roles in autoregressive video generation — writing key-value context for future predictions and denoising the current frames — while keeping them coupled through causal attention. The authors say both roles are jointly optimized with the original generation objective, without auxiliary losses, with context writing supervised through its contribution to future predictions, improving visual quality and long-horizon consistency over evaluated baselines in both framewise and chunkwise generation. They also report that a model trained on only 5-second rollouts can generate continuously for up to 24 hours without long-video fine-tuning.