SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning

2026-08-24Artificial Intelligence

Artificial Intelligence
AI summary

The authors present Self-Reflective Policy Optimization (SRPO), a method that helps large language models learn better by letting them think back on their own outputs to find mistakes. Instead of needing extra rewards or bigger teacher models, SRPO uses the model's reflections to turn sparse feedback into detailed training signals for each part of the generated text. This approach improves performance on math problems and long tasks, using much less computational power than traditional methods. Their experiments show that SRPO works well on multiple benchmarks while being more efficient.

Self-reflectionCredit assignmentLarge Language ModelsPolicy optimizationToken-level supervisionReinforcement learningSparse feedbackOn-policy rolloutsMathematical reasoning benchmarksAgentic benchmarks
Authors
Jialong Liu, Yuling Shi, Ning Yang, Xiaodong Gu, Zuchao Li
Abstract
Self-reflection is a powerful mechanism for credit assignment in human learning, converting sparse outcome feedback into actionable guidance. However, its potential for post-training Large Language Models (LLMs) remains underexplored. We propose Self-Reflective Policy Optimization (SRPO), a framework that internalizes this capability. SRPO enables LLMs to analyze their own completed trajectories, synthesize errors into concise "reflection patches," and use reflection-conditioned teacher scores on student on-policy rollouts as dense token-level training signals. This process effectively transforms sparse terminal supervision into dense, token-level learning signals without requiring external critics, separate reward models, or larger teacher models. We demonstrate that SRPO achieves state-of-the-art performance across mathematical reasoning and long-horizon agentic benchmarks with exceptional data efficiency. Using a Qwen3-8B base model, SRPO attains 73.3% on AIME'24 using only 8% (0.08x) of the training FLOPs required by scaled supervised fine-tuning, while significantly improving success rates on WebShop (64.7%), ALFWorld (76.8%), and SWE-Bench-Lite (31.2%). Code is available at https://github.com/Galleons2029/SRPO