Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
Lu Ma, Hao Liang, Meiyi Qiang, Lexiang Tang, Xiaochen Ma, Zhen Hao Wong, Junbo Niu, Chengyu Shen, Runming He, Bin Cui, Wentao Zhang
June, 2025
Abstract
ReLIFT interleaves reinforcement learning with online fine-tuning on questions beyond a model’s current capabilities, combining RL’s strength in refining existing reasoning with supervised learning’s ability to introduce new knowledge and reasoning patterns.
Publication
International Conference on Learning Representations 2026

Ph.D. Student in Computer Science and Engineering
I work on data infrastructure and high-performance distributed systems for large-scale LLM data preparation, with a focus on Ray and Apache Spark.