Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions

Abstract

ReLIFT interleaves reinforcement learning with online fine-tuning on questions beyond a model’s current capabilities, combining RL’s strength in refining existing reasoning with supervised learning’s ability to introduce new knowledge and reasoning patterns.

Publication
International Conference on Learning Representations 2026
Xiaochen Ma
Xiaochen Ma
Ph.D. Student in Computer Science and Engineering

I work on data infrastructure and high-performance distributed systems for large-scale LLM data preparation, with a focus on Ray and Apache Spark.