DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models

Abstract

DataFlex is a unified framework for data-centric dynamic LLM training. Built on LLaMA-Factory, it integrates sample selection, domain-mixture adjustment, and sample reweighting behind extensible trainer abstractions with support for large-scale distributed training.

Publication
arXiv preprint arXiv:2603.26164
Xiaochen Ma
Xiaochen Ma
Ph.D. Student in Computer Science and Engineering

I work on data infrastructure and high-performance distributed systems for large-scale LLM data preparation, with a focus on Ray and Apache Spark.