# Papers & Publications

## FLUX: Data Worth Training On — A Preprocessing Pipeline for Large Language Model Training

**Authors:** Gowtham, Sai Rupesh, Sanjay Kumar, Saravanan, Venkata Chaithanya  
**Published:** March 2026

**Abstract**  
FLUX is a preprocessing pipeline designed to improve the quality of large-scale web datasets used for training language models. The pipeline maximises token retention while maintaining strong filtering standards during dataset construction [More »](/content/Research/FLUX-Data/index.html)

[View on arXiv](https://arxiv.org/abs/2603.13972)

## Blu-WERP (Web Extraction and Refinement Pipeline): A Scalable Pipeline for Preprocessing Large Language Model Datasets

**Authors:** Gowtham, Sai Rupesh, Sanjay Kumar, Saravanan, Venkata Chaithanya  
**Published:** November 2025

**Abstract**  
Blubridge is proudly presenting the process behind "Blu-WERP", our pipeline that is setting a new industry standard for scalable, high-quality LLM pretraining data this month. In our paper, we are demonstrating training and evaluation details, including the data preparation pipeline, from JusText extraction to Benchmark-targeted classification... [More »](/content/Research/Blu-Werp/index.html)

[View on arXiv](https://arxiv.org/abs/2511.18054)
