Pretraining data
Improving corpus quality, mixture design, long-tail knowledge coverage, and multimodal alignment at scale.
Large Language Model Data
I work on data for large language models, spanning pretraining corpora, evaluation benchmarks, and the loops that turn model behavior into better training signals.
01 / About
I am a PhD student in Computer Science at Southeast University. My research focuses on data for large language models.
I work across the full data lifecycle: constructing pretraining corpora, designing evaluation benchmarks, and turning evaluation signals into targeted data for the next training cycle.
Improving corpus quality, mixture design, long-tail knowledge coverage, and multimodal alignment at scale.
Designing benchmarks that measure knowledge, multimodal understanding, and agent capabilities with high signal.
Turning model failures, expert feedback, and evaluation traces into targeted data for the next training cycle.
02 / Education
2026 - Present
PhD in Computer Science and Technology
2022 - 2026
Undergraduate degree in Computer Science and Technology
03 / Selected work
A 26K-question benchmark built through expert annotation and automated quality control.
Paper
Contributed to long-tail pretraining data and the multimodal caption data pipeline.
Paper
A benchmark for image-text alignment in naturally interleaved multimodal contexts.
PaperMore papers, citation details, and the complete publication record.
View Google Scholar04 / Experience
May 2026 - Present
Research Intern
Building evaluation-to-training loops for long-horizon agent capabilities.
Jul 2025 - May 2026
Research Intern
Delivered multimodal caption data and built evaluation-to-training loops for multimodal understanding.
Aug 2024 - May 2025
Research Intern
Led large-scale benchmark data construction and quality control for graduate-level LLM evaluation.
Apr 2024 - May 2025
Research Intern
Conducted research on machine learning and model evaluation.