Large Language Model Data

Bingli Wang.

I work on data for large language models, spanning pretraining corpora, evaluation benchmarks, and the loops that turn model behavior into better training signals.

01 / About

Data is where my
research begins.

I am a PhD student in Computer Science at Southeast University. My research focuses on data for large language models.

I work across the full data lifecycle: constructing pretraining corpora, designing evaluation benchmarks, and turning evaluation signals into targeted data for the next training cycle.

A

Pretraining data

Improving corpus quality, mixture design, long-tail knowledge coverage, and multimodal alignment at scale.

B

Evaluation data

Designing benchmarks that measure knowledge, multimodal understanding, and agent capabilities with high signal.

C

Evaluation-to-training loops

Turning model failures, expert feedback, and evaluation traces into targeted data for the next training cycle.

02 / Education

Academic background.

2026 - Present

Southeast University

PhD in Computer Science and Technology

2022 - 2026

Sichuan Agricultural University

Undergraduate degree in Computer Science and Technology

03 / Selected work

Research, selected.

All publications

More papers, citation details, and the complete publication record.

View Google Scholar

04 / Experience

Research in practice.

May 2026 - Present

Ubiquant ยท IQust Foundation Model

Research Intern

Building evaluation-to-training loops for long-horizon agent capabilities.

Jul 2025 - May 2026

Shanghai AI Laboratory

Research Intern

Delivered multimodal caption data and built evaluation-to-training loops for multimodal understanding.

Aug 2024 - May 2025

M-A-P

Research Intern

Led large-scale benchmark data construction and quality control for graduate-level LLM evaluation.

Apr 2024 - May 2025

Tsinghua University (Shenzhen)

Research Intern

Conducted research on machine learning and model evaluation.