Ru Peng (Perry)
"Only love endures the passage of time"
Iโm a final-year PhD student at Computer Science Department of Zhejiang University (ZJU), advised by Professors Junbo Zhao and Gang Chen. Iโm affiliated with DiLab-ZJU and the State Key Laboratory of Blockchain and Data Security. Early on, I was fortunate that Professors Tianyong Hao and Kehai Chen ushered me into research. I have interned at three LLM labs: Tencent Hunyuan (long-horizon agents), Ant Group Inclusion AI (data synthesis & RL from rubric rewards), and Alibaba Qwen (pre-training data management & synthesis for Qwen series models).
My research spans multiple AI areasโLLMs, Machine Learning, NLP, and Multimodal:
- LLM Agentic RL & Benchmarking (Now): focusing on credit assignment and context management for long-horizon agents, with end-to-end data and benchmark construction;
- LLM Pretrain Data: working on pre-training data management and data synthesis;
- Automated Model Evaluation: evaluating model performance without labels using unsupervised proxies, such as contrastive and energy-based automated model evaluation;
- Machine Translation: including multimodal (vision-language), sign language (video-text), text-only machine translation.
I also maintain two GitHub repositories: TableGPT (a table LLM) and LLM-Synthetic-Data (a reading list on data synthesis) โ welcome to follow!
I am open to opportunities across academia and industry โ feel free to get in touch!
Google Scholar ย ย
Twitter ย ย
Email ย ย
GitHub ย ย
ย
๐ฅ News
| Aug 28, 2026 | Tencent Hy4-preview is released now. |
|---|---|
| Aug 20, 2026 | Our paper โMedACE: A Time-Aware Asynchronous Clinical Environment for Long-Horizon Medical Treatment Decision-Makingโ is accepted at EMNLP Findings 2026! |
| Aug 20, 2026 | Our paper โTRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learningโ is accepted at EMNLP 2026! |
| Jun 30, 2026 | Honored to receive the Outstanding Graduate award of Zhejiang University! |
| May 21, 2026 | Tencent Hy3 is released now. |
| Apr 23, 2026 | Tencent Hy3-preview is released now. |
| Apr 15, 2026 | Our paper โDataXman: Selecting and Mixing Pretraining Data via Bilingual Mixture-of-Experts Data Managerโ is accepted at SCIENCE CHINA Information Sciences 2026! |
| Apr 07, 2026 | Our paper โHSS-Synth: Humanities and Social Sciences Data Synthesis for LLMsโ is accepted at ACL Findings 2026! |
| Mar 02, 2026 | Our paper โW2S: Weak-to-Strong Prompt Correction for Large Language Modelsโ is accepted at Machine Learning 2026! |
| Jan 26, 2026 | Our paper โOptimSyn: Influence-Guided Rubrics Optimization for Synthetic Data Generationโ is accepted at ICLR 2026! |
| Aug 18, 2025 | Ant RL technical report โReinforcement learning with rubric anchorsโ(extending RLVR with 10k+ Rubric rewards) is now released. |
| Feb 20, 2025 | Gave an invited talk on โDataMan: Data Manager for Pre-training Large Language Modelsโ at JIQIZHIXIN (ๆบๅจไนๅฟ)! |
| Feb 11, 2025 | Our paper โLLM-Enhanced Query Generation and Retrieval Preservation for Task-Oriented Dialogueโ is accepted at Findings of ACL 2025! |
| Feb 11, 2025 | Our paper โDataMan: Data Manager for Pre-training Large Language Modelsโ is accepted at ICLR 2025! |
| Dec 19, 2024 | Qwen2.5 technical report are released now. |
| Sep 20, 2024 | One paper โInference-Time Decontamination: Reusing Leaked Benchmarks for Large Language Model Evaluationโ is accepted at Findings of EMNLP 2024 and two paper โPredicting Rewards Alongside Tokens: Non-disruptive Parameter Insertion for Efficient Inference Intervention in Large Language Modelโ, โEmbedding and Gradient Say Wrong: A White-Box Method for Hallucination Detectionโ are accepted at EMNLP 2024! |
| Sep 19, 2024 | Qwen2.5 series foundation models are released now. |
| Jul 15, 2024 | Qwen2 technical report are released now. |
| Jul 04, 2024 | Release the paper of โDotamathโ for mathematical reasoning. |
| Jun 17, 2024 | Qwen2 series foundation models are released now. |
| May 16, 2024 | Our paper โDORY: Deliberative Prompt Recovery for LLMโ is accepted at Findings of ACL 2024! |
| Feb 04, 2024 | Qwen1.5 series foundation models are released now. |
| Jan 16, 2024 | Our paper โEnergy-based Automated Model Evaluationโ is accepted at ICLR 2024! |
| Oct 23, 2023 | I started my internship at Alibaba Qwen Team! Ping me if you want to meet up in HangZhou :) |
| Jul 15, 2023 | Our paper โCAME: Contrastive Automated Model Evaluationโ is accepted at ICCV 2023! |
| Oct 06, 2022 | Our paper โDistill The Image to Nowhere: Inversion Knowledge Distillation for Multimodal Machine Translationโ is accepted at EMNLP 2022 (Oral)! |
| Sep 10, 2022 | Started my PhDโs degree at College of Computer Science and Technology of Zhejiang University! |
| Apr 06, 2022 | Our paper โHybridVocab: Towards Multi-Modal Machine Translation via Multi-Aspect Alignmentโ is accpeted at ICMR 2022 (Oral)! |
๐ผ Work Experience
Qingyun Program Research Intern on Long-horizon Agents
- data and benchmark construction for long-horizon agentic tasks;
- coding agentic RL with credit assignment and agent continual learning via context management.
Plan-A Research Intern on Data Synthesis and RL
Mentor: Junbo Zhao- Built open-ended instruction-following and preference alignment data.
- Contributed to Rubicon-preview: reinforcement learning from rubric rewards for closed- and open-ended tasks.
Research Intern on Pre-training Data Management and Synthesis
Mentor: Dayiheng Liu, Junyang Lin and Chang Zhou- Contributed to the Qwen 1.5/2/2.5 series base models.
- Developed Data Manager (DataMan and DataXman) for data selection and mixing in LLM pre-training, adopted in the Qwen base models and reported in JIQIZHIXIN (ๆบๅจไนๅฟ).
- Synthesized open-ended task data for mid-training of Qwen series models.
๐ Selected Publications
๐ Academic Services
- Conference Reviewer: ICLR 2024, 2025; ICML 2023, 2025; NeurIPS 2022, 2023, 2024; AAAI 2026; CVPR 2025; ICCV 2023, 2025; ECCV 2024; ACL 2024, 2026; AISTATS 2025; COLM 2024.
- Journal Reviewer: IEEE Transactions on Big Data (TBD), Transactions of Machine Learning Research (TMLR).
- Publication Chair: International Conference on Natural Language Processing (ICNLP) 2025.
๐ Miscellaneous
I love music ๐ต, sports (basketball ๐, football โฝ, running ๐โโ๏ธ, etc.), traveling the world ๐บ๏ธ, hanging out with friends ๐ป, and trying anything new. Click on Totoro below โ๏ธ to hear one of my favorite songs: ๅคฑๆใฝใณใฐๆฒขๅฑฑ่ดใใฆๆณฃใใฆใฐใใใฎ็งใฏใใใ