Ru Peng (Perry)
"Only love endures the passage of time"
I’m a final-year PhD student at Computer Science Department of Zhejiang University (ZJU), advised by Professors Junbo Zhao and Gang Chen. I’m affiliated with DiLab-ZJU and the State Key Laboratory of Blockchain and Data Security. Early on, I was fortunate that Professors Tianyong Hao and Kehai Chen ushered me into research. I have interned at three LLM labs: Tencent Hunyuan (long-horizon agents), Ant Group Inclusion AI (data synthesis & RL from rubric rewards), and Alibaba Qwen (pre-training data management & synthesis for Qwen series models).
My research spans multiple AI areas—LLMs, Machine Learning, NLP, and Multimodal:
- LLM Agentic RL & Benchmarking (Now): focusing on credit assignment and context management for long-horizon agents, with end-to-end data and benchmark construction;
- LLM Pretrain Data: working on pre-training data management and data synthesis;
- Automated Model Evaluation: evaluating model performance without labels using unsupervised proxies, such as contrastive and energy-based automated model evaluation;
- Machine Translation: including multimodal (vision-language), sign language (video-text), text-only machine translation.
I also maintain two GitHub repositories: TableGPT (a table LLM) and LLM-Synthetic-Data (a reading list on data synthesis) — welcome to follow!
I am open to opportunities across academia and industry — feel free to get in touch!
Google Scholar
Twitter
Email
GitHub
🔥 News
| Sep 19, 2026 | Two papers “From “Weak” Signals to Strong Models: Preference Delta Aggregation with LoRA Merging” and “From Structural Feedback to Prompt Policies: Learning Faithful Text-to-Image Prompt Editors” are accepted at NeurIPS 2026! |
|---|---|
| Aug 28, 2026 | Tencent Hy4-preview is released now. |
| Aug 20, 2026 | Our paper “MedACE: A Time-Aware Asynchronous Clinical Environment for Long-Horizon Medical Treatment Decision-Making” is accepted at EMNLP Findings 2026! |
| Aug 20, 2026 | Our paper “TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning” is accepted at EMNLP 2026! |
| Jun 30, 2026 | Honored to receive the Outstanding Graduate award of Zhejiang University! |
| May 21, 2026 | Tencent Hy3 is released now. |
| Apr 23, 2026 | Tencent Hy3-preview is released now. |
| Apr 15, 2026 | Our paper “DataXman: Selecting and Mixing Pretraining Data via Bilingual Mixture-of-Experts Data Manager” is accepted at SCIENCE CHINA Information Sciences 2026! |
| Apr 07, 2026 | Our paper “HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs” is accepted at ACL Findings 2026! |
| Mar 02, 2026 | Our paper “W2S: Weak-to-Strong Prompt Correction for Large Language Models” is accepted at Machine Learning 2026! |
| Jan 26, 2026 | Our paper “OptimSyn: Influence-Guided Rubrics Optimization for Synthetic Data Generation” is accepted at ICLR 2026! |
| Aug 18, 2025 | Ant RL technical report “Reinforcement learning with rubric anchors“(extending RLVR with 10k+ Rubric rewards) is now released. |
| Feb 20, 2025 | Gave an invited talk on “DataMan: Data Manager for Pre-training Large Language Models” at JIQIZHIXIN (机器之心)! |
| Feb 11, 2025 | Our paper “LLM-Enhanced Query Generation and Retrieval Preservation for Task-Oriented Dialogue” is accepted at Findings of ACL 2025! |
| Feb 11, 2025 | Our paper “DataMan: Data Manager for Pre-training Large Language Models” is accepted at ICLR 2025! |
| Dec 19, 2024 | Qwen2.5 technical report are released now. |
| Sep 20, 2024 | One paper “Inference-Time Decontamination: Reusing Leaked Benchmarks for Large Language Model Evaluation” is accepted at Findings of EMNLP 2024 and two paper “Predicting Rewards Alongside Tokens: Non-disruptive Parameter Insertion for Efficient Inference Intervention in Large Language Model”, “Embedding and Gradient Say Wrong: A White-Box Method for Hallucination Detection” are accepted at EMNLP 2024! |
| Sep 19, 2024 | Qwen2.5 series foundation models are released now. |
| Jul 15, 2024 | Qwen2 technical report are released now. |
| Jul 04, 2024 | Release the paper of “Dotamath” for mathematical reasoning. |
| Jun 17, 2024 | Qwen2 series foundation models are released now. |
| May 16, 2024 | Our paper “DORY: Deliberative Prompt Recovery for LLM” is accepted at Findings of ACL 2024! |
| Feb 04, 2024 | Qwen1.5 series foundation models are released now. |
| Jan 16, 2024 | Our paper “Energy-based Automated Model Evaluation” is accepted at ICLR 2024! |
| Oct 23, 2023 | I started my internship at Alibaba Qwen Team! Ping me if you want to meet up in HangZhou :) |
| Jul 15, 2023 | Our paper “CAME: Contrastive Automated Model Evaluation” is accepted at ICCV 2023! |
| Oct 06, 2022 | Our paper “Distill The Image to Nowhere: Inversion Knowledge Distillation for Multimodal Machine Translation” is accepted at EMNLP 2022 (Oral)! |
| Sep 10, 2022 | Started my PhD’s degree at College of Computer Science and Technology of Zhejiang University! |
| Apr 06, 2022 | Our paper “HybridVocab: Towards Multi-Modal Machine Translation via Multi-Aspect Alignment” is accpeted at ICMR 2022 (Oral)! |
💼 Work Experience
Qingyun Program Research Intern on Long-horizon Agents
- data and benchmark construction for long-horizon agentic tasks;
- coding agentic RL with credit assignment and agent continual learning via context management.
Plan-A Research Intern on Data Synthesis and RL
Mentor: Junbo Zhao- Built open-ended instruction-following and preference alignment data.
- Contributed to Rubicon-preview: reinforcement learning from rubric rewards for closed- and open-ended tasks.
Research Intern on Pre-training Data Management and Synthesis
Mentor: Dayiheng Liu, Junyang Lin and Chang Zhou- Contributed to the Qwen 1.5/2/2.5 series base models.
- Developed Data Manager (DataMan and DataXman) for data selection and mixing in LLM pre-training, adopted in the Qwen base models and reported in JIQIZHIXIN (机器之心).
- Synthesized open-ended task data for mid-training of Qwen series models.
📝 Selected Publications
📚 Academic Services
- Conference Reviewer: ICLR 2024, 2025; ICML 2023, 2025; NeurIPS 2022, 2023, 2024; AAAI 2026; CVPR 2025; ICCV 2023, 2025; ECCV 2024; ACL 2024, 2026; AISTATS 2025; COLM 2024.
- Journal Reviewer: IEEE Transactions on Big Data (TBD), Transactions of Machine Learning Research (TMLR).
- Publication Chair: International Conference on Natural Language Processing (ICNLP) 2025.
😊 Miscellaneous
I love music 🎵, sports (basketball 🏀, football ⚽, running 🏃♂️, etc.), traveling the world 🗺️, hanging out with friends 🍻, and trying anything new. Click on Totoro below ↘️ to hear one of my favorite songs: 失恋ソング沢山聴いて泣いてばかりの私はもう。