Ru Peng (Perry)

Hi 😃! 彭儒, PhD @ Zhejiang University.

rupeng.jpg

"Only love endures the passage of time"

I’m a final-year PhD student at Computer Science Department of Zhejiang University (ZJU), advised by Professors Junbo Zhao and Gang Chen. I’m affiliated with DiLab-ZJU and the State Key Laboratory of Blockchain and Data Security. Early on, I was fortunate that Professors Tianyong Hao and Kehai Chen ushered me into research. I have interned at three LLM labs: Tencent Hunyuan (long-horizon agents), Ant Group Inclusion AI (data synthesis & RL from rubric rewards), and Alibaba Qwen (pre-training data management & synthesis for Qwen series models).

My research spans multiple AI areas—LLMs, Machine Learning, NLP, and Multimodal:

I also maintain two GitHub repositories: TableGPT (a table LLM) and LLM-Synthetic-Data (a reading list on data synthesis) — welcome to follow!

I am open to opportunities across academia and industry — feel free to get in touch!

Google Scholar Google Scholar    Twitter Twitter    Email Email    GitHub GitHub   

 

🔥 News

Sep 19, 2026 Two papers “From “Weak” Signals to Strong Models: Preference Delta Aggregation with LoRA Merging” and “From Structural Feedback to Prompt Policies: Learning Faithful Text-to-Image Prompt Editors” are accepted at NeurIPS 2026!
Aug 28, 2026 Tencent Hy4-preview is released now.
Aug 20, 2026 Our paper “MedACE: A Time-Aware Asynchronous Clinical Environment for Long-Horizon Medical Treatment Decision-Making” is accepted at EMNLP Findings 2026!
Aug 20, 2026 Our paper “TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning” is accepted at EMNLP 2026!
Jun 30, 2026 Honored to receive the Outstanding Graduate award of Zhejiang University!
May 21, 2026 Tencent Hy3 is released now.
Apr 23, 2026 Tencent Hy3-preview is released now.
Apr 15, 2026 Our paper “DataXman: Selecting and Mixing Pretraining Data via Bilingual Mixture-of-Experts Data Manager” is accepted at SCIENCE CHINA Information Sciences 2026!
Apr 07, 2026 Our paper “HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs” is accepted at ACL Findings 2026!
Mar 02, 2026 Our paper “W2S: Weak-to-Strong Prompt Correction for Large Language Models” is accepted at Machine Learning 2026!
Jan 26, 2026 Our paper “OptimSyn: Influence-Guided Rubrics Optimization for Synthetic Data Generation” is accepted at ICLR 2026!
Aug 18, 2025 Ant RL technical report “Reinforcement learning with rubric anchors“(extending RLVR with 10k+ Rubric rewards) is now released.
Feb 20, 2025 Gave an invited talk on “DataMan: Data Manager for Pre-training Large Language Models” at JIQIZHIXIN (机器之心)!
Feb 11, 2025 Our paper “LLM-Enhanced Query Generation and Retrieval Preservation for Task-Oriented Dialogue” is accepted at Findings of ACL 2025!
Feb 11, 2025 Our paper “DataMan: Data Manager for Pre-training Large Language Models” is accepted at ICLR 2025!
Dec 19, 2024 Qwen2.5 technical report are released now.
Sep 20, 2024 One paper “Inference-Time Decontamination: Reusing Leaked Benchmarks for Large Language Model Evaluation” is accepted at Findings of EMNLP 2024 and two paper “Predicting Rewards Alongside Tokens: Non-disruptive Parameter Insertion for Efficient Inference Intervention in Large Language Model”, “Embedding and Gradient Say Wrong: A White-Box Method for Hallucination Detection” are accepted at EMNLP 2024!
Sep 19, 2024 Qwen2.5 series foundation models are released now.
Jul 15, 2024 Qwen2 technical report are released now.
Jul 04, 2024 Release the paper of “Dotamath” for mathematical reasoning.
Jun 17, 2024 Qwen2 series foundation models are released now.
May 16, 2024 Our paper “DORY: Deliberative Prompt Recovery for LLM” is accepted at Findings of ACL 2024!
Feb 04, 2024 Qwen1.5 series foundation models are released now.
Jan 16, 2024 Our paper “Energy-based Automated Model Evaluation” is accepted at ICLR 2024!
Oct 23, 2023 I started my internship at Alibaba Qwen Team! Ping me if you want to meet up in HangZhou :)
Jul 15, 2023 Our paper “CAME: Contrastive Automated Model Evaluation” is accepted at ICCV 2023!
Oct 06, 2022 Our paper “Distill The Image to Nowhere: Inversion Knowledge Distillation for Multimodal Machine Translation” is accepted at EMNLP 2022 (Oral)!
Sep 10, 2022 Started my PhD’s degree at College of Computer Science and Technology of Zhejiang University!
Apr 06, 2022 Our paper “HybridVocab: Towards Multi-Modal Machine Translation via Multi-Aspect Alignment” is accpeted at ICMR 2022 (Oral)!

💼 Work Experience

Hunyuan LLM Team, Tencent Dec. 2025 - Now

Qingyun Program Research Intern on Long-horizon Agents

  • data and benchmark construction for long-horizon agentic tasks;
  • coding agentic RL with credit assignment and agent continual learning via context management.
Inclusion AI Team, Ant Group April 2025 - Oct 2025

Plan-A Research Intern on Data Synthesis and RL

Mentor: Junbo Zhao
  • Built open-ended instruction-following and preference alignment data.
  • Contributed to Rubicon-preview: reinforcement learning from rubric rewards for closed- and open-ended tasks.
Qwen Pre-training Team, Alibaba Group Oct 2023 - April 2025

Research Intern on Pre-training Data Management and Synthesis

Mentor: Dayiheng Liu, Junyang Lin and Chang Zhou
  • Contributed to the Qwen 1.5/2/2.5 series base models.
  • Developed Data Manager (DataMan and DataXman) for data selection and mixing in LLM pre-training, adopted in the Qwen base models and reported in JIQIZHIXIN (机器之心).
  • Synthesized open-ended task data for mid-training of Qwen series models.

📝 Selected Publications

  1. Hy4P Tech Report
    hy4_preview_technical_report.png
    Hy4-preview model card
    Tencent Hy Team
    Aug 2026
  2. Ant RL Tech Report
    Ant_RL_TechReport_Rubicon.png
    Reinforcement learning with rubric anchors
    Zenan Huang, Yihong Zhuang, Guoshan Lu, Zeyu Qin, Haokai Xu, Tianyu Zhao, Ru Peng , Jiaqi Hu, and 3 more authors
    arXiv preprint arXiv:2508.12790, Aug 2025
  3. ICLR 2025
    ICLR2025_DataMan.jpg
    DataMan: Data Manager for Pre-training Large Language Models
    In The Thirteenth International Conference on Learning Representations, Aug 2025
  4. Qwen2.5 Technical Report
    qwen2_5_technical_report.jpg
    Qwen2.5 Technical Report
    An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, and 33 more authors
    arXiv preprint arXiv:2412.15115, Aug 2024
  5. Qwen2 Technical Report
    qwen2_technical_report.jpg
    Qwen2 technical report, 2024
    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, and 3 more authors
    arXiv preprint arXiv:2407.10671, Aug 2024
  6. Qwen1.5 Blog
    qwen1_5_blog.jpeg
    Introducing qwen1. 5
    Qwen Team
    Online Blog, Aug 2024

📚 Academic Services

  • Conference Reviewer: ICLR 2024, 2025; ICML 2023, 2025; NeurIPS 2022, 2023, 2024; AAAI 2026; CVPR 2025; ICCV 2023, 2025; ECCV 2024; ACL 2024, 2026; AISTATS 2025; COLM 2024.
  • Journal Reviewer: IEEE Transactions on Big Data (TBD), Transactions of Machine Learning Research (TMLR).
  • Publication Chair: International Conference on Natural Language Processing (ICNLP) 2025.

😊 Miscellaneous

I love music 🎵, sports (basketball 🏀, football ⚽, running 🏃‍♂️, etc.), traveling the world 🗺️, hanging out with friends 🍻, and trying anything new. Click on Totoro below ↘️ to hear one of my favorite songs: 失恋ソング沢山聴いて泣いてばかりの私はもう。


Totoro Bottle