Profile UCAS · ISCAS

Portrait of Xueru Wen

Xueru Wen 温学儒

Ph.D. student in Computer Application Technology

University of Chinese Academy of Sciences

Chinese Information Processing Laboratory

Institute of Software, Chinese Academy of Sciences

Advisor: Prof. Ben He

Email: wenxueru2022@iscas.ac.cn

Biography

I am a Ph.D. student at the University of Chinese Academy of Sciences (UCAS), advised by Prof. Ben He. I conduct research in the Chinese Information Processing Laboratory at the Institute of Software, Chinese Academy of Sciences (ISCAS). I expect to graduate in 2028.

I received my B.S. degree from Jilin University in 2023, where I was in the Tang Aoqing Honors Program in Science, Computer Science Track. Since September 2023, I have been interning at Xiaohongshu (RedNote), and became an Ace Top Intern in July 2025.

Research Interests

My research focuses on LLM post-training, particularly reward modeling, agents, and recursive self-improvement (RSI). I study how reliable reward signals and reinforcement learning improve agent reasoning and tool use. For RSI, I explore how agents can automate data generation, evaluation, and training to support iterative self-improvement. As models become more capable, I am interested in scalable oversight that keeps this process reliable and aligned.

Internship & Project Experience

Xiaohongshu (RedNote)Sep. 2023 – Present

Ace Top Intern since July 2025. LLM post-training, with a focus on reward modeling, agent and recursive self-improving systems.

  • Reward model iteration. Iterated RLHF reward models with preference feedback, hallucination detection, and process rewards; evaluated downstream effects and explored scalable oversight mechanisms.
  • Math Verification Agent. Built an agent with Lean, Python, and natural-language verifiers, then strengthened it with SFT and agentic RL to improve ProofBench performance.
  • SWE Coding Agent. Built Docker/Kubernetes sampling pipelines for large-scale trajectory collection and validation; used SFT to evaluate data effectiveness. Implemented TITO, R3, and partial rollout for agentic RL, improving SWE-bench Verified, Multilingual, and Pro performance.
  • Automated Posttrain. Built large-scale data synthesis pipelines on cloud GPU clusters, used world models for targeted synthesis, and implemented segment-wise tokenization for training and inference to improve PostTrainBench performance.
Chinese Information Processing Laboratory, ISCASMar. 2023 – Aug. 2023

Worked on LLM pre-training with Megatron–DeepSpeed, including LLaMA training support, checkpoint conversion, numerical alignment, fused kernels, and Flash Attention.

Publications

My name is in bold. BibTeX · Google Scholar

  1. Coupled Variational Reinforcement Learning for Language Model General Reasoning Paper Code
    Xueru Wen, Jie Lou, Yanjiang Liu, Hongyu Lin, Ben He, Xianpei Han, Le Sun, Yaojie Lu, Debing Zhang.
    ICML 2026.
  2. Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards Paper
    Zhengzhao Ma, Xueru Wen, Boxi Cao, Yaojie Lu, Hongyu Lin, Jinglin Yang, Min He, Xianpei Han, Le Sun.
    ICML 2026.
  3. Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewards Paper
    Mengjie Ren, Jie Lou, Boxi Cao, Xueru Wen, Hongyu Lin, Xianpei Han, Le Sun, Xing Yu, Yaojie Lu.
    COLM 2026.
  4. Rethinking Reward Model Evaluation: Are We Barking up the Wrong Tree? Paper
    Xueru Wen, Jie Lou, Yaojie Lu, Hongyu Lin, Xing Yu, Xinyu Lu, Ben He, Xianpei Han, Debing Zhang, Le Sun.
    ICLR 2025 Spotlight
  5. Cheems: A Practical Guidance for Building and Evaluating Chinese Reward Models from Scratch Paper Code
    Xueru Wen, Jie Lou, Zichao Li, Yaojie Lu, Xing Yu, Yuqiu Ji, Guohai Xu, Hongyu Lin, Ben He, Xianpei Han, Le Sun, Debing Zhang.
    ACL 2025.
  6. On-Policy Self-Alignment with Fine-grained Knowledge Feedback for Hallucination Mitigation Paper
    Xueru Wen, Jie Lou, Xinyu Lu, Yuqiu Ji, Xinyan Guan, Yaojie Lu, Hongyu Lin, Ben He, Xianpei Han, Debing Zhang, Le Sun.
    ACL 2025.
  7. Transferable Post-training via Inverse Value Learning Paper
    Xinyu Lu, Xueru Wen, Yaojie Lu, Bowen Yu, Hongyu Lin, Haiyang Yu, Le Sun, Xianpei Han, Yongbin Li.
    NAACL 2025.
  8. AutoAlign: Get Your LLM Aligned with Minimal Annotations Paper
    Xinyu Lu, Dong Xu, Chunkang Zhang, Xinyan Guan, Junxiang Wang, Qingyu Zhang, Pengbo Wang, Yingzhi Mao, Hao Xiang, Xueru Wen, Zichao Li, Yaojie Lu, Hongyu Lin, Le Sun, Xianpei Han.
    ACL 2025.
  9. Critic-CoT: Boosting the Reasoning Abilities of Large Language Model via Chain-of-Thought Critic Paper
    Xin Zheng, Jie Lou, Boxi Cao, Xueru Wen, Yuqiu Ji, Hongyu Lin, Yaojie Lu, Xianpei Han, Debing Zhang, Le Sun.
    ACL 2025.
  10. The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward Models Paper Code
    Zichao Li, Xueru Wen, Jie Lou, Yuqiu Ji, Yaojie Lu, Xianpei Han, Debing Zhang, Le Sun.
    ICML 2025.
  11. Offline Pseudo Relevance Feedback for Efficient and Effective Single-Pass Dense Retrieval Paper
    Xueru Wen, Xiaoyang Chen, Xuanang Chen, Ben He, Le Sun.
    SIGIR 2023.
  12. End-to-End Entity Detection with Proposer and Regressor Paper
    Xueru Wen, Changjiang Zhou, Haotian Tang, Luguang Liang, Hong Qi, Yu Jiang.
    Neural Processing Letters 2023.
  13. Type-supervised Sequence Labeling Based on the Heterogeneous Star Graph for Named Entity Recognition Paper
    Xueru Wen, Changjiang Zhou, Haotian Tang, Luguang Liang, Yu Jiang, Hong Qi.
    arXiv, 2022.

Tools & Resources

A few small tools I’ve developed that you might find useful. Feel free to have your agent give them a try.

Skynet

A native macOS workspace for Codex and Claude Code, bringing together local and SSH sessions, queued follow-ups, and code review.

Python Hot Patch Skill

An agent skill for hot-patching live CPython processes while preserving in-memory state, with safe-point execution and independent verification.

Wait Skill

Event-driven waiting and task coordination for coding agents, supporting passive monitoring, timed loops, and workflows organized by task dependencies.

JSONLine Viewer

A VS Code extension for inspecting JSONL and NDJSON records, with formatted JSON previews and indexed navigation through large files.

Awards & Honors

Contact

Email: wenxueru2022@iscas.ac.cn
Chinese Information Processing Laboratory
Institute of Software, Chinese Academy of Sciences, Beijing, China