About Me
I am currently a Ph.D. student in Data Science and Artificial Intelligence at The Hong Kong Polytechnic University, supervised by Prof.
KC Tan. My research lies at the intersection of Reinforcement Learning and LLM Reasoning. More broadly, I am interested in Interpretability and Evolutionary Computation with RL (EC + RL).
I received my M.S. in Applied Statistics from Fudan University and my B.S. in Statistics from
Sun Yat-sen University. I was fortunate to enjoy research experiences at Tencent AI Lab, advised by Dr.
Haobo Fu, where I focused on Policy Diversity in RL; at ByteDance Seed, advised by Xiongcai Luo, where I worked on efficient reasoning and on-policy distillation; and at Amazon Web Services (AWS), advised by
Tong He and
Tianjun Xiao, where I worked on video object segmentation.
I am always happy to discuss research ideas around RL, LLM Reasoning, Interpretability, and EC + RL. Please feel free to contact me at nigelyaoj@gmail.com.
我目前是香港理工大学数据科学与人工智能方向的博士生,由 KC Tan 教授指导。我的研究兴趣主要集中在强化学习(RL)与大语言模型(LLM)推理的交叉方向,也关注可解释性以及演化计算与强化学习(EC + RL)。
我硕士毕业于复旦大学应用统计专业,本科毕业于中山大学统计学专业。关于研究经历,我有幸在 Tencent AI Lab 参与研究,在 Haobo Fu 指导下关注 RL 中的策略多样性;在 ByteDance Seed 参与 efficient reasoning 和 on-policy distillation 相关研究,由 Xiongcai Luo 指导;也曾在 Amazon Web Services(AWS) 参与 video object segmentation 相关研究,由 Tong He 和 Tianjun Xiao 指导。
欢迎交流 RL、LLM Reasoning、可解释性和 EC + RL 相关研究问题,也可以通过 nigelyaoj@gmail.com 联系我。
Publications
🤖 LLM Reasoning
- [ICML 2026] SLAT: Segment-Level Adaptive Trimming for Efficient CoT Reasoning
This work studies how to reduce structural redundancy in chain-of-thought reasoning. It proposes segment-level adaptive trimming to selectively suppress low-utility redundant reasoning segments, improving the accuracy-efficiency trade-off for large reasoning models.
- [NeurIPS 2025 Spotlight] Diversity-Aware Policy Optimization for Large Language Model Reasoning
This work studies the relationship between solution diversity and reasoning potential in LLM reasoning, and proposes a diversity-aware policy optimization method for reinforcement learning training.
- [arXiv 2025] VAR-MATH: Probing True Mathematical Reasoning in LLMs via Symbolic Multi-Instance Benchmarks
This work introduces VAR-MATH, a symbolic multi-instance benchmark that probes whether LLM mathematical gains reflect genuine reasoning rather than benchmark-specific overfitting or memorization.
🎮 Diversity in RL
- [ICLR 2025] Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation Learning
This work focuses on recovering stylistically diverse policies from expert trajectories by weighting state-action pairs with pointwise mutual information.
- [NeurIPS 2023] Policy Space Diversity for Non-Transitive Games
This work studies diversity in non-transitive games and proposes a policy-space diversity measure that is more aligned with Nash approximation quality.
- [ICLR 2023] Quality-Similar Diversity via Population Based Reinforcement Learning
This work studies how to learn user-controllable and task-relevant diverse policy sets while maintaining similar policy quality.
🧩 Others (CV & GNN)
- [NeurIPS 2022 Spotlight] Self-supervised Amodal Video Object Segmentation
This work proposes a self-supervised video segmentation framework for inferring complete object shapes in occluded scenes using temporal information.
- [NeurIPS 2021] GRIN: Generative Relation and Intention Network for Multi-agent Trajectory Prediction
This work combines conditional generative modeling with graph neural networks to model agent intentions and social relations for multi-agent trajectory prediction.
Academic Services
- Reviewer for NeurIPS, ICLR, and ICML, 2023-present
- Reviewer for AAAI, 2025-present
- Reviewer for IEEE Transactions on Evolutionary Computation (TEVC)
- Reviewer for IEEE Transactions on Cognitive and Developmental Systems (TCDS)
Teaching Assistant
- ENG 2003: Information Technology, The Hong Kong Polytechnic University, Fall 2024
- COMP 6707: Computational Intelligence, The Hong Kong Polytechnic University, Spring 2025
- DSAI 4205: Big Data Analysis, The Hong Kong Polytechnic University, Fall 2025 and Spring 2026
