👨🏻🎓 Biography
I am currently a Second-year Ph.D. student at the Show Lab,
National University of Singapore, advised by Prof. Mike Zheng Shou.
I obtained my Master of Science degree from the AAIS at
Peking University, advised by Prof. Yuxin Peng.
Previously, I have interned at TikTok and Kuaishou.
My research interests include 🤖 Agent, 🦾 Robotics, and 🖼️ Multimedia.
I’m open to collaborations and discussions. Feel free to drop me an email~
🔥 News
Sep. 2026🦾 We released Show-Harness [Website, Code], and a Survey on robot-use agent.Sep. 2026🎉 Three papers got accepted by CoRL 2026.May 2026🎉 Code2Video and ACA got accepted by ICML 2026.Jan. 2026🎉 We won the third place in RoCo Challenge @AAAI Embodied AI Workshop 2026.Nov. 2025🎉 UniAPO got accepted by AAAI 2026.Aug. 2025🎓 Joined Show Lab @ NUS to start my Ph.D. journey!
📝 Publications
⭐ Selected Publications
📚 All Publications
- Code2Video: A Code-centric Paradigm for Educational Video Generation
ICML 2026 Paper - Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation
ICML 2026 Paper - Where Success Breaks: Failure-Boundary Learning for Robust Vision-Language-Action Models
CoRL 2026 Paper - Supervise What Survives: Geometry-Guided VLA Adaptation from Synthetic Robot Videos
CoRL 2026 Paper - UniAPO: Unified Multimodal Automated Prompt Optimization
AAAI 2026 Paper - Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution
NLPCC 2026 Paper - PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress
ICMLW AI4S 2026 · Best Poster Award Paper - Show-Harness: Just a VLM Agent Can Play Robots
arXiv 2026 Paper - Survey on Multimodal Embodied Agents: A Unified Capability-centric Perspective from Computer-Use to Robot-Use
TechRxiv 2026 Paper - ActionMap: Robot Policy Learning via Voxel Action Heatmap
arXiv 2026 Paper - Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance
arXiv 2026 Paper - Spectral Surgery: Training-Free Refinement of LoRA via Gradient-Guided Singular Value Reweighting
arXiv 2026 Paper - MAI: A Multi-turn Aggregation-Iteration Model for Composed Image Retrieval
ICLR 2025 Paper - UniCode²: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
Technical Report 2025 Paper - FashionERN: Enhance-and-Refine Network for Composed Fashion Image Retrieval
AAAI 2024 Paper - SPIRIT: Style-guided Patch Interaction for Fashion Image Retrieval with Text Feedback
TOMM 2024 Paper - Real20M: A Large-scale E-commerce Dataset for Cross-domain Retrieval
ACM MM 2023 Paper - PKU_WICT at TRECVID 2022: Disaster Scene Description and Indexing Task
TRECVID 2022 · Ranked 1st Paper
🏫 Education
National University of Singapore2025.08 – Present
Ph.D., School of Computer (SoC)
Peking University2022.09 – 2025.06
Master, Academy for Advanced Interdisciplinary Studies (AAIS)
- Outstanding Graduate of the Wangxuan Institute of Computer Technology2025.06
- Merit Student2024.11
- Leo KoGuan Scholarship2024.11
Wuhan University2018.09 – 2022.06
Undergraduate, School of Computer Science
- Outstanding Undergraduate Graduate2022.06
- National Scholarship2020.11
📖 Service
- Conference Reviewer: ICLR, CoRL, ICML, AAAI, ACM MM, etc.
- Teaching Assistant: NUS EE4309 Robot Perception