Zheyu Zhuang 庄哲宇

Portrait of Zheyu Zhuang

I make a handful of demonstrations go a long way.

I’m a postdoctoral researcher at KTH, working with Professor Danica Kragic on data-efficient visuomotor policies for real-world manipulation. I previously completed a PhD in Robotics at the Australian National University with Professor Robert Mahony, where I combined learned perception and closed-loop control for visual robotic picking.

01 / Skills

A systems approach to robot learning.

I build robot-learning pipelines end to end, from data collection and ROS-based sensor integration to algorithm design, evaluation and deployment on physical systems.

01

Robotics systems

Integrating hardware and software through ROS, with multi-sensor synchronisation, communication, calibration, control and real-time execution.

ROSSensor synchronisationSystem integrationReal-time control
02

Robot learning

Designing, training and evaluating visuomotor and VLA policies across simulation and hardware, with attention to data efficiency, robustness and generalisation.

Algorithm designVLA fine-tuningImitation learningPolicy evaluation
03

Multimodal perception

Combining RGB, depth, event cameras, point clouds and proprioception with visual foundation models such as DINO and SAM.

DINO & SAMRGB & depthEvent camerasPoint clouds6D poseProprioception
04

Real-robot deployment

Designing data collection and evaluation pipelines, deploying policies and analysing the gap between simulated results and physical behaviour.

Data collectionManipulationSystem integrationFailure analysis

02 / Research

Research that learns what matters for action.

Recent papers develop action-grounded perception for robust, data-efficient manipulation.

2026 · CoRL

Seeker: Attention from Action, for Action

Learns where to look directly from action supervision.

Task and robot state guide visual attention without spatial labels. The learned regions support RGB cropping, augmentation and point-cloud filtering.

On real robots: 76.7% in-domain success vs 48.3% for the strongest baseline.

2026 · CoRL

FocusPool: Localized Visual Feature Aggregation

Finds control-relevant visual detail within intermediate CNN features.

Proprioception selects local features within a CNN map, helping policies focus on action-relevant detail without manual image crops.

2026 · RA-L

PALM: Perception Alignment for Local Visuomotor Policies

Generalises by aligning perception around the robot hand.

Local hand-centred perception and proprioception help manipulation skills transfer across workspaces, camera views and robot embodiments.

More research

2024–2025
2025 · CoRL

MirrorDuo

Mirrored demonstrations transfer skills across space with consistent images, state and 6 DoF actions.

2025 · ICRA

Task-aware visual encoders

End-to-end motor supervision changes what a visual encoder attends to, even in the same scene.

2024 · CoRL

RoboSaGA

Policy saliency guides augmentation to make manipulation robust to visual distractions.

2024 · IROS

Robot-centric Pooling

Image–proprioception alignment helps a policy recognise the body it controls.

Seeker ROI across real-robot tasks
Enlarged Seeker regions of interest across two real-robot manipulation tasks
FocusPool Local visual features for policy learning Open full image ↗
Enlarged comparison of FocusPool, average pooling and spatial softmax using CNN feature maps
PALM Perception alignment across domains
Enlarged PALM perception alignment figure showing changes in workspace position, camera viewpoint, and robot embodiment with a TCP-centric crop
RoboSaGA Robust manipulation across visual shifts
Enlarged RoboSaGA manipulation policies operating across varied visual backgrounds
Task-aware visual encoders Same scene, different policy objective
Enlarged comparison of task-dependent encoder attention for reach-spam and reach-mug policies
MirrorDuo Reflection-consistent visuomotor learning
Enlarged MirrorDuo reflection-consistent visuomotor learning animation
Robot-centric Pooling Body ownership through image and proprioception alignment
Enlarged Robot-centric Pooling comparison of alignment scores and saliency maps

03 / Systems

Beyond policies: complete robot systems.

Earlier systems work brought the same perception-and-control focus to deployed robots.

GoferBot collaborating with a person during furniture assembly

GoferBot

I co-developed GoferBot, a vision-based human–robot interaction system for collaborative furniture assembly. It combines human action recognition, visual servoing and handover behaviour so the robot can coordinate naturally with a person.

Watch GoferBot
Cartman warehouse picking robot Semantic segmentation of warehouse objects for the Cartman project

Cartman

I contributed as a member of the ACRV team that built Cartman, combining custom hardware with robotic perception for warehouse picking. The system won the 2017 Amazon Robotics Challenge, and the team leader later spun this work out into Lyro Robotics.

Amazon announcement Paper Lyro Robotics

04 / Mentoring

Supporting research beyond my own work.

I contribute through mentoring, teaching, editorial work and peer review.

Mentoring

Developing independent researchers

Mentored student projects contributing to CoRL, RA-L, and ICRA publications, and supervised undergraduate work resulting in an IROS paper.

Teaching

Connecting theory to hardware

Tutored five graduate courses and one undergraduate course, and led a Robotic Vision Summer School workshop later incorporated into teaching at Monash University.

Service

Editorial and peer-review service

Associate Editor for IROS 2025 and 2026 and ICRA 2027; reviewer for IEEE RA-L, IEEE T-RO, RSS, CoRL, ICRA, and IROS.

05 / Publications

Selected publications across robot learning, perception and control.

Recent work focuses on data-efficient visuomotor learning; earlier papers span closed-loop control, 3D perception, human–robot collaboration and integrated manipulation systems.

Show publication list Hide publication list
GoferBot: A Visual Guided Human-Robot Collaborative Assembly System Zheyu Zhuang, Yizhak Ben-Shabat, Jiahao Zhang, Stephen Gould, Robert Mahony IROS · 2022
A General Approach to State Refinement Gerard Kennedy, Jin Gao, Zheyu Zhuang, Xin Yu, Robert Mahony IROS · 2021
Stereo Hybrid Event-Frame Cameras for 3D Perception Ziwei Wang, Liyuan Pan, Yonhon Ng, Zheyu Zhuang, Robert Mahony IROS · 2021
End-to-end Multi-Instance Robotic Reaching from Monocular Vision Zheyu Zhuang, Xin Yu, Robert Mahony ICRA · 2021
6DoF Object Pose Estimation via Differentiable Proxy Voting Regularizer Xin Yu, Zheyu Zhuang, Piotr Koniusz, Hongdong Li BMVC · 2020
LyRN: A Real-Time Closed Loop Approach from Monocular Vision Zheyu Zhuang, Xin Yu, Robert Mahony ICRA · 2020
Learning Real-time Closed Loop Robotic Reaching from Monocular Vision by Exploiting a Control Lyapunov Function Structure Zheyu Zhuang, Jürgen Leitner, Robert Mahony IROS · 2019
Cartman: The Low-cost Cartesian Manipulator That Won the Amazon Robotics Challenge D. Morrison, A. W. Tow, M. McTaggart, R. Smith, N. Kelly-Boxall, S. Wade-McCue, J. Erskine, R. Grinover, A. Gurman, T. Hunn, D. Lee, A. Milan, T. Pham, G. Rallos, A. Razjigaev, T. Rowntree, K. Vijay, Zheyu Zhuang, C. Lehnert, I. Reid, P. Corke, J. Leitner ICRA · 2018
Semantic Segmentation from Limited Training Data A. Milan, T. Pham, K. Vijay, D. Morrison, A. W. Tow, L. Liu, J. Erskine, R. Grinover, A. Gurman, T. Hunn, N. Kelly-Boxall, D. Lee, M. McTaggart, G. Rallos, A. Razjigaev, T. Rowntree, T. Shen, R. Smith, S. Wade-McCue, Zheyu Zhuang, C. Lehnert, G. Lin, I. Reid, P. Corke, J. Leitner ICRA · 2018

06 / Off the clock

When I'm not chasing robots around a lab.

A few things that get me outside and away from a screen.

Zheyu climbing

Climbing

Helps me completely empty my head as the survival instinct screams, but oftentimes it's back to the familiar pattern of stopping half way and rethinking my approach.

Zheyu hiking

Hiking

Often question myself why I put myself through this — days without a shower — but in the end, the cleared mind, the breathtaking views, and being off the grid and off the internet make it all worth it.

Zheyu fishing

Fishing

Maybe it's the 10th time getting skunked, but all it takes is just one fish to keep me going.