Robotics systems
Integrating hardware and software through ROS, with multi-sensor synchronisation, communication, calibration, control and real-time execution.
Zheyu Zhuang 庄哲宇
I’m a postdoctoral researcher at KTH, working with Professor Danica Kragic on data-efficient visuomotor policies for real-world manipulation. I previously completed a PhD in Robotics at the Australian National University with Professor Robert Mahony, where I combined learned perception and closed-loop control for visual robotic picking.
01 / Skills
I build robot-learning pipelines end to end, from data collection and ROS-based sensor integration to algorithm design, evaluation and deployment on physical systems.
Integrating hardware and software through ROS, with multi-sensor synchronisation, communication, calibration, control and real-time execution.
Designing, training and evaluating visuomotor and VLA policies across simulation and hardware, with attention to data efficiency, robustness and generalisation.
Combining RGB, depth, event cameras, point clouds and proprioception with visual foundation models such as DINO and SAM.
Designing data collection and evaluation pipelines, deploying policies and analysing the gap between simulated results and physical behaviour.
02 / Research
Recent papers develop action-grounded perception for robust, data-efficient manipulation.
Learns where to look directly from action supervision.
Task and robot state guide visual attention without spatial labels. The learned regions support RGB cropping, augmentation and point-cloud filtering.
On real robots: 76.7% in-domain success vs 48.3% for the strongest baseline.
Finds control-relevant visual detail within intermediate CNN features.
Proprioception selects local features within a CNN map, helping policies focus on action-relevant detail without manual image crops.
Generalises by aligning perception around the robot hand.
Local hand-centred perception and proprioception help manipulation skills transfer across workspaces, camera views and robot embodiments.
Mirrored demonstrations transfer skills across space with consistent images, state and 6 DoF actions.
End-to-end motor supervision changes what a visual encoder attends to, even in the same scene.
Policy saliency guides augmentation to make manipulation robust to visual distractions.
Image–proprioception alignment helps a policy recognise the body it controls.
03 / Systems
Earlier systems work brought the same perception-and-control focus to deployed robots.
04 / Mentoring
I contribute through mentoring, teaching, editorial work and peer review.
Mentored student projects contributing to CoRL, RA-L, and ICRA publications, and supervised undergraduate work resulting in an IROS paper.
Tutored five graduate courses and one undergraduate course, and led a Robotic Vision Summer School workshop later incorporated into teaching at Monash University.
Associate Editor for IROS 2025 and 2026 and ICRA 2027; reviewer for IEEE RA-L, IEEE T-RO, RSS, CoRL, ICRA, and IROS.
05 / Publications
Recent work focuses on data-efficient visuomotor learning; earlier papers span closed-loop control, 3D perception, human–robot collaboration and integrated manipulation systems.
06 / Off the clock
A few things that get me outside and away from a screen.
Helps me completely empty my head as the survival instinct screams, but oftentimes it's back to the familiar pattern of stopping half way and rethinking my approach.
Often question myself why I put myself through this — days without a shower — but in the end, the cleared mind, the breathtaking views, and being off the grid and off the internet make it all worth it.
Maybe it's the 10th time getting skunked, but all it takes is just one fish to keep me going.