Zheyu Zhuang 庄哲宇 Roboticist

Portrait of Zheyu Zhuang

I make a handful of demonstrations go a long way.

Behaviour cloning turns demonstrations into a policy, and that policy stays tied to the conditions it was trained under. Change the background, move the object, shift the camera, swap the arm, and it starts to fail. Those are usually treated as four different problems. I treat them as one, and use the structure a task already contains instead of more data.

Current
Postdoctoral researcher · KTH
Focus
Generalisation from less data
Systems
Simulation to real robots

roboticist.zheyu@gmail.com

01 / About

Building embodied intelligence that generalises.

I am a postdoctoral researcher in Robotics, Perception and Learning at KTH Royal Institute of Technology, working with Professor Danica Kragic. I received my PhD in Robotics from the Australian National University with Professor Robert Mahony.

My research asks how visuomotor policies can learn action-relevant representations that transfer across changes in visual appearance, spatial configuration, viewpoint and robot embodiment. By grounding perception in the task, interaction and the robot's own body, I aim to make policies more data-efficient and reliable beyond the conditions in which they were trained.

I approach embodied AI as an end-to-end robotics systems problem. My work spans simulation, real-robot data collection, multimodal sensing and calibration, ROS integration, policy learning, real-time control, evaluation and deployment. This systems perspective helps me connect policy design to physical constraints, diagnose failures across the pipeline, and investigate how embodied agents can learn from limited data and transfer across environments and robot platforms.

02 / Research

Perception becomes meaningful through action.

A useful representation of the world is not necessarily a complete one. For a robot, meaning comes from what it can do, what it is trying to achieve and how it understands its own body. My research explores how these connections can be learned and how they can endure as the world changes.

2026 · Preprint

Seeker: Attention from Action, for Action

Learns where to look directly from action supervision.

Seeker uses the task and robot state to produce task-aware regions of interest, without spatial labels. Instead of treating every pixel equally, it identifies the evidence needed for the next action. The learned attention supports RGB cropping, background augmentation and point-cloud filtering across changing scenes.

2026 · RA-L

PALM: Perception Alignment for Local Visuomotor Policies

Generalises by aligning perception around the robot hand.

PALM focuses on local visual evidence around the robot hand, then connects it with proprioception. The representation stays meaningful across changes in workspace, camera viewpoint and robot embodiment, allowing fine manipulation skills to transfer without relearning the entire scene.

2025 · CoRL

MirrorDuo: Reflection-Consistent Visuomotor Learning

Transfers skills across space through consistent data mirroring.

MirrorDuo reflects images, proprioception and full 6 DoF actions as one consistent transformation. Each demonstration becomes a physically coherent paired example, enabling transfer to mirrored workspaces with little additional data.

2025 · ICRA

Feature Extractor or Decision Maker? Rethinking Visual Encoders

The same visual scene produces different representations when the task changes.

When trained end to end, the encoder absorbs the policy objective rather than behaving like a generic visual backbone. Even for similar images, motor supervision changes where it attends and which features it preserves. This reveals task aware visual representations.

2024 · CoRL

RoboSaGA: Saliency-Guided Augmentation for Visual Robustness

Builds visual robustness from the encoder’s own attention.

RoboSaGA uses policy saliency to locate and protect the visual evidence needed for control, while aggressively changing irrelevant appearance during training. This improves robustness to lighting, shadows, distractors, textures and backgrounds without manually defined masks.

2024 · IROS

Robot-Centric Pooling: Body Ownership in Visuomotor Policies

Learns visual self identification in multi robot environments.

Robot-centric Pooling connects visual features with proprioception so a policy can recognise the body it controls among similar-looking robots. This robot-centred perception remains reliable around unseen embodiments, self-distractors and large pixel shifts.

Seeker ROI across real-robot tasks
Enlarged Seeker regions of interest across two real-robot manipulation tasks
PALM Perception alignment across domains
Enlarged PALM perception alignment figure showing changes in workspace position, camera viewpoint, and robot embodiment with a TCP-centric crop
RoboSaGA Robust manipulation across visual shifts
Enlarged RoboSaGA manipulation policies operating across varied visual backgrounds
Task-aware visual encoders Same scene, different policy objective
Enlarged comparison of task-dependent encoder attention for reach-spam and reach-mug policies
MirrorDuo Reflection-consistent visuomotor learning
Enlarged MirrorDuo reflection-consistent visuomotor learning animation
Robot-centric Pooling Body ownership through image and proprioception alignment
Enlarged Robot-centric Pooling comparison of alignment scores and saliency maps

03 / Approach

A systems approach to robot learning.

I build robot-learning pipelines end to end, from data collection and ROS-based sensor integration to algorithm design, evaluation and deployment on physical systems.

01

Robotics systems

Integrating hardware and software through ROS, with multi-sensor synchronisation, communication, calibration, control and real-time execution.

ROSSensor synchronisationSystem integrationReal-time control
02

Robot learning

Designing, training and evaluating visuomotor and VLA policies across simulation and hardware, with attention to data efficiency, robustness and generalisation.

Algorithm designVLA fine-tuningImitation learningPolicy evaluation
03

Multimodal perception

Combining RGB, depth, event cameras, point clouds and proprioception with visual fundation models such as DINO and SAM.

DINO & SAMRGB & depthEvent camerasPoint clouds6D poseProprioception
04

Real-robot deployment

Designing data collection and evaluation pipelines, deploying policies and analysing the gap between simulated results and physical behaviour.

Data collectionManipulationSystem integrationFailure analysis

04 / Systems

Two examples of complete robotic systems.

Each brought perception, control, software and hardware together around a different robotics problem.

GoferBot collaborating with a person during furniture assembly

GoferBot

I co-developed GoferBot, a vision-based human–robot interaction system for collaborative furniture assembly. It combines human action recognition, visual servoing and handover behaviour so the robot can coordinate naturally with a person.

Watch GoferBot
Cartman warehouse picking robot Semantic segmentation of warehouse objects for the Cartman project

Cartman

I contributed as a member of the ACRV team that built Cartman, combining custom hardware with robotic perception for warehouse picking. The system won the 2017 Amazon Robotics Challenge, and the team leader later spun this work out into Lyro Robotics.

Amazon announcement Paper Lyro Robotics

05 / Mentoring

Supporting research beyond my own work.

I contribute through mentoring, teaching, editorial work and peer review.

Mentoring

Developing independent researchers

Mentored student projects contributing to CoRL, RA-L, and ICRA publications, and supervised undergraduate work resulting in an IROS paper.

Teaching

Connecting theory to hardware

Tutored five graduate courses and one undergraduate course, and led a Robotic Vision Summer School workshop later incorporated into teaching at Monash University.

Service

Editorial and peer-review service

Associate Editor for IROS 2025 and 2026; reviewer for IEEE RA-L, IEEE T-RO, RSS, CoRL, ICRA, and IROS.

06 / Publications

Earlier work across perception, control and robotic systems.

Before focusing on data-efficient visuomotor policy generalisation, I worked across closed-loop control, 3D perception, human–robot collaboration and integrated manipulation systems.

GoferBot: A Visual Guided Human-Robot Collaborative Assembly System Zheyu Zhuang, Yizhak Ben-Shabat, Jiahao Zhang, Stephen Gould, Robert Mahony IROS · 2022
A General Approach to State Refinement Gerard Kennedy, Jin Gao, Zheyu Zhuang, Xin Yu, Robert Mahony IROS · 2021
Stereo Hybrid Event-Frame Cameras for 3D Perception Ziwei Wang, Liyuan Pan, Yonhon Ng, Zheyu Zhuang, Robert Mahony IROS · 2021
End-to-end Multi-Instance Robotic Reaching from Monocular Vision Zheyu Zhuang, Xin Yu, Robert Mahony ICRA · 2021
6DoF Object Pose Estimation via Differentiable Proxy Voting Regularizer Xin Yu, Zheyu Zhuang, Piotr Koniusz, Hongdong Li BMVC · 2020
LyRN: A Real-Time Closed Loop Approach from Monocular Vision Zheyu Zhuang, Xin Yu, Robert Mahony ICRA · 2020
Learning Real-time Closed Loop Robotic Reaching from Monocular Vision by Exploiting a Control Lyapunov Function Structure Zheyu Zhuang, Jürgen Leitner, Robert Mahony IROS · 2019
Cartman: The Low-cost Cartesian Manipulator That Won the Amazon Robotics Challenge D. Morrison, A. W. Tow, M. McTaggart, R. Smith, N. Kelly-Boxall, S. Wade-McCue, J. Erskine, R. Grinover, A. Gurman, T. Hunn, D. Lee, A. Milan, T. Pham, G. Rallos, A. Razjigaev, T. Rowntree, K. Vijay, Zheyu Zhuang, C. Lehnert, I. Reid, P. Corke, J. Leitner ICRA · 2018
Semantic Segmentation from Limited Training Data A. Milan, T. Pham, K. Vijay, D. Morrison, A. W. Tow, L. Liu, J. Erskine, R. Grinover, A. Gurman, T. Hunn, N. Kelly-Boxall, D. Lee, M. McTaggart, G. Rallos, A. Razjigaev, T. Rowntree, T. Shen, R. Smith, S. Wade-McCue, Zheyu Zhuang, C. Lehnert, G. Lin, I. Reid, P. Corke, J. Leitner ICRA · 2018

07 / Off the clock

When I'm not chasing robots around a lab.

A few things that get me outside and away from a screen.

Zheyu climbing

Climbing

Helps me completely empty my head as the survival instinct screams, but oftentimes it's back to the familiar pattern of stopping half way and rethinking my approach.

Zheyu hiking

Hiking

Often question myself why I put myself through this — days without a shower — but in the end, the cleared mind, the breathtaking views, and being off the grid and off the internet make it all worth it.

Zheyu fishing

Fishing

Maybe it's the 10th time getting skunked, but all it takes is just one fish to keep me going.