Seeker: Attention from Action, for Action
Learns where to look directly from action supervision.
Seeker uses the task and robot state to produce task-aware regions of interest, without spatial labels. Instead of treating every pixel equally, it identifies the evidence needed for the next action. The learned attention supports RGB cropping, background augmentation and point-cloud filtering across changing scenes.