FoldNet++: a Large-Scale Synthetic Dataset for Robotic T-Shirt Folding and Unfolding
Abstract
Due to the highly deformable nature of garments, training a generalizable policy for robotic T-shirt folding and unfolding remains a significant challenge. In this work, we present a large-scale synthetic dataset for robotic T-shirt folding and unfolding, covering 6 robotic embodiments, 1K T-shirts, 1K environmental assets, and 120K episodes with rich annotations, which can be used to train a wide range of manipulation policies. We first follow the FoldNet pipeline to generate a large-scale dataset of physically simulatable T-shirts with diverse appearances and annotated semantic keypoints. Based on these semantic keypoints, we then generate manipulation demonstrations for different robotic embodiments through a unified rule-based framework. We use these demonstrations to train visuomotor policies, and experimental results demonstrate that models trained solely on our synthetic data can achieve over 90% end-to-end task success rates when directly deployed to unseen real-world environments and previously unseen T-shirts from arbitrary initial configurations.
Method
Overview of FoldNet++. (a) Our dataset includes 6 robots, 1K T-shirts, and 1K table textures and HDRI environment maps. (b) Using these assets, we import them into an FEM-based physics simulator. Combined with predefined manipulation rules and domain randomization, we generate 120K episodes containing multi-view observations from three camera viewpoints, robot actions, and subtask annotations. (c) Leveraging the FoldNet++ dataset, we train different categories of models (VA, VLA, and WAM) and evaluate them across multiple embodiments in both simulated and real-world environments under zero-shot and few-shot settings.
Results
Deployment on Galbot-G1
Deployment on ARX-R5
More Videos
Galbot-G1
ARX-R5
Large-scale real-world testing on Galbot-G1
Large-scale real-world testing on ARX-R5