Robot T-Shirt Folding Looks General—Until the Gripper Meets a Sleeve

Robot laundry demonstrations have become more capable, but their meaning has not fundamentally changed: folding a T-shirt is evidence of difficult manipulation, not proof that a machine can handle arbitrary household work. Newer systems can transfer learned behavior across more tasks and robot bodies, yet their published results remain tied to particular hardware, training data and evaluation setups.
The most useful change since the earlier wave of folding videos is that researchers are now testing beyond neatly presented shirts. Keys, doors, socks, tools and slippery objects expose limitations that a successful fold can conceal—especially inconsistent execution and grippers that physically cannot complete a task.
Why folding a T-shirt occupies robotics’ sweet spot
A shirt is much harder for a robot than a rigid box. Its shape changes whenever it is lifted, one layer can hide another, and the same garment can begin crumpled, inside-out, partly folded or tangled with other items. A robot must infer where useful edges and corners are while its own actions keep changing the scene.
That combination makes folding technically meaningful and visually legible. Viewers can immediately recognize the goal, follow the sequence and judge the final result. The task also fits on a table, can be repeated without a hazardous environment and allows developers to constrain garment type, starting position and workspace when necessary.
This is why folding sits in an optimal demonstration zone: difficult enough to exercise perception, two-arm coordination and corrective control, but bounded enough to produce a concise presentation. It should not be confused with a formal industry benchmark or with evidence that the same system can independently perform every stage of laundry.
A broad academic review of robotic textile manipulation explains that progress is often specialized to particular textiles or textile categories. The review of cloth-manipulation research identifies variations in shape, weight, stiffness, elasticity and friction as central obstacles to generalization, while also noting gaps in datasets and standardized evaluation.
A completed fold proves a chain of skills—not universal competence
Even a controlled fold can require the robot to locate grasp points, separate fabric layers, coordinate two manipulators and monitor whether the material moved as expected. Starting with a crumpled pile raises the difficulty because the system cannot simply replay one geometric trajectory. Recovery after a poor grasp or a disturbed garment is particularly valuable evidence because it shows closed-loop behavior rather than a perfectly staged motion.
But the visible result does not reveal all the conditions that produced it. A video alone may not show how much task-specific demonstration data was collected, which garments were excluded, how often the robot failed, whether the sequence was selected from multiple attempts or whether a human reset the workspace between trials. Autonomy during one recorded run also does not mean that learning, setup or recovery from every possible failure was autonomous.
The correct conclusion is therefore narrower: a successful fold demonstrates that a specific combination of model, sensors, actuators, end effectors and training procedure can execute that version of the task. Claims of versatility require evaluations across unfamiliar garments, initial states, environments and robot bodies, with failures included and success criteria stated.
The sleeve test reveals where software ends and hardware begins
Physical Intelligence supplied an unusually clear example in December 2025 by applying a π0.6-based policy to a set of everyday “Robot Olympics” tasks. Its published task evaluation reported a 52% average success rate and 72% average task progress; the company also said most tasks used under nine hours of collected data.
The system demonstrated initial solutions at the highest proposed level in three of five categories, including passing through a self-closing door, using a key and cleaning a greasy pan. In laundry, however, it reached the sock task rather than the highest-level shirt challenge. The stated reason was concrete: its gripper was too wide to enter the sleeve of the inside-out dress shirt that had to be corrected and hung.
That failure is more informative than another clean T-shirt fold. It shows that a capable learned policy cannot compensate for an end effector that cannot reach the required contact point. The orange-peeling challenge produced a similar boundary: adding a metal tool made the action feasible, but violated the proposed task rules, so the company did not count it as a successful highest-level result.
Hardware and intelligence are therefore not competing explanations. The model must choose and adjust useful actions, while the body must supply suitable reach, force control, contact geometry and sensing. A demonstration evaluates the combined system; attributing the result to the AI model alone discards part of the causal chain.
Newer generalist models broaden the test, but the scores stay task-specific
The frontier has continued moving beyond laundry. Google DeepMind’s current Gemini Robotics 2 model page describes whole-body control across tabletop and humanoid platforms, but lists the model in private preview rather than as a generally available household product.
Its displayed evaluations also illustrate why one headline number cannot summarize versatility. For an Apollo humanoid with multi-finger hands, the listed accuracies range from 32% for a dustpan task to 92% for unscrewing a bulb; other hand and gripper configurations have their own task sets. These results show meaningful breadth, but they compare distinct behaviors performed with specified embodiments—not a universal score for “general intelligence.”
This newer evidence changes the practical interpretation of folding videos. Laundry is no longer the outer boundary of what learned robot policies can demonstrate. It remains one informative member of a larger evaluation portfolio, while model access, reliability and transfer to uncontrolled homes are separate questions.
What convincing versatility would look like
A stronger test would keep the clarity of folding while varying the conditions that make real laundry unpredictable. Different materials, garment sizes, tangled starting states, sleeves, fasteners and partly completed folds would test whether the system recognizes changes that matter mechanically. Requiring the same policy to recover from slips and misplaced folds would reveal whether it can correct its own work.
Evaluation should also separate three ideas that polished demonstrations often compress into one:
- Task completion: Did the configured robot finish the defined job under the stated conditions?
- Robustness: How often did it succeed across repeated trials, disturbances and plausible starting states?
- Transfer: Did learned capability carry to new objects, tasks, environments or robot bodies without extensive retraining?
No single household chore can establish all three. A T-shirt fold remains a legitimate technical achievement because cloth is perceptually and mechanically difficult. Its value increases when it is reported alongside failure rates, garment diversity, training requirements and tests that stress different physical capabilities.
The enduring lesson is not that folding demonstrations are trivial or deceptive. It is that they are a carefully bounded window into physical intelligence. When a robot moves from a flat shirt to an obstructed sleeve, a key that must be reoriented or a wet sponge that changes contact conditions, the gap between a learned routine and dependable versatility becomes visible.
Also read:
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.