A pair of dexterous hands that can spin walnuts and make ice creamis already memorable enough.Sharpa:

is an embodied intelligence startup founded by a team that set out again from Hesai Technology. Over the past period, its most distinctive label has been high-DOF tactile dexterous hands.
At the ongoingIROS, Sharpa pushed this one step further, with the arrival of its firstfully in-house developed general-purpose humanoid robot D01:

Along with a new-generation dexterous hand W02:

and the exoskeleton data glove AE01:

Hand, body, and data entry point, all laid out on the booth at once.
Why has Sharpa started building the robot body itself, and even the data entry point for 'how to learn manipulation'?
Looking at these three products together, the answer actually points to the question almost everyone in the embodied intelligence field is debating right now:
What does it really take for robots to move beyond trade-show and stage demos and generate real commercial value?
What the three products are at the IROS site
First look at D01, which Sharpa describes as an integrated tactile-perception dexterous manipulation robot.
The most noteworthy numbers for this robot all revolve around dexterous manipulation: the arm payload-to-self-weight ratio is close to 1:1, maximum end-effector speed exceeds 10.5m/s, communication frequency reaches 1000Hz, repeat positioning accuracy is 0.2mm, and the spherical wrist joint uses a 1:1 human-scale design.
Among these, 'high-speed motion' means the robot can quickly complete dynamic tasks such as catching a baton or tool operation:
'high payload capacity' gives the arms sufficient operating margin during movements; and the 0.2mm repeat positioning accuracy pushes the capability further toward fine manipulation.
When a robot faces complex tasks, speed, strength, and precision rarely exist in isolation, and D01's design focus lies precisely here.
What's even more special is the touch。
D01's electronic skin covers the whole body, with tactile coverage on the upper body, a tactile sampling rate of 100Hz, a force sensing range of 0.1—20N, and a force resolution of 0.2N. The robot can therefore directly perceive light touches, collisions, and external forces.

The reason Sharpa is so obsessed with 'touch' becomes clearer only when placed in the context of humanoid robots actually entering workspaces — that is, dexterous manipulation environments.:
Vision can tell the robot in advance that 'there is something ahead,' but once contact occurs, the robot still needs to know 'where it was touched, how much force was applied, and whether the contact has changed.'
For example, when catching a fast-falling object, the robot needs to adjust the state of its arms and fingers within an extremely short time; or when a person working nearby bumps into the robot's arm, the robot needs to determine where this external force came from and whether it needs to change its current trajectory. For these instantaneous physical events, vision alone can hardly provide complete information.
The human body is actually doing this all the time. When fingers sense an object slipping, they subconsciously increase grip strength; when an arm hits a table, it changes the next movement; when the body receives an external force, it immediately adjusts posture.
A large number of actions do not need to be choreographed into a textual 'chain of thought' to guide execution; the body itself provides real-time feedback. D01 lays electronic skin over the arms and chest, effectively incorporating this kind of 'bodily feedback' into the robot's perception system.
With full tactile coverage of D01's core interaction areas, what changes is the role of the robot's body in the control system: elevated from mere anomaly detection or after-the-fact logging to a part of motion adjustment.

Now look atW02, the "next-generation fully tactile, ultra-compact, lightweight dexterous hand". Here a somewhat surprising change appears:The previous-generation W01 had 22 active degrees of freedom, while W02 has actually been reduced to 21。
. What was removed is one degree of freedom at the CMC joint at the base of the little finger.
According to Sharpa's application research, this degree of freedom is used relatively infrequently in common grasping and in-hand manipulation tasks and contributes little to them, so W02 simplified the corresponding structure, reducing both mechanical complexity and potential points of failure.
This is also one of the most noteworthy aspects of W02:continuing to increase a dexterous hand's capabilities does not necessarily mean the number of degrees of freedom must keep climbing。

. In addition, W02 packs 21 active degrees of freedom into a more compact structure: the overall hand size is about 30% smaller than the W01, and the weight has also been further reduced. This directly brings two changes: forearm payload drops, and it becomes easier to fit into workspaces designed around human hand dimensions.
Being smaller and lighter also enables some tasks that were previously hard to accomplish.
For example, W02 achieves a "zero grip radius," enabling more stable handling of small-diameter objects such as chopsticks and thin threads. The fingertips use high-resolution visual-tactile sensors with a force sensing range of 5mN–30N and a spatial resolution of 1mm; the other areas of the palm are covered by electronic skin with a force sensing range of 0.1–20N and a spatial resolution of 5mm.

In other words, the W02's tactile coverage has expanded from "a few fingertips" to the entire hand, allowing the robot to simultaneously obtain contact information from the fingertips, fingers, and palm, which is especially important for small objects and continuous in-hand manipulation.
What is noteworthy here is that Sharpa has placed "miniaturization" and "tactile coverage" into the same engineering constraint: a hand must both fit into a human workspace and retain sufficient contact sensing.
Finally, there isAE01, this is a 'high-fidelity exoskeleton tactile data glove,' mainly used for precise teleoperation and first-person-view data collection。
The AE01 precisely captures the operator's natural hand movements through 22 encoders and maps them to the robot in real time, achieving more natural, more responsive, high-fidelity teleoperation.

The key lies in "feedback": when a person operates an object, the eyes see position and shape, while the fingers can also feel force, contact, and sliding. AE01 lets the robot feed this information back to the human, so the operator can adjust movements based on touch.
At the same time, the person's natural hand movements can be captured by the robot in real time.
Therefore, the value of AE01 mainly lies in the step of "turning human manipulation experience into data usable by robots": the human motion trajectories, the contact feedback on the robot side, and the mapping between the two are put into a single collection process.
What exactly makes Sharpa's dexterous manipulation full stack different
If we look at the three products in terms of specific technical problems, what is interesting about Sharpa's three products at IROS is the attempt to move "dexterous manipulation" from several isolated hardware capabilities toward a technical system built aroundcontact perception, closed-loop control, and data learning.
But to understand this technical system, one should break it down starting from the question "when does a robot need touch".
What traditional vision-based strategies are best at is the stage before contact occurs: recognizing objects, estimating positions, planning an arm trajectory, and then reaching the hand over. The problem arises precisely in the last few centimeters, or even the last few millimeters.

For example, when inserting, twisting, aligning, or grasping soft objects, the variables that determine whether the task can be completed are often hidden after contact: has the object shifted, has it slipped, is the pressure too high, has the contact point changed. Vision can see part of the result, but it is hard to reliably observe these instantaneous physical states.
Sharpa's CraftNet model handles this problem very clearly: System 2 is responsible for understanding the task and long-horizon planning, System 1 is responsible for pre-contact motion planning, and System 0 enters the contact phase, processing tactile feedback and fine movements at a higher frequency. System 0 not only executes fine movements, but also feeds state information back to System 1, allowing the higher-level motion to readjust when execution fails or the physical state changes.
This means thatthe role of touch here is actually to change the time scale of the control loop。
Before contact, the robot can use a slower model to think "what do I want to grasp, from which direction to approach"; after contact, what is needed is "is it slipping now, should I grip a bit more, which way should the finger move a little". The latter requires the control system to operate at a higher frequency.
And then taking one step further, the question becomes: how can a robot remember a single contact experience?
This is also where Sharpa's latest WM-Craftnet research is more noteworthy than simple "vision + touch fusion".
Compared with previously simply concatenating depth maps, touch, and joint states and feeding them to a policy, the new research trains a World Synesthesia Model, that is, a world model with temporal memory.
The inputs include wrist depth, touch, proprioception, and previously executed actions, which are compressed through a Dreamer-style recurrent state-space model into a hidden state that changes over time. This state then serves as context for the policy, used to judge the object's current geometry, contact, and motion state.
What the model learns is not just "what it is touching now"; it must also combine what happened in the past to infer "why it has become like this now".
The significance of this step is that itmoves one layer further from "touch participating in control": the robot begins attempting to form an internal representation of the contact state。
Returning to the original question: why has Sharpa always been "obsessed" with touch, placing it at the core of dexterous manipulation?
Sensors provide raw contact signals, the control system reacts at the millisecond to hundreds-of-milliseconds level, and the world model attempts to organize these instantaneous events into "physical states" that the robot can continuously use—
this is precisely what distinguishes dexterous manipulation from ordinary robotic arm trajectory execution.
AE01 completes the last piece of the puzzle on the data side.
Traditional teleoperation mainly records human motion trajectories, but it is difficult to fully preserve the physical experience during the contact process.
The AE01 captures natural hand movements through encoders while feeding tactile state from the robot side back to the operator, forming a bidirectional interface: human actions go into the robot, and the robot's state after contacting objects in turn influences the human's next move.
In this way, the data left by a single human operation no longer contains only "where the hand moved," but also the moment of contact, changes in force, sliding states, and the corrective actions that followed. Precise control, synchronized tactile capture, and calibration-free operation lower the barrier for human-robot mapping and multimodal data synchronization.
So the new progress at IROS really boils down to one sentence:
What Sharpa is trying to solve is a very practical problem in embodied intelligence: high-quality manipulation data is hard to obtain, and the data that is truly valuable happens precisely at those moments when contact is most complex。
Viewed from this angle, every part of Sharpa's dexterous manipulation "full stack" capability is indispensable: the robot body brings the robot into real workspaces, the dexterous hand handles the fine-grained contact between humans and objects, and the data entry point preserves both human operational experience and the robot's physical feedback.
Finally, the model reuses this information, turning each operation into capability for the next round.
The Top Player in Dexterous Manipulation
The hand is, first of all, the most reliable organ humans use to accomplish all kinds of complex tasks.
It is also the most direct medium through which robots perform fine-grained physical interaction with the real world.
So it is not surprising that, starting from the W01, Sharpa has long focused on high degrees of freedom, touch, and fine manipulation.
At the CES demo early this year, it made a stunning debut,as Sharpa showed the industry a ceiling-level live demo of dexterous manipulation.
But once a dexterous hand reaches a certain level, the questions naturally extend to both ends of the body: what kind of physical platform does this hand need in order to enter various work environments? And what kind of data does it need so that a robot can draw inferences and achieve “commercial generalization” across different task scenarios?
The problems involved include control precision under high-speed motion, real contact sensing, human tool manipulation, and the ability to handle complex objects such as thin sheets, flexible objects, and liquids, while also making the body, wrist, fingers, and tactile feedback work together in continuous tasks.
The opportunity to truly string these capabilities together will only appear in real tasks.
At the end of the 8th month of this year,the Blizzard robot restaurant opened by Sharpa in partnership with DQ on Wujiang Road in Shanghaidid not build a separate dedicated workstation for the robot; instead, it directly used DQ's existing equipment, ingredients, and production process, with the robot completing 55 continuous steps from opening the cabinet door, putting on the cup ring, catching ice, adding toppings, stirring, to pouring into the cup—
Any single step taken alone is far from “science fiction,” but the difficulty is that they happen in succession, and errors left by the previous step carry over into the next.
The biggest challenge is completing a series of highly interrelated operations in succession in an unmodified environment.
And this corresponds exactly to the technology chain Sharpa built up earlier.
Long-horizon tasks are decomposed and state-managed by the model; System 1 handles the specific motions; once contact begins, System 0 uses tactile and force feedback for continuous correction; and the world model combines vision, tactile and force feedback, and proprioceptive state to determine where the task currently stands and what to do next.
And this is alsothe second time within the year that Sharpa has raised the ceiling of dexterous manipulation.。
From dexterous hand hardware, to tactile control, and then to the world model and real store tasks, Sharpa is gradually expanding dexterous manipulation from an end-effector capability into a full set of system capabilities.
Among these, vision solves “what is seen,” while touch supplements “what was touched, how much force was applied, and whether slipping occurred,” jointly supporting motion adjustment in continuous tasks such as grasping, plugging, twisting, and squeezing.
In terms of technical architecture, touch exists as a relatively independent computational layer, and the upper-layer models can be replaced, which means the tactile capability can keep accumulating without being tied to any particular model generation; what the robot thereby gains is a set of contact information that can continuously participate in perception and control.
At the other end is the data.
Sharpa co-founderLi Yifanonce told QbitAI: “Purchased data cannot build a gap.”
This statement is easier to understand in today's product layout: rather than any single model or any single data collection method,what Sharpa cares about is having robots continuously produce “action—contact—feedback—correction” experience in real tasks, and then feeding that experience back into the models.。
This is a route that starts from tasks and in turn defines the hardware and models.
This is also the “commercial generalization” that Sharpa's founding team, tempered by the large-scale mass production of autonomous driving, continues to emphasize in embodied intelligence: whether things keep working after changing scenes and product categories, and whether efficiency gains can cover the costs of procurement, deployment, maintenance, and data.
From this perspective, the most noteworthy change in Sharpa's showing at IROS this time may not be “three more products released.”
Whereas the W01 dexterous hand solved “whether a robot can hold and manipulate things well,” with the real deployments of the D01, W02, AE01 and DQ, what Sharpa continues to question and explore is a more fundamental layer:
For different task scenarios, distill general hardware and technical requirements, providing a set of general operational capabilities capable of handling real positions, continuous delivery, data accumulation, and iterative improvement.