Today for AI

量子位官网 · 10/9/2026, 12:15:05

Sharpa Unveils D01 Humanoid at IROS: Full-Body Tactile Skin and 10.5m/s Dexterous Hands Enable Closed-Loop Manipulation

By 贾浩楠Original title: 灵巧操作头号玩家:Sharpa把指尖「触觉」进化到全面「体感」
78AI Score
Executive Summary

Embodied AI startup Sharpa launched its first fully self-developed humanoid robot D01 at IROS, alongside the new dexterous hand W02 and exoskeleton data glove AE01. The D01 features full upper-body electronic skin (100Hz sampling, 0.2N force resolution) and high dynamic performance (>10.5m/s end-effector speed, 0.2mm repeatability), aiming to address real-world generalization challenges through an integrated hardware stack of hands, body, and data entry.

SOURCE COVERAGEOriginal coverage

Contents3 sections

What does it take for robots to move beyond trade shows and stage demos to generate real commercial value?

A dexterous hand capable of spinning walnuts and making ice cream is enough to make people remember Sharpa:

This embodied AI startup, founded by the team behind Hesai Technology, has recently been most distinctly characterized by its high-degree-of-freedom tactile dexterous hands.

At the ongoing IROS, Sharpa took this further with the launch of its first fully self-developed general-purpose humanoid robot, D01:

Alongside the new generation dexterous hand W02:

And the exoskeleton data glove AE01:

Hands, body, and data entry points—all displayed on the booth at once.

Why did Sharpa start building the robot body itself, and even develop the data entry point for "how to learn manipulation"?

Viewing these three products together, the answer points to a proposition that nearly everyone in the embodied AI sector is debating:

What exactly does it take for robots to escape trade show demos and generate real commercial value?

At IROS: What Are the Three Products?

First, let's look at D01, which Sharpa describes as an integrated tactile-sensing, dexterous-manipulation robot.

The key metrics worth noting about this robot all revolve around dexterous manipulation: the arm payload-to-weight ratio approaches 1, the maximum end-effector speed exceeds 10.5 m/s, the communication frequency reaches 1000 Hz, the repeatability accuracy is 0.2 mm, and the spherical wrist joint features a 1 human-scale design.

"High-speed motion" means the robot can quickly complete dynamic tasks such as catching batons or tool operations:

"High payload capacity" ensures the arm retains sufficient operational margin during movement; the 0.2 mm repeatability accuracy pushes capabilities further toward fine manipulation.

When facing complex tasks, speed, strength, and precision rarely exist in isolation. This is precisely where D01's design focus lies.

Tactile sensing is particularly distinctive.

D01 features electronic skin covering the entire body, with tactile coverage across the upper body. The tactile sampling rate reaches 100 Hz, the force sensing range is 0.1–20 N, and the force resolution is 0.2 N. As a result, the robot can directly perceive light touches, collisions, and external forces.

The reason Sharpa is so committed to "tactile sensing" becomes clearer when considering humanoid robots entering actual workspaces—specifically, dexterous manipulation environments:

Vision can tell the robot in advance that "there is something ahead," but after contact occurs, the robot still needs to know "where it touched, how much force was applied, and whether the contact changed."

For example, when catching a rapidly falling object, the robot must adjust the state of its arms and fingers within an extremely short time frame. Similarly, if a person working nearby bumps into the robot's arm, the robot needs to determine the source of the external force and decide whether to alter its current trajectory. For such instantaneous physical events, vision alone struggles to provide complete information.

The human body constantly performs this task. When fingers sense an object slipping, grip force increases subconsciously; when an arm touches a table, the next movement changes; when the body experiences an external force, posture adjusts immediately.

Many actions do not require pre-orchestrated textual "chain-of-thought" guidance for execution. The body itself provides real-time feedback. By deploying electronic skin on the arms and chest, D01 effectively incorporates this "body feedback" into the robot's perception system.

Full tactile coverage of D01's core interaction areas transforms the role of the robot's body within the control system: elevating it from simple anomaly detection or post-hoc logging to becoming an integral part of action adjustment.

Next, consider W02, the "new-generation fully tactile, ultra-compact, lightweight dexterous hand." Here, a somewhat surprising change emerges: while the previous generation W01 had 22 active degrees of freedom (DoF), W02 reduced this to 21.

The removed DoF is one at the CMC (carpometacarpal) joint of the little finger.

According to Sharpa's application research, this specific degree of freedom has low usage frequency in common grasping and in-hand manipulation tasks, contributing minimally to overall performance. Therefore, W02 simplified the corresponding structure, reducing mechanical complexity and potential failure points.

This is one of the most noteworthy aspects of W02: enhancing dexterous hand capabilities does not necessarily mean the number of degrees of freedom must continuously increase.

Additionally, W02 packs 21 active degrees of freedom into a more compact structure. The overall hand size is approximately 30% smaller than W01, and the weight is further reduced. This directly results in two benefits: decreased forearm load and easier integration into workspaces designed for human hand dimensions.

Being smaller and lighter also enables tasks that were previously difficult to accomplish.

For instance, W02 achieves a "zero grasp radius," allowing it to handle small-diameter objects like chopsticks and thin wires more stably. The fingertips utilize high-resolution visual-tactile sensors with a force sensing range of 5 mN–30 N and spatial resolution of 1 mm; other palm areas are covered by electronic skin with a force sensing range of 0.1–20 N and spatial resolution of 5 mm.

In other words, W02’s tactile coverage has expanded from "a few fingertips" to the entire hand. The robot can simultaneously acquire contact information from fingertips, fingers, and the palm, which is particularly critical for handling small objects and performing continuous in-hand manipulation.

What is noteworthy here is that Sharpa integrates "miniaturization" and "tactile coverage" into a single engineering constraint: the hand must fit within human-scale workspaces while retaining sufficient contact perception capabilities.

Finally, there is AE01, a "high-fidelity exoskeleton haptic data glove" primarily used for precise teleoperation and first-person perspective data collection.

AE01 uses 22 encoders to precisely capture the operator's natural hand movements and map them in real-time to the robot, enabling more natural and sensitive high-fidelity teleoperation.

The key lies in "feedback": when humans manipulate an object, their eyes perceive position and shape, while their fingers sense force, contact, and slippage. AE01 allows the robot to feed this information back to the human, enabling the operator to adjust actions based on tactile cues.

Simultaneously, the human's natural hand movements are captured by the robot in real-time.

Therefore, the value of AE01 primarily resides in "transforming human operational experience into data usable by robots": human action trajectories, robotic-side contact feedback, and the mapping between the two are integrated into a single acquisition process.

What Makes Sharpa’s Full-Stack Dexterous Manipulation Different?

If we examine these three products through specific technical challenges, the interesting aspect of Sharpa’s IROS lineup is its attempt to evolve "dexterous manipulation" from isolated hardware capabilities into a comprehensive technical system built around contact perception, closed-loop control, and data learning.

To understand this system, one should start by breaking down the question: "When does a robot need touch?"

Traditional vision-based strategies excel at the pre-contact phase: identifying objects, estimating positions, planning arm trajectories, and reaching out. The problem arises precisely in the final few centimeters, or even millimeters.

For example, during insertion, twisting, alignment, or grasping flexible objects, the variables determining task success often lie hidden after contact occurs: whether the object has shifted, slipped, experienced excessive pressure, or changed contact points. Vision can observe part of the outcome but struggles to stably monitor these instantaneous physical states.

Sharpa’s CraftNet model addresses this issue with clear delineation: System 2 handles task understanding and long-horizon planning; System 1 manages pre-contact action planning; and System 0 enters the contact phase, processing tactile feedback and fine motor actions at higher frequencies. System 0 not only executes fine-grained actions but also feeds state information back to System 1, allowing upper-level actions to be re-adjusted upon execution failure or changes in physical state.

This means that the role of touch here effectively alters the time scale of the control loop.

Before contact, the robot can use slower models to reason about "what to grasp" and "where to approach from"; after contact, it needs to determine "has it slipped," "should I grip harder," or "which way should the fingers move." The latter requires the control system to operate at higher frequencies.

Taking it a step further, the question becomes: How does a robot remember a single contact experience?

This is why Sharpa’s latest WM-Craftnet research is more noteworthy than simple "vision + touch fusion."

Compared to previous approaches that simply concatenated depth maps, tactile signals, and joint states to feed into policies, the new research trains a World Synesthesia Model—essentially a world model with temporal memory.

Inputs include wrist depth, tactile signals, proprioception, and previously executed actions. Through a Dreamer-style recurrent state-space model, this information is compressed into a time-varying latent state. This state serves as context for the policy, helping it infer the object’s current geometry, contact status, and motion state.

The model learns not just "what is being touched now," but combines past events to infer "why the current state exists."

The significance of this step moves beyond "touch participating in control" to another level: the robot begins to form an internal representation of contact states.

Returning to the initial question: Why has Sharpa consistently been "obsessed" with placing touch at the core of dexterous manipulation?

Sensors provide raw contact signals, control systems handle reactions at millisecond to hundred-millisecond scales, and world models attempt to organize these instantaneous events into "physical states" that robots can continuously utilize—

This is precisely what distinguishes dexterous manipulation from standard robotic arm trajectory execution.

AE01 completes the final piece of the puzzle on the data side.

Traditional teleoperation primarily records human action trajectories but struggles to fully preserve the physical experiences during contact processes.

AE01 captures natural hand motions via encoders while feeding robotic-side tactile states back to the operator, forming a bidirectional loop. This reduces the friction of mapping human intuition to machine execution and lowers the barrier for collecting high-quality, contact-rich demonstration data.

Sharpa is tackling a very real problem in embodied intelligence: high-quality manipulation data is hard to obtain, yet the most valuable data occurs precisely during the moments of most complex contact.

From this perspective, Sharpa’s “full-stack” capabilities for dexterous manipulation are indispensable: the robot body brings the machine into real-world workspaces, the dexterous hand handles fine-grained contact between humans and objects, and the data entry point preserves human operational experience alongside the robot’s physical feedback.

Finally, the model leverages this information to transform each operation into capability for the next iteration.

The Leading Player in Dexterous Manipulation

The hand is, first and foremost, the most reliable organ humans use to complete various complex tasks.

It is also the most direct medium through which robots engage in precise physical interactions with the real world.

Therefore, it comes as no surprise that since W01, Sharpa has long focused on high degrees of freedom, tactile sensing, and fine manipulation.

The demo at CES earlier this year made a splash, showcasing a ceiling-level live demo of dexterous manipulation for the industry.

However, once dexterous hands reach a certain level of maturity, questions naturally extend to both ends of the system: What kind of physical embodiment does this hand need to enter diverse working environments? And what kind of data is required to enable robots to generalize from one example to many, achieving “commercial generalization” across different task scenarios?

These challenges include control precision during high-speed motion, perception of real-world contact, human-like tool usage, and the ability to handle complex objects such as thin sheets, flexible items, and liquids—all while ensuring that the body, wrist, fingers, and tactile feedback coordinate seamlessly throughout continuous tasks.

True opportunities to integrate these capabilities only arise in real-world tasks.

In late August this year, the Blizzard Robot Restaurant, a collaboration between Sharpa and DQ, opened on Wujiang Road in Shanghai. Rather than designing dedicated stations for the robots, the setup directly utilized DQ’s existing equipment, ingredients, and production processes. The robots completed 55 consecutive steps—from opening cabinet doors and fitting cup rings to dispensing ice, sprinkling toppings, stirring, and pouring cups—

Viewed individually, none of these steps seem particularly “sci-fi,” but the difficulty lies in their continuity, where errors from previous steps propagate to subsequent ones.

The biggest challenge was completing a highly interdependent sequence of operations continuously within an unmodified environment.

This scenario perfectly corresponds to the technical chain Sharpa had previously built.

Long-horizon tasks are decomposed and managed by the model; System 1 executes specific actions; upon entering contact, System 0 uses tactile and force feedback for continuous correction; meanwhile, the world model integrates vision, touch, force feedback, and proprioceptive state to determine the current stage of the task and plan the next step.

This marks Sharpa’s second breakthrough in raising the ceiling of dexterous manipulation within the year.

From dexterous hand hardware to tactile control, and further to world models and real-store tasks, Sharpa is gradually expanding dexterous manipulation from a capability of end-effectors into a comprehensive system capability.

Vision solves “what is seen,” while touch supplements “what was contacted, how much force was applied, and whether slipping occurred.” Together, they support action adjustments in continuous tasks such as grasping, plugging/unplugging, twisting, and squeezing.

Architecturally, tactile sensing exists as a relatively independent computational layer. Since upper-layer models can be replaced, tactile capabilities can accumulate continuously without being tied to a specific generation of models. Consequently, the robot gains access to a set of contact information that persistently participates in perception and control.

On the other end lies data.

Sharpa co-founder Li Yifan once told QbitAI: “Purchased data cannot build a competitive moat.”

This statement becomes easier to understand in the context of today’s product layout: Compared to any single model or data collection method, what Sharpa prioritizes is enabling robots to continuously generate “action–contact–feedback–correction” experiences in real tasks, and then feeding these experiences back into the model.

This is a path that starts with tasks and subsequently defines hardware and models.

It also reflects the “commercial generalization” emphasized by Sharpa’s founding team—who have been tempered by the mass-scale deployment of autonomous driving—in the realm of embodied intelligence: Can the system continue to function after switching scenarios or categories? And can efficiency improvements cover procurement, deployment, maintenance, and data costs?

From this perspective, the most noteworthy change in Sharpa’s recent IROS presentation may not be “launching three new products.”

Previously, the W01 dexterous hand addressed whether “robots could grasp and manipulate objects well.” With the introduction of D01, W02, AE01, and the real-world deployment at DQ, Sharpa is continuing to probe and explore a more fundamental layer:

Extracting universal hardware and technical requirements for diverse task scenarios, providing a suite of capabilities capable of handling real-world operations.