Author丨Wu Simeng
Editors丨Cen Feng, Ma Xiaoning
If you have played 《Mario Tennis》 or 《TopSpin》, you probably would not think tennis is a particularly complicated thing. The ball flies over, you move the character, judge where it will land, and press the hit button at the right moment. Especially in 《Mario Tennis Aces》, the game even breaks tennis down further into a very intuitive language of movements: forehand, backhand, topspin, slice, lob, and even special shots that require precise timing to execute. Mario Tennis Aces game screenshot. Real-world tennis is of course far more complicated than all this; what the game does is“action mapping”, essentially mapping a very abstract “button input” from the player into an entire complex set of the character's bodily movements. The player does not need to know how the motion is performed, because the motor control in between is taken over by the game system. When the ball lands, when the body rotates its hips, which way the feet should step, how fast the arm should swing—these things are almost entirely hidden in the game. What the player sees is just a simple input, and the game engine is responsible for translating it into a whole chain of complete bodily movements. No matter how fast the ball flies, the character will never become unsteady from failing to shift its center of gravity in time, let alone lose its balance entirely because its feet were a step slow. But if you replace this “press a button and swing” character with a real humanoid robot, “embodiment” will make things completely different.Embodied intelligence, by contrast, requires the robot to take back on the bodily process that the game hid—that is, the gap from intent to action.It must first judge where the ball is coming from, then decide when to move and how much; as it approaches the contact point, its legs, hips, torso, shoulders, arms, and wrists must coordinate. The racket may touch the tennis ball for only a very brief instant, and within that time the robot must also handle a series of changing factors: its own posture, joint velocities, friction, racket weight, and the incoming ball's speed. Tennis therefore becomes a very interesting robotics problem. It is not like grasping a stationary cup in a laboratory. The ball is always in flight, the contact point is always changing, and the robot cannot just get one motion right once—it has to get it right repeatedly amid constantly changing states. What it needs is no longer just “knowing what to do,” but making its entire body form a coordinated response within a few hundred milliseconds or even less. This is also the question that IROS 2026 work from the Galbot (银河通用) team, Tsinghua University, Peking University, the Shanghai Qi Zhi Institute, and the Shanghai Artificial Intelligence Laboratory research team—LATENT—tried to answer. The paper's title itself is quite interesting:Training Athletic Humanoid Robots in Tennis Skills Based on Imperfect Human Motion Data. The keyword is not “perfect,” but rather “imperfect.” The research team did not prepare for the robot a set of carefully curated, complete, and accurate data from professional human tennis matches. Instead, they deliberately started from imperfect human motion data, letting the robot learn the most fundamental bodily capabilities of tennis, and then combining these capabilities, through reinforcement learning and simulation, into policies that can actually hit the ball. In other words, what they wanted to solve was not “how to completely copy human motions to a robot,” but a question closer to reality:if the demonstrations humans give the robot are inherently incomplete and imprecise, what can it still learn from them?
01
Five hours of imperfect data
If you really wanted a robot to learn from a human tennis player, the most intuitive approach would be to record that player in full. From the ready position, to moving, swinging, hitting, recovering, and then to the next shot. Ideally there would also be large amounts of data with different incoming balls, different landing points, different body postures, even complete matches. That way, what the robot sees would not just be “how the hand should swing,” but a complete set of tennis behavior. The problem is that such data is extremely hard to collect. Not to mention that if the target is a humanoid robot, the human body and the robot's own body structure are not the same. A human's arm length, joint ranges, body weight, and muscle strength all differ from the robot's; even if a motion capture system accurately records a person's movements, there is no guarantee the robot can execute them exactly the same way. LATENT took a different path. The research team recruited 5 amateur tennis players and, in a motion capture area of roughly 3 meters × 5 meters, recorded about 5 hours of human motion data. This space was not even a small fraction of a standard tennis court. Rather than asking the players to play complete matches, the researchers focused on some of the most fundamental bodily motions in tennis, including forehands, backhands, and movement actions such as side shuffles and crossover steps. Human tennis motion capture. These motions look fragmentary. A person takes two steps to the side, a person completes one forehand swing, a person does one backhand motion. They cannot be pieced together into a complete tennis match, let alone tell the robot “where the next ball should be hit.” But the research team believes these fragments still contain one very important thing:how humans use their own bodies when facing the sport of tennis.This is also what the paper calls “imperfect human motion data.” The “imperfection” exists on at least two levels.First, it is not precise enough.The human motion obtained through motion capture does not mean the robot can copy it without modification, especially in local motions such as the wrist, where there are clear differences between human data and the robot's body.Second, it is not complete enough.What the research team collected were basic tennis motions, not a complete set of match behaviors. The data can tell the robot how to move and how to swing, but it will not directly tell it what strategy to choose when facing different incoming balls. The traditional mindset easily treats these shortcomings as problems: if the data is not accurate enough, keep collecting; if the data is incomplete, collect more. LATENT's thinking, however, takes a slight step sideways. They did not try to patch this data into a “perfect manual of human tennis motion,” but treated it as a prior. The robot does not need to become a replica of some specific human athlete; what it first needs to know is roughly how humans use their bodies in a high-speed sport like tennis. This distinction matters. Because for reinforcement learning, exploring all possible bodily motions from scratch is extremely costly. A humanoid robot with 29 degrees of freedom does not have only two choices, “swing” and “don't swing.” At every moment it could change the state of its legs, hips, torso, shoulders, and arms. Starting entirely from scratch, it would not even know what counts as a reasonable bodily movement. Human motion data, even if just fragments, can first draw a rough boundary for it. It is like a person picking up a tennis racket for the first time. You do not need to give them a complete professional training course first; they only need to know “this is roughly how humans stand, move, and swing,” and only then is it possible to find their own style through continuous trial and error. What LATENT does can roughly be understood as giving the robot this kind of “bodily intuition” first.02
What the robot needs is a set of “body language”
The research team did not simply feed the collected human motions to it joint by joint. The real difficulty lies precisely here:human motion is not robot motion.When a person swings, the brain does not explicitly compute “how many degrees the left knee bends, how many degrees the right hip rotates, how many degrees the torso tilts.” What ultimately forms is a kind of bodily coordination. If the robot had to directly decide how all 29 degrees of freedom should move every time, the learning space would be enormous, and unnatural or even unstable motions could easily arise during reinforcement learning. LATENT method framework diagram. LATENT therefore first trained a motion tracker, enabling the robot to learn the bodily movement patterns contained in human motions. Then the research team further compressed this motion tracker into a continuous latent action space, the so-called latent action space. A more intuitive analogy: it is a bit like a “motion dictionary.” The robot does not have to compute from scratch how the 29 joints should coordinate every time; instead, it can first pick an already learned bodily movement from this dictionary, and then adjust it according to the ball in front of it. This action space has another very important feature: it does not lock the robot to human motions. The human demonstrations only provide a prior for motion; when actually facing the ball, the robot still has to learn on its own. This is why the paper does not treat “imitating humans” as the ultimate goal. Suppose a human player has a wrist motion when swinging, and this motion is not reliable for the G1—then the robot has no need to cling to it. The research team even deliberately processed the right wrist: lowering its reliability during the motion tracking stage, so that the robot cannot over-rely on this error-prone local motion, and then letting the high-level policy correct it when actually hitting the ball. This design is like deliberately leaving a slightly loose screw in the robot. If a system can only work when that screw is perfectly accurate, then a little real-world error could make the entire motion fail. Conversely, if the robot is forced from the start to learn to accomplish motions through the legs, hips, torso, shoulders, and other body parts working together, then when the wrist deviates, it still has a chance to adjust its whole body and hit the ball back. This is also one of the interesting aspects of LATENT:human data is not the answer the robot ultimately pursues, but rather more like a starting point for the robot when learning bodily movement.Next, the research team also had to face a very typical problem in reinforcement learning. Once the robot possesses a “motion dictionary,” will the reinforcement learning, in order to raise the return success rate, start arbitrarily jumping between completely different motions in the dictionary? The answer is: possibly. So the research team proposed the Latent Action Barrier, or LAB. It constrains the actions generated by the high-level policy, keeping the policy from easily deviating from the original action distribution. Latent Action Barrier. The “Barrier” here does not lock the robot down; rather, it draws a boundary during exploration. The paper uses the Mahalanobis distance to measure action deviation, rather than the simple Euclidean distance. The reason is easy to understand: the scales and ranges of variation of different action dimensions are not the same, so changes in all directions cannot be treated as the same degree of “deviation.” Experiments also demonstrated the value of this constraint. In the forehand hitting task, the complete LATENT policy achieved a success rate of 96.52%; after removing LAB, the success rate dropped to 93.12%. But the more obvious change appeared in motion quality: the motion smoothness metric rose from 25.61 to 37.64, and joint torques also rose from 7.40 to 12.53. In other words, what LAB truly improves is not just “whether the ball can be hit back,” but whether the robot, while completing the task, can have fewer sudden, violent, mechanical movements. It is much like the difference in how humans move. A person may occasionally hit the ball back, but if every shot relies on violently flinging the body, abrupt stops, and constant corrections, that kind of motion is hard to sustain. Truly useful athletic ability is not just succeeding once, but being able to turn success into a stable bodily habit.03
Robots also can't just learn in a "clean world."
At this point, the robot has both movement and a strategy. But there is still a bigger problem: the simulated world and the real world are not the same. In the simulator, researchers know the ball's mass, elasticity, and air resistance, as well as the robot body's mass and friction parameters. Where the ball flies in from and at what speed can all be precisely controlled. In reality, a tennis ball's bounce is never the same twice, a racket's center of gravity may have subtle errors, and the ground's friction also varies. The robot's joints, actuators, and sensors all have errors too. Not to mention that the real system also has latency, noise, and occasionally missing observations. If the robot only learns to play in a "perfect simulated world," the movements it learned may fail immediately once it arrives in reality. LATENT's approach is not to try to make the simulated world as perfect as possible, but the reverse:deliberately stuffing uncertainty into the simulated world.The research team randomized parameters such as the robot body's mass, center of mass, foot friction, joint friction, and actuator inertia, and also randomized dynamic parameters of the tennis ball such as mass, coefficient of restitution, and air resistance. Figure 2(c+d): Dynamics Randomization + Observation Noise. At the same time, they also added factors such as observation noise, frame drops, and latency. In other words, what the robot faces during simulated training is not a fixed tennis-ball world, but an entire ensemble of "possible real worlds." Sometimes the ball is a bit heavier, sometimes lighter; sometimes friction is a bit greater, sometimes smaller; sometimes the sensor's data arrives a bit late, and sometimes a frame is simply missing. What the robot is forced to learn is: "Even if I don't know which situation the world is actually in, I should still try to hit the ball back." This is also a very interesting difference between robot learning and traditional programmed control. A program can tell the machine: ball speed is A, friction is B, latency is C, so execute action D. But the real world rarely actually gives you such tidy conditions as A, B, and C. So LATENT did not attempt to eliminate uncertainty, but let the robot get used to uncertainty during the training stage. The ablation experiments in the paper illustrate this well. If tennis ball dynamics randomization is removed, the real-world forehand success rate drops markedly to 16.67%; if observation noise is removed, the backhand experiment even fell to 0%. These two results precisely point to one issue:when a robot fails in reality, sometimes it is not because it failed to learn the movements, but because the world it learned was too clean.This is one of the core difficulties of sim-to-real. Making the simulator more and more like reality is of course important, but there is another route: letting the robot get used to the fact that reality will never completely resemble the simulator. LATENT chose the latter. This approach was ultimately brought to a real Unitree G1. However, a special clarification is needed here: this was not a robot standing on an ordinary tennis court completing all the experiments with just its own cameras. The current real-robot validation still relies on a motion capture system. The research team used motion capture cameras to obtain motion information about the robot and the ball, and the real experiments used a large-scale motion capture environment. According to official project materials, the experiments used more than 50 motion capture cameras, with a motion capture area of 19 meters × 15 meters. This forms an interesting contrast with the earlier data collection. When collecting human body motions from 5 amateur players, the research team only needed a small 3-meter × 5-meter area; yet actually making the robot stably play in the real world required a whole set of complex motion capture infrastructure. That is to say,"imperfect data" does not mean the whole experimental system is simple.On the contrary, the research team shifted part of the problem from "how to collect perfect data" to "how to let the robot learn stable capability from imperfect data," but the real-robot validation itself remains an expensive, complex engineering system. The paper also explicitly treats the dependence on the motion capture system as a limitation of the current work, and points to further use of active vision as the next direction. This means that today's LATENT is not yet an "autonomous tennis robot" that has escaped laboratory conditions. But it has at least proven one thing:imperfect data can be the starting point for a robot to acquire dynamic motor capabilities.04
It really hit the ball back
What is ultimately most tangible is still the ball. In simulation experiments, the research team conducted large-scale tests evaluating different incoming ball conditions. LATENT's forehand return success rate reached 96.52%, with an average landing error of 1.32 meters. As a comparison, the forehand success rates of methods such as AMP, ASE, and PULSE were 41.32%, 63.47%, and 71.85% respectively, with correspondingly much larger average landing errors. These numbers show that LATENT really is not simply making the robot "perform a swing motion." It is already able to connect body movements, incoming ball conditions, and task goals. With incoming balls at different positions and speeds, the robot needs to move its body, complete the hit at the right position, and then send the ball back to the target area as much as possible. And in the real world, the results still held. The paper conducted 20 consecutive human-robot rally experiments. Under forehand conditions, the paper reports a return success rate of 90.90%, and 77.78% for backhand; front-court and back-court conditions were 88.89% and 81.82% respectively. There is one very important qualification here: these numbers are not the robot's "real tennis match win rate." The success criterion in the real experiments was whether the robot could return the ball inside the opponent's court; and the experiment scale was only 20 consecutive rallies. Therefore, 90.90% is more accurately understood as: under the real experimental conditions set by the paper, the robot can complete forehand returns quite stably. What it proves is a capability, not professional competitive level. In fact, the task LATENT currently solves is still fairly well-defined: the robot needs to hit the incoming ball back to a target position, rather than complete a full singles or doubles match like a real tennis player. Real tennis matches involve many more complex questions. When should you hit down the line, and when cross-court? Where is the opponent standing? How should it recover position after this shot? If the opponent keeps changing the landing spot, how does the robot adjust its tactics? If the ball does not land in the expected area, what should the next shot be? These questions are no longer just "how to swing." They require the robot to understand a continuously changing adversarial environment. The robot-robot self-play in the paper provides a supplement. In simulation, two robots using this policy can play against each other, completing up to 25 consecutive rallies. But a boundary must be drawn here too:the 25 rallies come from simulated robot-versus-robot experiments, not real-world human-robot matches.So, if LATENT is placed in the longer history of robotics, what it has actually accomplished is not "teaching robots to play tennis matches," but pushing forward a previously hard-to-solve problem: can a robot learn a transferable bodily capability from limited fragments of human motion, and then actually use it in a high-speed, dynamic, error-filled environment? The answer is currently yes. At least within the tasks set by the paper, it already can. And the reason this deserves attention is precisely not tennis itself.05
Tennis is just an entry point
For a long time, demonstrations of robot locomotion ability tended to take place in relatively deterministic environments. Walking, grasping, carrying, opening doors—every task had a clear goal, and object positions usually did not change suddenly. But sports are not like that. The ball flies, the target moves, contact time is short, and movements must be performed continuously. The robot not only has to control its own body, but also continuously re-decide its next step based on changes in the external world. Sports therefore actually provide an extremely extreme testbed for embodied intelligence. Soccer requires the robot to run, turn, kick, and also understand other players' positions; badminton requires it to complete footwork, swings, and landing-point control against high-speed incoming shots; tennis concentrates these problems further into a very specific instant: the robot must hit a ball flying at high speed at the correct position, with the correct body posture, at the correct time. This is also why at this year's IROS 2026, athletic humanoid robotics has become a direction discussed in its own right. The IROS 2026 "Perception and Decision Making for Athletic Humanoid Robotics" workshop directly lists agile locomotion, sports, dynamic manipulation, real-time planning, embodied AI, and more as discussion topics, focusing precisely on how robots perform perception, decision-making, and whole-body control in high-speed, contact-rich, and unpredictable environments. The workshop also scheduled demonstrations of athletic robots from Unitree, Phybot, and Booster Robotics, covering tasks such as sports, badminton, and soccer. LATENT's value lies right here. It did not attempt to turn human motion data into an extremely precise answer. The 5 amateur players left behind only some motion clips, and even those clips themselves contain errors. But this imperfect data still tells the robot something: how humans move in this world, how they keep balance, how they swing their bodies to chase a fast-moving object. Then, the robot continues forward on its own. It compresses human motion into a latent action space, then uses reinforcement learning to find action combinations suited to its own body; it does not fully trust human wrist motion, so it learns to let the whole body participate in compensation; it does not assume the real world's parameters are always accurate, so it actively adds noise, latency, and dynamics variations in the simulator; finally, it brings all of this onto a real robot body. It is trying to answer a bigger question:when we cannot give the robot a perfect instruction manual for the world, can it build its own understanding of its body and the world from imperfect information?This is precisely what makes tennis such an interesting testing ground for embodied intelligence: it requires the robot to act in a world with no fixed answers, one that is constantly changed by other agents, continuously exposing problems and continuously optimizing. The ball will not wait for the robot. When it comes, the body must react before language. The sensor has latency, the movement has error, and the original trajectory is no longer an instruction that can simply be followed. So in the end, what really stands on the court is not a machine that copies human motion frame by frame. In its feet are motion fragments left by 5 amateur players, there is imperfect human data, deliberately exposed uncertainty, and bodily strategies formed through reinforcement learning's constant trial and error. What humans provide are traces of knowledge, but what the robot gains is not knowledge itself—the remaining part must be completed by its own body.To make it easier for everyone to gather and chat and exchange conference information, Leifeng Network (official account: Leifeng Network) has specially created an IROS 2026 conference attendee exchange group! Join the group and you'll unlock these "perks"?
Conference Intel Station: submission DDLs, formatting requirements, review progress... key milestones delivered the moment they happen, so you'll never miss a deadline again—your submission prep is rock solid.
Peer Tea Room: peers from all over are in the group, sharing experiences, exchanging ideas, and building connections—on the research road, we walk together.
Frontier Supply Station: the latest research results and hot topics synced in real time, helping you clarify your submission approach in line with the conference themes and fueling your paper inspiration.
On-Site Support Team: (exclusive during the conference):No need to panic once you get to the venue—
Route Navigation: venue layout, how to get to the lecture halls, ask in the group anytime;
Shared Session Reading: discussions of hot sessions and Keynotes—catch up even if you missed them;
Poster Compilation: curated highlights of on-site posters, saved in one click so nothing is missed;
Meetup Plaza: dinner meetups, interview meetups, exchange meetups... want to chat, collaborate, or meet in person? Just start it in the group.
? Group entry portal: Scan the code to join the group or add WeChat Luyoyo_2026, with the note: ECCV + institution/school + name + research direction.
In research or tech, information gaps matter a lot.
Come, let's stay one step ahead together!
Hop on board and let us show you the best of global top AI conferences
Exclusive access to:
Expert presentation slides
Full text of conference reports
Interpretations of popular papers
Interviews with rising academic stars
Scan the QR code above
or click "Read Original" to follow the channel.
This is an original article by Leiphone. Unauthorized reprinting is prohibited. For details, seeReprint Notice。