雷峰网 AIUpdated

After studying 1933 IROS papers, we have seen six new shifts in robotics

Large models have not "eaten" robotics. Author: Ma Xiaoning Editor: Cen Feng On September 27, IROS 2026 will open in Pittsburgh. After an interval of 31 years, IROS returns to this city of steel. As one of the world's…

Large models have not "eaten" robotics.

    Author: Ma Xiaoning

Editor丨Cen Feng

                                                                                                       

On September 27, IROS 2026 will open in Pittsburgh. After an interval of 31 years, IROS returns to the Steel City. As one of the largest robotics academic conferences in the world, this year's main program includes 1933 papers. Spread out, these nearly two thousand papers read more like a cross-section of robotics research in 2026: from reinforcement learning and imitation learning to path planning, visual perception, motion control, and dexterous manipulation, and on to humanoid robots, tactile sensing, and large models — nearly every important technical direction in robotics from recent years appears simultaneously in this map. But what truly deserves attention is not keywords like Humanoid, VLA, or World Model, which more easily attract industry and media attention. After a multi-label categorization of the 1933 papers, we can see that Robot Learning / Embodied AI related papers number about 809, Navigation / Planning 564, Perception / Vision 556, Control / Dynamics 546, Manipulation 520, while Humanoid / Legged related papers number about 213. Since a single paper may span multiple directions, these numbers cannot simply be added together, but they demonstrate one thing: robotics research has not been re-consolidated into a single "AI problem" because of the arrival of large models. On the contrary, learning, perception, planning, control, and manipulation are becoming interwoven with unprecedented density. Against the backdrop of the widely watched AGI frenzy, these nearly two thousand papers send a signal:Large language models have not“Eat up”Robotics.On the contrary, after two years of sprinting on end-to-end approaches, robotics research is undergoing a 'middle-layer revival' of 3D geometry, symbolic reasoning, and low-level control, trying to recover in Pittsburgh the physical soul obscured by pixels and tokens." The real question worth asking about IROS 2026 may not be 'whether large models are replacing traditional robotics,' but whether robotics research is entering a more complex systems phase.

01


  AI has not replaced traditional robotics,

it is being re-embedded into it

If we look further at the cross-relationships among these papers, what actually changes when large models enter robotics becomes clearer. About 276 papers feature Robot Learning and Manipulation together, 269 intersect with Perception, 225 with Navigation / Planning, and 221 with Control / Dynamics. Learning methods have not simply squeezed out robotics' original technical modules; instead, they are increasingly embedded in these traditional problems. This change is first reflected in planning. Jorge Mendez-Mendez from Stony Brook University, in "A Systematic Study of Large Language Models for Task and Motion Planning with PDDLStream", the study tested the performance of large language models embedded into a traditional Task and Motion Planning framework. The research does not point to "large models directly replacing the planner"; instead, it shows that geometric constraints, motion feasibility, and execution logic still require the participation of traditional planning systems. This is quite representative. An LLM can understand what "take the cup to the other side of the table" means, but when the robot actually executes it, the system must still answer whether the robotic arm will collide, whether the joints are reachable, whether the path can be executed, and whether the current environment has changed. The question therefore shifts from "whether the LLM can replace the Planner" to how the two should divide their roles. A similar phenomenon also appears on the perception side of VLA. A finalist for the Best Paper Award in Cognitive Robotics at IROS 2026, GeoVLA: Empowering 3D Representations Within Vision-Language-Action Modelsa finalist for the IROS 2026 Cognitive Robotics Best Paper award, was completed by a team from Tianjin University, Yuanli Lingji, and Tsinghua University, with authors including Yuanli Lingji'sXie Bin, Tsinghua University'sShi Hao, and others. It targets precisely the problem that existing VLAs rely excessively on two-dimensional vision: traditional VLAs often generate actions directly from RGB images and language instructions, but real robots face a three-dimensional world, where object height, size, and camera viewpoint changes can all directly affect grasping. GeoVLA did not continue simply scaling up the model, but explicitly added depth and 3D geometric information, using 3D representation together with vision-language features for action generation. In other words, once large models became strong, robotics' most traditional "spatial geometry" was filled back in. A similar phenomenon has also appeared on the control side. Training Fast Robot Policies Using Slow Foundation Models, from a team at the University of Central Florida, the Indian Institute of Technology Kanpur, the University of Bath, and the University of Maryland, College Park, studies a very practical problem: large Foundation Models can provide stronger knowledge and generalization, but their inference speed is too slow to run continuously in the robot's high-frequency control loop. Therefore, the researchers let the large model participate more in training, then learn a lighter, faster robot policy responsible for actual execution. When a language model answers a question a few hundred milliseconds late, users usually just find it slightly slow; when a robot acts a few hundred milliseconds late while grasping, making contact, or maintaining balance, the task may already have failed. Model capability and real-time control have never been the same problem. Tactile research, which is closer to the physical world, shows a similar direction.DexTact: A Visuo-Tactile Foundation Model for Dexterous Hand Grasping, completed by Fang Xinminand others from the University of Colorado Denver in collaboration with a team from Kennesaw State University, puts vision and touch together in the dexterous hand grasping problem. What it studies is no longer just "images tell the robot where the object is," but "after the hand contacts the object, what else can the system perceive." Here, Foundation Models, touch, dexterous manipulation, and control are placed into the same system to be understood. Therefore, the change truly worth noting in 2026 is not that robotics is being replaced by AI. More precisely, AI is beginning to go deep into robotics' most traditional parts, while also being reshaped by these traditional problems.

02


VLA moves from "proving it can do it" to "patching weaknesses"

If the core question for VLA over the past two years was still 'can a single unified model understand vision and language and directly generate robot actions', then by IROS 2026 the research focus has clearly shifted in another direction: now that models can do it, what needs to be solved is how to make them faster, more stable, and better suited to real robots. In our multi-label statistics, there are roughly 162 papers related to VLA, VLM, LLM, and Foundation Model, accounting for 8.4% of all papers; among them, about 82 papers explicitly involve VLA. This scale is already enough to form an independent research mainline. But the questions they focus on have not continued to dwell on 'can models get even bigger'; instead they have rapidly diverged into efficiency, 3D understanding, long-term memory, viewpoint changes, real-world adaptation, and safety. The most direct example is Shallow-π: knowledge distillation for flow-based VLAs. This paper comes from 42dot, Seoul National University, and Samsung Research, and has been shortlisted for the IROS 2026 Best Paper / Best Student Paper Award. What it solves is not how to train a bigger VLA, but precisely the problem that models are too large and inference too slow; by compressing VLAs through knowledge distillation, it makes the models more suitable for deployment on real robots and edge devices. This exposes the first real constraint VLAs face after entering robotics: models must not only be 'smart', but also fast enough. Another shortcoming is 3D spatial understanding. The aforementioned GeoVLA from the Tianjin University, Yuanli Lingji, and Tsinghua team is essentially supplementing VLAs with robotics' most traditional geometric capabilities. Viewpoint change has itself even become an independent problem.AnyCamVLA: zero-shot camera adaptation for viewpoint-robust vision-language-action models comes from a team at Seoul National University and MIT. It studies why a well-trained VLA tends to fail after the camera position changes, and whether it is possible to adapt to a new camera viewpoint without retraining the entire model. In real deployments, camera installation errors, robot platform differences, and viewpoint changes are all common; a policy that only works under a fixed viewpoint can hardly truly generalize. Long-horizon tasks expose another kind of gap: models can execute actions but do not necessarily accumulate experience.Long-term memory for agents based on VLAs in open-world task execution was completed by teams from Nanjing University, Waseda University, LimX Dynamics, and Zhejiang University, among others, connecting long-term memory directly to a VLA Agent so the system can save trajectories and experiences it has already completed and recall them later. The question here is no longer whether a single grasp succeeds, but whether the robot can remember 'what happened before' in long-horizon multi-step tasks. From Sungkyunkwan University, RoboBRIDGE: a modular framework for bridging policies to robust real-world robotic agents goes a step further. It does not attempt to retrain a stronger VLA, but instead adds a Monitor, Perceptor, Planner, Controller, and Robot Interface around the pre-trained policy. When execution fails, the Monitor identifies the problem; when the environment changes, the Planner re-plans; the Perceptor updates the scene; and the Controller turns the high-level policy into actual actions. This structure illustrates very well what is happening in the next stage of VLA: research is shifting from 'building a stronger action predictor' toward how to build a complete system around it. So the coming competition for VLA may not just be about scaling. The bigger question is becoming system engineering: how a VLA works together with perception, memory, planning, control, and hardware, and ultimately transforms from a model capable of outputting actions into a robotic system capable of running over the long term.

03


Robots are regrowing a 'middle layer'

If the first stage of VLA was an attempt to unify perception, language, and action into a single model, then another increasingly clear research thread at IROS 2026 is: researchers have begun actively adding new structures in the middle of that end-to-end chain again. Among the 1,933 papers, there are roughly 119 papers related to Reasoning / Memory, accounting for 6.2% of all papers; among them, 64 explicitly involve Reasoning and 36 involve Memory. They have already formed a batch of very clear Sessions, including Agents and Memory for Robust VLA Manipulation, Constraint-Guided Reasoning for Long-Horizon Manipulation, LLM Agents Reasoning Through Robot Tasks, and Representing 3D Space for Robot Reasoning. This fact alone is worth probing: if end-to-end models can already generate actions directly from vision and language, why do robots still need Reasoning, Memory, Planner, or even reintroduced Symbolic Constraints and explicit 3D representations? One reason is that once robot tasks become longer, 'what is seen in the current frame' is no longer enough.TempoFit: a plug-and-play layer-wise temporal KV memory for long-horizon VLA manipulation comes from a team at Xi'an Jiaotong-Liverpool University. It does not retrain a bigger VLA, but instead uses the model's internal temporal KV memory, allowing existing policies to save and retrieve historical information across time. But what robots really need to remember is not just 'what was seen in the past'. From a team at Tokyo University of Science, GaussMemory: task-driven 3D Gaussian scene memory for long-horizon robotic manipulationfurther places memory in 3D space. Even if an object is temporarily occluded, it must not disappear from the robot's 'world' just because the camera cannot see it at the moment; an object that has already been moved also needs its position updated in time. Therefore, memory in robots is not entirely the same as contextual memory in language models. It must continuously correspond to real 3D space. The same goes for reasoning.DualCoT-VLA: a visual-linguistic chain of thought via parallel reasoning for vision-language-action models was completed by a collaboration between HKUST (Guangzhou) and Huawei. It is not satisfied with having a VLA directly output actions; before action generation, it adds a visual Chain-of-Thought and a linguistic Chain-of-Thought: one handles spatial relations, the other handles high-level task logic. The 'high-level–low-level' division of labor that end-to-end models originally tried to eliminate has come back in another form. VL-Nav: a neuro-symbolic approach for reasoning-based vision-language navigation, a finalist for the Cognitive Robotics best paper award, was mainly completed by a team from the University at Buffalo, State University of New York, with first authorDu Yi. VL-Nav combines the semantic understanding of vision-language models with explicit spatial and symbolic reasoning for vision-language navigation in complex environments. The reappearance of symbolic methods does not mean robotics has returned to a fully rule-based era. What it undertakes is the part of the task that large models are still not good at completing stably: making spatial relations, execution conditions, and constraints more explicit. From Memory and Reasoning to Geometry, Planner, and Symbolic Constraint, these works use seemingly different techniques, but all solve the same structural problem: there is a huge distance between high-level semantic understanding and low-level continuous actions. Therefore, the Reasoning, Memory, Planner, and Symbolic Constraint appearing at IROS 2026 are less about being 'anti end-to-end' and more about robotics reinventing themiddle layerbetween large models and low-level control. End-to-end has not simply failed, and modularization has not simply been restored. The more accurate description of the change is: as tasks move from single-step grasping toward long-horizon manipulation, and from laboratory environments toward the open world, hierarchies are reappearing in robotic systems.

04


Manipulation 

has become one of the most densely contested battlefields in embodied intelligence

If VLA, Reasoning, and World Models represent more how robots 'think', then Manipulation is closer to another question: how these capabilities ultimately act on the physical world through a hand, a gripper, or a robotic arm. There are roughly 520 papers related to Manipulation; after further filtering in directions such as Dexterous Manipulation, Tactile, In-hand Manipulation, and Contact-rich Manipulation, about 225 papers remain relevant. Among them, about 91 papers explicitly involve Tactile, and 52 explicitly involve Dexterous Manipulation. Among these 225 Dexterous/Tactile-related papers, about 165 also involve Manipulation, 79 involve Learning, 57 involve Perception, and 52 involve Control. Tactile sensing, learning, perception, and control are continuously converging within the same batch of manipulation tasks. IROS 2026 Award Candidate VTAP Gripper: synergizing fingertip sensing and a visuo-tactile active palm for dexterous in-hand manipulation, completed by a collaboration between teams from Purdue University and Columbia University, includes Purdue University'sHu Zhixianand Columbia University'sHuang Binghao, Li Yanzhuand other researchers. Instead of continuing to pursue an anthropomorphic dexterous hand whose degrees of freedom get closer and closer to the human hand, it designed three reconfigurable soft fingers and an active visual-tactile palm surface, completing grasping and in-hand manipulation through fingertip touch and active contact with the palm. The most interesting point here is that high-level dexterous manipulation does not necessarily depend on a highly anthropomorphic structure with many degrees of freedom. The synergy among mechanical structure, active contact, and multimodal sensing can likewise produce complex manipulation capabilities. Another Robotics Mechanism Design Award candidate PDS Joint: A Parametric Double-Spiral Joint Designed Specifically for Dexterous Hands, from a five-person team at the Beijing Institute of Technology. They started from a lower-level joint structure, jointly incorporating compliant structures, joint stiffness, proprioception, and learning-based calibration into dexterous hand joint design. From VTAP to the PDS Joint, we can see that the hardware itself is still far from converged, but hardware is also becoming increasingly difficult to separate from sensing, control, and learning. At the same time, touch is moving from "installing a sensor" toward "entering the policy."TacVLA: Contact-Aware Tactile Fusion for Robust Manipulation with Vision-Language-Action Models was completed through collaboration between Purdue University and the Italian Institute of Technology, with the Purdue team includingXu Zhengtong, Zhang Zhiyuanand other researchers. It feeds touch directly into the VLA so that, when contact actually occurs, tactile information participates in action decisions. This is because in scenarios involving occlusion, fine insertion and unplugging, or contact that has already been established, the information provided by vision drops rapidly. Another route attempts to reduce the dependence on tactile hardware during the deployment stage.HapticVLA: Contact-Rich Manipulation Through a Vision-Language-Action Model Without Tactile Sensing at Inference Time comes from a team at the Skolkovo Institute of Science and Technology. They use tactile information during training, letting the policy learn contact knowledge, but at inference time no longer strictly depend on real-time tactile input. TacVLA and HapticVLA happen to represent two different lines of thinking: one connects touch into the policy in real time, while the other uses touch during training and then tries to compress the tactile experience into the model. Data issues are also beginning to emerge. Among the finalists for the Entertainment and Amusement Paper Award, PianoFingering-1.5K: A Large-Scale Dataset of Expert Piano Fingering for Dexterous Robot Learning comes from a team at the University of Science and Technology of China, with authors includingWang Ruoyuand others. They organized expert piano fingering into a structured dataset. Piano performance may seem specialized, but it is in fact an extremely demanding dexterous manipulation task: multiple fingers must complete precise, continuous contacts under strong temporal constraints within very short time intervals. Systematizing such expert-level movements is essentially providing high-quality human priors for robotic dexterous manipulation. From hardware and touch to data, several routes ultimately converge on the same question: how robots can form truly closed-loop manipulation capabilities. Dexterous manipulation is therefore expanding from "building a better hand" toward establishing a complete closed loop ofperception—touch—data—policy—control.

05


Humanoid is expanding from "being able to walk" to "working while walking"

Humanoid robots may be one of the robot forms that has attracted the most industry and media attention over the past two years, but placing them back into the full paper landscape of IROS 2026 paints a more restrained picture. According to the statistics, there are approximately 89 papers related to Humanoid / Loco-Manipulation, accounting for 4.6% of all papers. Among them, 46 papers intersect with Control, 44 with Learning, 34 involve Manipulation, and 32 involve Planning. What is truly noteworthy is that the problems within these papers are changing. Over the past few years, the most important technical advances in humanoid robots were largely concentrated in locomotion; now we can see another batch of work beginning to put "mobility" and "manipulation" into the same system. A typical example is ULTRA: A Unified Multimodal Control Framework for Autonomous Whole-Body Loco-Manipulation in Humanoid Robots. This paper was completed by Xiaolin He of the University of Illinois Urbana-ChampaignXiaolin Heand others, together with researchers from Shanghai Jiao Tong University and Tsinghua University, and was selected as a candidate for the IROS 2026 Best Paper related to Mobile Manipulation. It is not about training yet another stronger walking policy, but about enabling humanoid robots to accomplish continuous tasks such as approaching objects, picking them up, carrying them and putting them down. What is most noteworthy here is not that the robot finally "can carry things," but that the structure of the problem has changed. In pure locomotion, the core objective is maintaining body stability; once the robot holds an object in its hands, arm movements shift the center of gravity, the load affects the gait, and foot contacts, upper-limb motion and the grasping task must all be satisfied simultaneously. "Walking" and "manipulation" can no longer be completely separated.SteadyTray: Learning Object Balancing Tasks for Humanoid Tray-Carrying Through Residual Reinforcement Learning comes from a team at the University of California San Diego, with first authorAnlun Huang. It magnifies this contradiction into a very intuitive task: having a humanoid robot walk while carrying a tray, with the objects on the tray not allowed to fall. SteadyTray adds a residual policy on top of an existing locomotion policy, specifically handling the disturbances that the gait introduces to the end effector.DreamMimic: Learning Visuomotor Whole-Body Loco-Manipulation by Means of a World Model was completed by Jie Yin of Shanghai Jiao Tong UniversityJie Yinin collaboration with a Tsinghua University team, bringing vision and World Models into whole-body loco-manipulation, using predictive models to help the robot maintain an understanding of state during long-horizon visual control. VLAs are also beginning to enter the same problem.Learning Humanoid Loco-Manipulation Using Responsibility-Induced Specialized Experts Within Vision-Language-Action Models comes from a Seoul National University team. Addressing the problem that the action distributions of locomotion and manipulation differ greatly, this work introduces specialized experts, letting different experts take on different control responsibilities. This again shows that "general" does not mean "nothing is differentiated internally." An even more extreme example is the Award Candidate Learning Athletic Tennis Skills for Humanoid Robots from Imperfect Human Motion Data. This work was led by Zhikai Zhang of Tsinghua UniversityZhikai Zhangand others, completed jointly with researchers from Peking University, the Shanghai AI Laboratory, Galbot and Delft University of Technology. Playing tennis requires the robot to observe a fast-moving ball, predict its trajectory, move its body, swing the racket, all while maintaining balance. It may look far removed from factory material handling, but it exposes the same problem: a humanoid robot cannot merely have one stable gait and then simply "add two hands." True whole-body intelligence requires the legs, torso and arms to move together around the task. Therefore, summarizing the IROS 2026 humanoid robot papers simply as "Humanoids are getting hotter" actually misses the real change. A more accurate description is:the research focus is expanding from pure locomotion toward whole-body loco-manipulation.

06


World Models are hot, but have not yet become mainstream

If you only look at the discussion buzz in the robotics field over the past year, World Models easily give the impression of having "fully entered the main battleground." But among all the papers, only 19 explicitly involve World Models, about 1% of all papers. "Hot" does not mean "many." What is truly worth noting is not how many World Model papers there already are, but where they are starting to appear: quadruped locomotion, dexterous manipulation, precision insertion, humanoid whole-body control, medical robots and navigation. Best Paper / Best Student Paper Award Candidate DAWN: Noise-Robust Parkour for Quadruped Robots by Means of Depth-Denoising World Models comes from a team at Korea University of Technology and Education. It does not have the quadruped robot "imagine" more realistic videos; instead, it uses a World Model to model noisy depth perception, helping the robot complete parkour in real environments. Here the World Model has begun to become a state-modeling tool in the perception-control chain. In dexterous manipulation,Scaling Up Cross-Embodiment World Models for Dexterous Manipulation was led byZihao Heand completed jointly with researchers from the University of California San Diego, MIT and other institutions. It attempts to bypass the completely different joint and action spaces between different robot hands, learning more general laws of "how actions change the physical world." Behind this lies a very interesting hypothesis: what different robots can truly share in the future may not be the actions themselves, buthow actions change the world。Achieving Generalizable Robotic Insertion Tasks With World Models was completed jointly by teams from UC San Diego, NVIDIA, the University of Washington and the University of Southern California. It puts World Models into precision insertion tasks. Rather than having the policy memorize "how to insert this kind of part," the researchers want the model to learn "how my actions will change the current state," then use the prediction to complete control. As for humanoid robots, the aforementioned Shanghai Jiao Tong University'sJie Yinand, together with the Tsinghua team, DreamMimic has again brought World Models into whole-body loco-manipulation. While RoDyn: Taming a 2.5D World Model with Interactive Robot Dynamics for Robotic Manipulation was completed by a joint team from Nanyang Technological University and Tsinghua University. RoDyn attempts to solve an obvious problem with purely 2D video prediction: a future frame looking plausible does not mean the physical relationships are correct. It further introduces depth and robot dynamics information, making the prediction results closer to a state representation truly usable for robot control. This precisely points to the most critical contradiction of World Models in robotics. What robots truly need is not a visually plausible future, buta future under given actions that is sufficiently credible geometrically and physically. So this year, the question for World Models is no longer just: "If the robot does this, what will the future look like?" It has started to become: "Since I can predict the future, can this prediction truly change what I do next?" In this sense, World Models currently look more like a technical route in transition frompredicting the worldtoparticipating in control, rather than an approach that already dominates robotics research.

07


Robotics enters the "era of systems problems"

Looking at these studies together, the most noteworthy thing about IROS 2026 is not that some technical route is "winning." VLA has not replaced traditional planning and control, World Models have not yet become mainstream, and Humanoids are only one part of the overall robotics research landscape. The real change taking place is that once these new methods enter the physical world, problems such as perception, memory, planning, control, touch, safety, and recovery become unavoidable again. The success of language models over the past few years has largely come from unifying a large number of tasks into a single model. But robots are different. They must not only "understand" but also truly act, and the physical world exposes errors, latency, contact, and failure in full. So IROS 2026 is more like answering a new question: once models are strong enough, how do you reassemble these capabilities into a robotic system that truly works. This does not mean end-to-end has failed, nor that traditional modules have regained dominance. More precisely, robotics is re-searching for the boundaries between large models, planning, control, and hardware. Future competition may no longer be about whose model is bigger or more capable, but about who can truly combine these capabilities in a stable way.Once robots get smarter and smarter, what truly becomes difficult is: how to keep them from making mistakes.

To make it convenient for everyone to gather, chat, and exchange conference information, we have specially created an IROS 2026 2026 attendee discussion group! Joining the group unlocks these "perks"?

  • Conference intel station: submission DDLs, formatting requirements, review progress... key milestones delivered first-hand, so you never miss a deadline and your submission prep is rock solid.

  • Peer tea chat: colleagues from all over are in the group, sharing insights, exchanging ideas, and building connections—we walk the research road together.

  • Cutting-edge supply station: the latest research results and hot topics synced in real time, helping you clarify your submission direction in line with the conference themes and fully fueling your paper inspiration.

  • On-site support team: (exclusive during the conference):No need to panic when you arrive at the venue—

    Route navigation: venue layout, how to get to lecture halls—ask in the group anytime;

    Session co-reading: discussions of popular sessions and Keynotes—catch up even if you missed them;

    Poster compilation: highlights of on-site posters organized for one-click collection so nothing is missed;

    Meetup plaza: dinner gatherings, interview meetups, exchange meetups... want to chat, collaborate, or meet face to face—just start it in the group.

? Portal to join the group: Scan the code to join, or add WeChat Luyoyo_2026 with note: ECCV + institution/school + name + research direction.

In research and technology, information gaps matter.

Come, stay one step ahead together!

Hop aboard for a tour of the highlights of top global AI conferences

Exclusive access to:

Expert presentation slides

Full conference reports

Interpretations of popular papers

Interviews with rising academic stars

Scan the QR code above

or click "Read Original" to follow the Leiphone zone (WeChat account: 雷峰网).

This is an original Leiphone article; unauthorized reproduction is prohibited. For details, seeReprint Notice。

Original source

雷峰网 AI

Content notes

Original publication and rights belong to the source.

Machine translation · Refer to the original