雷峰网 AI

Wu Jiajun IROS 2026 Speech: Structured Representation is Not the Answer, But a Scaffold for Robots to Learn Reasoning | IROS 2026

From structured representation to demonstration-free learning. Author: Wu Simeng, Editor: Cen Feng. A plate, a faucet, and a small lump of blue detergent. Wu Jiajun started his speech at IROS 2026 with such a kitchen…

From structured representation to learning without demonstrations.

Author | Wu Simeng

|| Cen Feng

                                                                                                       

A plate, a faucet, and a small lump of blue detergent. Wu Jiajun began his speech at IROS 2026 with such a kitchen scene that requires almost no technical expertise. A student squeezed the detergent, turned on the faucet, washed half of it, and left temporarily. Now, a robot takes over. At first glance, the task is straightforward: clean the plate and put it back. But what the robot actually needs to handle is not the phrase “cleaning plates”. It needs to first know which object in front of it is a plate; it needs to determine how dirty this plate is; it needs to notice whether there is a small lump of blue detergent on the plate. It also needs to know if the faucet is currently open, if the plate has been washed, and whether a new problem appears on the table surface after it picks up the plate. The same task, with a different initial state, may have a completely different answer. If the plate has already been washed, there’s no need to wash it again. Without detergent, it’s necessary to squeeze some onto it first. If there is dirt pressed under the plate, “washing the plate clean” doesn’t mean putting it back on the rack. When people face such changes, they often don’t feel like they have done any complex reasoning. We just take a look, and then continue with our task. But for a robot, every “look and continue” involves a question: what exactly does it understand about the world in front of it? On September 27, on the opening day of IROS 2026, Wu Jijun, an assistant professor in the Department of Computer Science at Stanford University, delivered an invited presentation titled “Building Physical Agents via Structured Representations” at the RoBoWoMo seminar. Over the past few years, he has been researching how to recover the world hidden behind the visual image from images: geometry, materials, physical properties, and the relationships between objects. Now, when these issues are placed in the physical world of a robot’s real life, other problems arise as well. How should a robot be described for a task? Can these descriptions be learned on their own? When dealing with objects that have never appeared in the training data, when does the model need post-training training? If planning itself can also be learned, is a traditional planner still necessary? Finally, the question even becomes not “has the robot seen how humans do it?” Without demonstrations, without remote operation data, and even when faced with a faucet it has never seen before, can the robot still figure out on its own how it might be used? Here is a condensed version of Wu Jijun’s speech presented at IROS 2026. Lei Feng Network (public account: Lei Feng Network) · AI Technology Review has edited it based on the original English speech, reorganizing the spoken expressions and technical terms without altering the original meaning.

▎Building Physical Agents via Structured Representations

Speaker: Wu Jiajun, Assistant Professor in Computer Science at Stanford University (also Assistant Professor in Psychology)

01


Why a plate causes trouble for robots

Today, I want to talk about how we can use structured representations to build agents that can act in the physical world. Let’s start with vision. In the end, you only see images, but behind each image lies a whole world: geometry, materials, and various physical properties. Our research focus has always been on how much of these elements can be restored; and once they are restored, you can recreate that world. In recent years, there has also been a very active area: agents in digital spaces. They receive instructions and interact with the entire world. This naturally leads to another question: Is there a similar structure in the physical world? In the digital world, you have directories, which contain files, and files have their properties; in the physical world, you have scenes, which contain objects, and objects also have their properties. So, can we use this structure to bring those agents from the digital space into the physical space? This is the topic of today’s presentation. Once you really want robots and humans to exist in the same world, things become difficult. This world is complex; it’s dynamic and diverse, and you need to handle a large number of different tasks. Of course, you need basic models, and they are becoming stronger by the day—so strong that it’s hard to imagine not using them. But the way you use them can be very different: whether to perform post-training, whether to introduce more structure, whether to incorporate some additional knowledge, and how to combine that knowledge with the existing parts afterward. These are all pending choices. There is also a aspect that is particularly related to robots and differs greatly from the digital world. The digital world is, to a large extent, unified by language; on the robot side, it is a different scenario. Different fields, different tasks – often proprietary scenarios and proprietary data – each robot is unique. These data do not appear in the massive corpus used for training large models. So the question becomes: how to cross that barrier. After talking about so much abstraction, let’s start with a completely non-specific scenario—a very ordinary kitchen scene—and see what is needed for a robot system to actually work alongside humans in such a situation. A student pours dish soap onto a plate, turns on the faucet, and then has to leave. Now let the robot take over. Of course, it first needs to perform a semantic reasoning step: among this pile of ceramic tableware, which one is the one I need to help wash? It’s this plate. This step can be done relatively easily today using a visual model. Then it needs to be determined that this plate is for being washed. The detergent has been squeezed on, and the faucet is open. So what I need to do is pick up the plate, wash it, put it on the rack, and turn off the faucet. It seems very clear. But the real challenge is: how to stay consistent between different environments, different objects, and different goals. Even in this already simple scenario, several completely different situations can arise. The first one is that you need to adapt to others' actions. For example, if someone comes in and washes the plates and puts them back on the table, you should know that the plates have already been washed, so there’s no need to wash them again. All you need to do is put them back, and there’s no need to turn off the tap, as it has already been turned off. The second type involves a very subtle difference in the initial state. Assume everything remains the same as before, with the only difference being that there is no dish soap on this plate. If you look at the entire image, this difference is very small—it’s just a tiny bit of blue material that’s hard to see. But it isn’t something that can be carried over with just a “minor difference”: it is a qualitative difference in state representation. Because if there is no detergent applied on the plate, you probably shouldn’t wash it at all; washing directly will create a mess. You need to apply detergent first, then wash, then place it on the shelf, and finally do the rest of the tasks. The difference in input is minimal, while the difference in state is qualitative, and the actions required are completely different. The third type is that you need to follow the constantly evolving observations. For example, when the robot picks up a plate and places it on the rack, that’s enough because it is already clean, and there’s no need to turn off the faucet. But there are some dirty things pressing under the plate. At this point, you should realize that if my goal is really to keep everything clean, then I should wipe the table. No one has instructed you to do this; you figured it out by following the new observations. Ideally, you should also put that plate back; but we can’t even do that step because the robot’s body and the skills it possesses limit it. So you need to be adaptable to all these different situations. The question is, how can you do that?

02


If robots also need a language

In the past few years, we have been trying to develop a set of paradigms. Their specific implementation has always changed, as the basic model and various possibilities also evolve. But the framework has remained the same: one end is the raw input—video or language; in the middle, the basic model helps discover a set of abstract concepts that can be reused and recombined to handle general tasks; the other end is control actions, which should ideally be discovered and learned by ourselves. When it comes to learning actions, we hope to learn from what is called natural supervision. That is to say, the entire process is learned and discovered; the less people are involved, the better. What you provide to it might just be a video, a demonstration, or some language, and now these inputs can already be processed. The less supervision, the better the versatility. But when it’s drawn out like this, a series of questions appear on the paper. What does this combinatorial abstraction look like? How are they discovered? Where does that specialized language for describing tasks come from? What if you don’t need it at all? How should planning be solved? Use a traditional task planner, or learn a planning system? Also, regarding this world model itself, should it be used as is, or require fine-tuning or post-training? There are different implementations for each of these approaches, and we have also tried various versions along the way. Let me first explain the general structure at a high level. The earliest implementation was done two years ago, and the specific components were much more crude than now, but the paradigm is clear. Suppose you have a set of demonstrations. It could be a robot that does many different tasks over the course of time. You can identify the key points throughout the entire process. For example, in the first step, it grabs the bowl; in the second step, it runs under the faucet; and in the third step, it places it on the shelf. It could also be something else—video data, for instance. In any case, you have a set of such data in hand. Your goal is to learn the set of descriptions behind this task. It’s somewhat like a symbolic world model, which consists of roughly two groups of modules. The first group includes the attributes and spatial relationships that you care about. Is this object dirty or clean? What is the relationship between these two objects—two-dimensional or three-dimensional? Such things need to be determined; this is a set of predicates. Of course, determining them cannot rely on human-written rules; instead, we use networks and basic models: allowing them to learn how to determine any set of predicates, whether it’s about dirtiness or certain spatial relationships. The second group are actions. You have a set of atomic actions, such as “wash”. Each action must clearly state its preconditions and postconditions. For this action of washing, you first need to hold the object in your hand. The person must be in the kitchen at this time. The effect is that if there is dish soap on the object, it becomes clean. As for how to know whether the robot is holding an object in its hand, or how to determine whether an object is clean or dirty, these are determined by the previous set of predicates. Additionally, to actually perform an action, you need a strategy, that is, how the joints should exert force. This part is also a neural network, with a diffusion strategy for each atomic action. In this way, the symbol layer defines combinability: as long as the preconditions are met, these atomic actions can be linked together to form longer and more complex tasks. Both perception and control are performed by neural networks. But in a certain sense, this is a specialized language, a domain-specific language (DSL). Thus, a very fundamental question arises: at the beginning, these things were either manually written or directly copied from sources like PDDL, and the DSL was already fixed. So, is it general enough? Why is it sufficient? What if there are things missing? So the next step naturally is: the DSL itself also needs to be learnable. You can have the basic model examine the data you have, whether it’s a demonstration video, operation records, or language, and then tell it that this is the task we want to focus on. Next, it should output: I think these attributes, these atomic actions, as well as their preconditions and postconditions, are important. But what if it makes a mistake? You can also verify it. On one hand, it can be verified symbolically; on the other hand, data can be used directly for verification. For example, the system says that before washing something, you must hold it in your hands first, otherwise it cannot be washed properly. Then you should go back and check your own data: every time I wash something, did I really hold it in my hands first? You can use these demonstrations or other data at hand—public or proprietary—to verify this specification. So this language itself was also proposed by another system. Last year, I talked about this section in a rather messy way. This year, I hope it will be better, because I’ve found that a good analogy can make things much clearer. The analogy is right on your own computer. If you paste the sentence “do this” into a programming agent, it will do ten or eight steps by itself. The first step is to create a file, the second step is to perform some operation, and so on. It has many different steps. Think about it carefully—it’s very similar to what we need to do. What we do is essentially moving a smart entity like Cursor that works on a computer into physical space: first, come up with a series of plans, then execute those actions, and there must also be a layer of verification to ensure that these actions are defined correctly and executed properly. Basically, that’s it.

03


What to do if you encounter something you haven’t seen before, what about the model

Next, let’s talk about when you should sit down and do post-training work. Of course, if you simply use the front-end model to push it all the way forward, as in the analogy just now, that is the easiest way. But there is a problem that cannot be avoided with robots: the things you have to deal with are likely something you have never seen before. Because the physical world is much more fragmented than the digital world. You may have seen objects of a certain shape, but more often, you have a part used on a production line that only exists in your factory and is not available elsewhere. I have never seen this part before. Now I need to determine whether it is dirty or not, and whether it produces any effect. Obviously, I cannot directly use the publicly available cutting-edge model at this time, because it doesn’t even know what defines “this part being good or bad”. So there will indeed be some situations where you have to perform some post-training work. The prerequisite is that you decide not to rely solely on those existing cutting-edge models. But how do you do that? Note something: although these effects seem logical, such as having dish soap on it and becoming clean after washing, in reality, they are written approximately. Logical operations mostly exist in the form of probability. So you can connect the whole chain: I say there is dish soap on this thing, I wash it, and it should become clean. But what if my criterion for “whether there is dish soap” is not good? Since this is a new object that I have never seen before, it cannot be judged, and the output is wrong; everything that follows is also wrong. At this point, my demonstration tells me another fact: this object is indeed clean. Thus, you get a label. Since all the execution steps are stacked incrementally, you can use this execution label from the demonstration for backpropagation, providing supervision for the model. So, post-training can be carried out in some cases. It is not necessary, but if you treat these models as modules within a system, and they need to be retrained later, this path is feasible. I won’t go into detail about the results from two years ago; that set of components is now outdated, and after changing the method, some aspects of it seemed quite clumsy at the time. However, the paradigm itself has great potential. Later, we replaced different components within it with more modern versions. By the way, one of the postdocs who worked on that project is now a professor at National University of Singapore. The other collaborator was a doctoral student at MIT, and he has also become a professor today.

04


Even planning starts being handed over to the model

Next, I will talk about a few more recent developments. Let’s go quickly. As mentioned before, the modules in that system are developed layer by layer, and verification, perception, and control have been handed over to the neural network. But there is one thing still missing: planning. When all the modules are available and the state has been verified, should I wash first, or do something else first? At that time, this step still used traditional methods, following the approach used in traditional task planning. Recently, we just published a paper that replaces this aspect as well. This is also why I don’t have very good slides today. Actually, I should have a video, and I really want a video. So, optimistically, we will have a beautiful video in a few days. For now, we can only use a picture from the paper as a substitute. The idea is as follows: you still have a set of states, which are verified by a neural network; but the decision of “given the current state, should I wash first, or pick up the plate first, or what else?” is no longer left to a traditional planner, but handed over to the network. To be more specific: that set of predicates is basically natural language now—true natural language. You feed both natural language and images as prompts to the base model so it can make plans. But it still needs post-training training. So we have to use an open-source weight model because it needs to be trained later; at the same time, since images are part of the input, the model also needs to be able to process images. And it turned out that some open-source models have very poor visual reasoning capabilities. At least, our tests showed that they are far inferior to contemporaneous closed-source models. So if you use them directly, they often output things that make no sense at all. In practice, we simply selected some models and performed post-training training instead. As a result, even the planner itself becomes something learned through post-training. It can use this kind of symbolic natural language to plan, and the generated action sequences have much higher execution probability. For such long-term tasks, this means much less manual intervention is required. There is also a challenge: all of this depends on a skill set, and each robot’s skill set may be different. If you directly use any cutting-edge model, it has no context from your skill set. For example, it might say, “Now you need to fill this mug with water using the faucet,” but the faucet hasn’t been turned on at all. This action cannot be executed. After the training, the planner produces actions that seem more feasible when dealing with such long-term operations. In a few days, we will have videos, but for now, let’s take a look at the latest results. So far, we have been assuming there is a demonstration available: either a video demonstration or human operation. What I want to talk about in the remaining time is: what if there is no demonstration? What if I encounter a completely new object that I have never seen before? The previous approach was: to demonstrate it to me and tell me how to deal with this object. But is it possible for me to come up with all the possible ways to interact with this object without even a demonstration video or any help from someone? For example, if there is a batch of new objects in front of you. We’ve already demonstrated how to turn on the faucet before. But I have never used this faucet before. Could you provide some examples and tell me what methods can be used to interact with it? This is also something we are exploring. Let’s formalize it roughly. Suppose we only have one stationary object—a faucet or a box. Can we develop a system that derives a distribution from these single stationary objects, and then sample a possible space from this distribution? If you give it a faucet, it samples forms that can rotate; if you give it a box, it samples opening and closing, as well as all other different forms. I’ll briefly mention the specific method today, to respect everyone’s time. But the result is: once you master this system, you can enter a completely new scenario. On a bunch of different objects, a model can infer possible interactions. Displays, USB drives, flowers—anything. And what it produces isn’t just memorized answers, but uses the simplest thing from what you’ve learned about distributions: how else can I approach this object? What about in the real world? When a robot enters a room, its perception is very messy. So we also did one thing: have a student use the iPhone to scan the room. The iPhone can indeed perform 3D scanning, but the accuracy is not high. What it produces is very noisy and bulky, like blobs of material. Even if we feed this highly noisy output into our system, you’ll find that it actually has a bit of zero-shot generalization: it can infer how you should interact with a dishwasher, a refrigerator, or a microwave oven. This would be much more useful for robots. Because the robot will enter a room you have never seen before, even though the input is so rough, it will still say, “I have never seen this specific faucet, but I can reason out various ways to deal with it, imagine different methods of handling it, and then use this thing to devise a strategy.” The same is true for the office room; the reconstruction result is just as bulky, but it can reason out how to interact with a laptop. However, there is another problem with this reasoning. It can provide forty ways of interacting with this object, but you were originally aiming to achieve a certain goal. You don’t know which of these forty methods is truly useful for you. So we are also wondering whether we can add language conditions to this paradigm so that we can use it more effectively. For example, if I have a static scene and I say a sentence: “A child jumps on a seesaw,” can we generate a moving three-dimensional scene, or even a four-dimensional scene? This allows for interactions between people and objects, as well as actions on both sides when children jump onto them. This is much more useful, because you can then ask: How should the robot interact with this object? You can generate models in that direction or in another direction, and then select the action that actually helps solve the problem. How do you achieve this? It still depends on a large number of pre-trained generation models. In this task, the main model we rely on is the video diffusion model. Suppose you have this static scene, which can be very messy and in any situation, you can convert it into an intermediate representation with thousands of points, and then you can add a dimension—time—to it. Thus, you get a four-dimensional representation. But if you directly upgrade a three-dimensional object to a four-dimensional one, it remains static; how do you make it move? Here’s a way. If you only have a static shape and want to make it more realistic, there is a technique called Score Distillation Sampling (SDS). In conventional 3D, the approach is as follows: you render this shape to get videos from different perspectives, then feed them into a video diffusion model to update the representation forward in time. The same thing can be done in 4D. Practically, it’s very difficult; you need to find solutions for a variety of problems and design data structures for the 4D representation. But conceptually, it’s just about using video models and pre-trained diffusion models to guide you in creating sequences of interactions between real people and objects, or between robots and objects. Here are two more examples. For instance, if you have a brick and add the sentence “The robot is picking up the brick”, it will generate an interaction sequence where the robot picks up the brick. This shows you what the interaction looks like under a different language condition. You can also add some free demos and learn skills from them. Previously, you had to have human-operated videos or remote operation data before you could learn skills and incorporate them into planning; now I can do without any human demonstrations. I generate my own demonstrations, and then learn skills from these demonstrations. After learning, I put them back into the original training framework. It isn’t necessary to use it only on robots. For example, a basketball hitting the basket and bouncing off, or a cat jumping onto a cushion—these can all be done. But we have indeed applied it to robots as well. For instance, there’s a piece of cloth. You ask me to generate a sequence showing human interaction with this cloth and unfolding it. On the left is the generated result. Since it is generated, you get a set of free demonstrations, and you can obtain the dense trajectory of points on the cloth in three dimensions, which is tracked over time in three dimensions. These become a set of free monitoring targets, letting the robot know: I want this object to move like this. On the right is the real robot trying to do the same thing. Starting from imagination, using the diffusion model as a guide, and treating its output as a success goal, you learn a strategy to do this. Another example is closing the laptop computer. Recently, we have also been working on making it much faster. Because in such methods, if you deal with diffusion units, it is very slow, and fractional distillation in four dimensions is even slower. So we used caching and semantic rendering to increase the proportion of successful updates, so you don’t need to call the video model so frequently. This allows you to generate demonstration sequences with a longer duration. Previously, I might have generated just one demo—the robot picking up bricks or closing a notebook—which was to help you learn atomic skills. This isn’t a robot yet, but there’s no reason it can’t become one. You can generate very long-duration interleaved sequences and do completely different things. So what you get is not just a free demo of atomic skills, but a free demo of complex long-duration operation problems, which can provide unlimited data for your skill training. Okay, I’ll stop here.

05


Structure is not the answer, but a scaffold

Looking back, what we have been talking about is how to identify these structured representations. You will find that as time passes, there are more and more such things. Now, almost every component is available, but you still retain that intermediate representation layer. It now exists in the form of language, and of course, it corresponds to scene representations behind it, which makes the entire system understandable.

We have also been thinking about how to make it more general-purpose, so that it isn’t bound by existing objects or existing demonstrations.

Regarding the term “structure”, I would like to say more. When others talk about structure, it’s easy to misinterpret it as something else: you need to forcefully integrate that structure into your system. That isn’t good, and it’s not what we’re doing. What we want is to see if this interface can help the system operate more efficiently; it acts like an external scaffold, while also allowing humans to communicate with it. Fundamentally, you are using those data-rich modalities, especially videos, as well as language, and the three-dimensional and four-dimensional data you can obtain. But ultimately, what matters to you is another thing: how these structured representations enable us to have reasoning abilities. You have atomic-level representations; I will clean them, store them, and do these things. Next, what are the dependencies between these tasks, and what are their effects? What really enables you to solve long-term problems in a very general way is this. Thank you all.

To facilitate group communication and sharing of meeting updates, we have specifically created the IROS2026 participation and exchange group! What benefits will you get by joining the group?

  • Meeting Information Station: Submissions of DDL, format guidelines, review progress… key updates delivered immediately, so you never miss the deadline. Your submission preparation is well-prepared.

  • Peer Discussion Group: All peers from different places are in the group. We share insights, exchange ideas, and build connections. We walk together on the path of scientific research.

  • Frontier Supply Station: Latest research findings and hot topics are synchronized in real time. Combined with the conference theme, it helps you clarify the flow of your submission and provides inspiration for your paper.

  • Venue Support Group: (Available only during the meeting):Don’t panic when you get to the venue—

    Route Navigation: How to reach the venues and the lecture hall, ask in the group at any time;

    Topic Reading: Popular sessions and Keynotes are available for review even if you missed them;

    Poster compilation: A collection of key live poster content, easily saved with one click without any omission;

    Yueju Square: Meal gatherings, interview sessions, communication forums… If you want to chat, collaborate, or meet in person, it can be organized right there in the group.

? Group entry link: Scan the code to join the group, or add WeChat Luyoyo_2026, and note: ECCV + institution/school + name + field of study.

In scientific research or technology work, information gap is very important.

Come on, take a step ahead!

Get on the bus, and I’ll show you the highlights of global AI conferences

Exclusive access:

Expert Speech PPT

Full Report of the Conference

Interpretation of Popular Papers

Interview with Academic Star

Scan the QR code above

Or click “Read the original article” to follow the section.

Leifeng Net original article, unauthorized reproduction is prohibited. For details, see Reproduction Guidelines.

Original source

雷峰网 AI

Content notes

Original publication and rights belong to the source.

Machine translation · Refer to the original