雷峰网 AI

Zhi Jian Dynamic Feng Zongbao: Realizing robot making robot in a gap of only 0.5 millimeters | IROS 2026

Embodied intelligence is moving from laboratories to real factories and toward precision manufacturing. Author: Deng Zhemin Editor: Qi Chengyong. Have you ever thought that robots could also enter workshops and manually…

Embodied intelligence moves from laboratories to real factories and toward precision manufacturing.

Author | Deng Zhemin

Edit | Qi Caiyong

                                                                                                       

Have you ever thought that robots could also enter the workshop and manually produce the parts needed to assemble a robot? It sounds like a joke, but in reality, it is the most challenging implementation challenge for embodied intelligence. The joints, chucks, and precise structural components of robots are all manufactured through steps of CNC (Computer Numerical Control) lathes and milling machines. If robots can independently monitor these machines, perform grasping, loading, insertion, removal, and placement of workpieces, then the robot industry will be supported by robots themselves—the more robots there are, the stronger the ability to build robots becomes and the lower the cost, creating a positive feedback loop. However, achieving this is much more difficult than imagined. The parts required to build robots are precisely those with the highest demands for precision: the gap between the workpiece and the chuck is often only 0.5 millimeters, while the reflective metal surface, the nearly textureless shape, and the highly similar parts all challenge the robot’s vision, touch, and overall precision. No matter how beautiful the demonstrations in the laboratory, they must first pass through the three hurdles of tolerance, cycle time, and yield rate. On September 29, 2026, in Pittsburgh, USA, at the top-level robotics conference IROS 2026, a Tech Talk directly addressed this challenge: how to enable a robot to independently monitor two CNC machines and perform all the above operations within a 0.5-millimeter gap, using only 600 real robot trajectories for training. The speaker was Feng Zongbao, head of Simplest Power Reinforcement Learning. This was not just a technical demonstration; it was also a solution to how laboratory precision can be transformed into factory precision. Below is a condensed version of Feng Zongbao’s speech at IROS 2026, organized based on the English speech at the event. Leifeng Network (public account: Leifeng Network) has reorganized the spoken language and technical terms without changing the original meaning.

▎From Laboratory to Factory: Why Real-World CNC Lathe Manipulation Breaks Current Robot Learning Pipelines

Speaker: Feng Zongbao, Head of Reinforcement Learning at Simplexity Robotics

01

Building robots with robots

This sharing starts with a manufacturing goal—manufacturing robots with robots. CNC operations may seem simple, but making them stable and reliable in real factories is a completely different issue. It reveals the clear gap between the success of laboratory benchmark tests and factory-level precision. Today, I will show how we strive to bridge that gap. SimpleClaw Power is a full-stack embodied intelligence company. Our self-developed multimodal models combine understanding and generation; a closed-loop data flywheel allows us to continuously improve these models; a unified modular hardware platform enables them to operate on real robots; the SimpleClaw platform provides developers with an open platform for building and deploying robot applications. On the hardware side, we have four product lines. i7 Pro is our wheeled robot, and iX is our wheeled humanoid robot. We also offer two configurations of modular force-controlled arms, which are shown here with Dexter dexterous hands and grippers. Finally, our data collection gloves help us scale up human demonstrations, providing data for robot learning. Therefore, our approach connects three components: our full-stack self-developed solutions from hardware, software to models; the data collection gloves that make scaling up human demonstrations possible; and the multimodal models that combine understanding and generation. The goal is to learn more from this data and raise the ceiling of robots’ capabilities in precise CNC operations. This sharing focuses on three complementary methods: SimpleWAM generates action segments from multimodal observations; DRAM, namely Delta rule cyclic association memory, processes long-time-series contexts using a fixed-size memory matrix; DPE scores candidate actions with a separately trained evaluator while keeping the strategy frozen. Together, they address the action generation, memory, and action selection problems in our CNC system.

02

Real Factory Challenges

This is the manufacturing challenge we face. Our goal is for robots to create robots, and precision CNC operation is a representative example of this. Why is it difficult? First, the gap between the workpiece and the chuck is only 0.5 millimeters, leaving almost no room for alignment errors. Second, reflective surfaces, subtle textures, and parts with similar shapes make visual alignment difficult. Third, the robot must perform a series of actions within a limited working space, and the key contact interfaces are not always visible. Therefore, the system must maintain precision throughout the entire task process, even when the visual information is incomplete. These CNC constraints shape our architecture: weak textures require a sufficiently strong visual encoder; long-distance proximity requires multi-viewpoints provided by cameras on the head and wrist; precise contact requires conditional conditions for proximity perception. Thus, visual, linguistic, and state inputs are converted into tokens. SimpleWAM provides a multimodal backbone and Action DiT. We insert a DRAM module between them to supplement historical context beyond current observations. Action DiT proposes candidate action segments, and DPE scores these candidates with an independently trained evaluator, while keeping the trained action proposer unchanged. The selected action segments are then sent to i7 Pro.

03

Specific Response Measures

SimpleWAM is our unified world action model. Its idea is: learn from rich future supervision during training, and directly predict actions during reasoning. During the training phase, video DiT, 3D DiT, and action DiT jointly reconstruct future video frames, 3D states, and action segments under task instructions, enabling the model to learn dynamics of vision, space, and actions simultaneously. During the reasoning phase, we only input the current image, current 3D state, and language instructions, and the model directly outputs action segments without generating future images or 3D states. Before implementation, we extract features separately and then perform adaptive fusion on them. A shared frozen encoder retains dense image tokens from head and wrist cameras; subsequently, each view is read independently, so none of the views suppresses the others before fusion; gating balances global context and local alignment cues for each query and each feature channel. Finally, by fusing visual history with robust, state-conditionized action DiT, action segments are predicted. The attention diagram on the right shows a stronger boundary focus effect. In the nominal experiments of all methods, the fusion method succeeded 15 times, in contrast, Query-concay had 0 successes, and π0.5 SFT baseline had 8 successes. To maintain stable actions throughout a long task, a robot needs context from previous steps. For this, we introduce DRAM—Delta Rule Recurrent Associative Memory—to attach this capability to the pre-trained robot strategy. Here, the frozen SimpleWAM backbone extracts features from the current observation; DRAM first reads relevant history from these features in memory; action DiT combines the read history with the current context to generate action segments; afterward, the memory is updated using Delta rules via gating for use in the next step. The memory is a fixed-size associative matrix, so its size remains consistent as the task length increases. DRAM can be added without modifying the backbone architecture or retraining it, making long-time-series memory easier to integrate into existing strategies. During precise insertion, vision alone may not allow judgment of the state of the workpiece touching the chuck. Force and torque exactly fill the missing contact information. We use them in two ways: first, proximity perception conditioning adjusts the robot’s state input, allowing action DiT to consider contact factors when generating the next action; second, a real-time impedance controller uses real-time feedback to make execution more robust and adjusts motion based on changes in contact. Together, they enable recovery when direct insertion is blocked. In precise CNC insertion tasks, this combined approach achieved 100% task success rate. Common practice is to update the actioner with evaluator feedback, but this may lead to strategy drift. DPE first trains a behavior-cloned actioner, then uses a separately trained evaluator to score candidate actions during deployment; successful and failed rollouts are then updated only by the evaluator, forming a closed loop, while the actioner remains frozen on the robot. DPE improves task success rate and reduces execution steps.

04

Result Presentation

We finally achieved a complete operation plan for a single-device fully automated double CNC machine tool. The main demonstration video presented the entire autonomous operation process of the robot at five times the speed, and through multiple segmented segments, it showed three core independent processes: autonomous workpiece grasping, precise loading chuck operation, and picking up and positioning of the processed parts after processing. It is worth noting that throughout the entire model training process, we only used 600 real robot motion trajectories as training data. This is also the core value of our work this time: bringing embodied intelligence from laboratory benchmark testing results into high-precision, highly realistic industrial manufacturing scenarios.

To facilitate group communication and sharing of meeting updates, we have specifically created the IROS2026 participation and communication group. What benefits will you unlock by joining the group?

  • Conference Information Station: Submission DDLs, format guidelines, review progress… key updates delivered immediately, so you never miss the deadline—your submission preparation is well-organized.

  • Peer Discussion Meeting: All peers from all over are in the group. We share insights, exchange ideas, and build connections—we travel together on the path of scientific research.

  • Frontier Supply Station: Latest research findings and hot topics are synchronized in real time. It helps you clarify the flow of submissions based on the conference theme, providing plenty of inspiration for your paper.

  • On-site support group: (activated only during the meeting):You don’t need to panic when you get to the venue—

    Route Navigation: How to get to venues and the lecture hall, ask in the group at any time;

    Topic Discussion: Popular sessions and Keynotes are available for viewing even if you missed them;

    Poster compilation: A summary of key on-site posters, save them with one click without missing anything;

    Yueju Square: Meal gatherings, interview sessions, exchange meetings… If you want to chat, collaborate, or meet in person, it can be started right here in the group.

? Group entry link: Scan the code to join the group, or add WeChat Luyoyo_2026, and note: ECCV + institution/school + name + research direction.

In scientific research or technology work, information gap is very important.

Come on, take a step ahead!

Get on the bus, and I’ll show you the highlights of global AI conferences

Exclusive access:

Expert Speech PPT

Full Report of the Conference

Interpretation of Popular Papers

Interview with Academic Rising Star

Scan the QR code above

Or click “Read the original article” to follow the section.

Leifeng Network original article; reproduction is prohibited without authorization. For details, see Reproduction Guidelines.

Original source

雷峰网 AI

Content notes

Original publication and rights belong to the source.

Machine translation · Refer to the original