Chinanews
Other爱范儿Thu, 27 Aug 2026 10:34:37 +0000

After the Robot Games, X-Square Built a 'Virtual Training Ground' for Robots

After the Robot Games, X-Square Built a 'Virtual Training Ground' for Robots

My feeds have been flooded with robots for a week.

Robots broke the human 100-meter record, "took off" from a standing position for nearly three meters, a "shy run" posture went viral overseas, and a cyber "Zhang Sanfeng" executed a 540-degree aerial kick...

**This year's World Humanoid Robot Games were even livelier. A total of 666 teams from 16 countries across six continents brought 2,056 robots to Beijing. The competitions extended from track and field and martial arts to real-world scenarios such as homes, hotels, industrial settings, and emergency rescue. However, when robots mess up on the field, people just laugh it off. But once robots actually enter homes, they might knock over water glasses, crash into furniture, or drag freshly folded clothes back onto the floor—and that won't be funny anymore. In the real world, there are no fixed tracks or uniformly placed props. If a cup is moved slightly or a rag is left on the table, originally trained movements can easily fail. Moving these robot trial-and-error processes into the virtual world is exactly what current world models aim to do. They attempt to replicate an environment governed by physical laws, allowing robots to rehearse before actually acting: will moving the hand slightly to the left miss the grasp? Will closing the gripper early knock over the cup? But in reality, many world models cannot achieve reliable "virtual testing." They might understand the visual frame but fail to truly obey the actions. As the simulation time extends, objects and positions begin to drift. Strategies that perform better in the virtual world might not necessarily work on real robots. Just as robots begin to step out of the arena and into real-world scenarios, X-Square Robot released the autoregressive world model WALL-SS.

The "virtual training ground" built by WALL-SS has finally broken through the bottlenecks of world models. Whether actions can truly alter outcomes, whether coherence can be maintained in long-horizon tasks, and whether virtual results can withstand real-world testing—all now have new solutions and answers. It can simulate 60-second water pouring and object organizing tasks to test the outcomes of different actions: depending on how the cup is grasped, will the robot miss, knock it over, or successfully pick it up? The research team also placed multiple robot policies into both WALL-SS and real-world tests. Action strategies that performed well in the virtual world often yielded equally good results when transferred to the physical world. This is fascinating: the "dreams" robots have can now be used to practice and filter real-world actions.**

Related **Project Page: http://x2robot.com/pages/ss Paper Link: https://github.com/X-Square-Robot/wall-ss/blob/main/wall-ss-paper.pdf Github Link: https://github.com/X-Square-Robot/wall-ss

How a Robot "Dreams" Determines Whether It Fails in Reality

World models are currently one of the most closely watched directions in embodied AI. It can be understood as a rehearsal system inside the robot's brain. Given the current visual frame and a set of actions to be executed, it predicts what will happen next. If the robotic arm moves 1 centimeter to the left, will the cup be grasped or knocked over? The drawer is already open; should the next step be to continue retrieving items, or to adjust the wrist angle first? The robot can compare multiple futures before deciding which action path to take.

Robot companies can also test different versions of action policies within world models. The better-performing versions are then deployed on real robots, saving the exorbitant costs of testing action efficacy in the real world. The problem is that many world models often act more like an overly optimistic director. Robot training data typically consists mostly of successful demonstrations. If the model sees that most gripper-closing actions are followed by the object being picked up, it tends to take a shortcut: as long as the gripper closes, the object should be picked up. Even if there is a visible gap between the gripper and the object, the object in the generated frame might actively snap into the gripper. The paper refers to this phenomenon as a "magnetic grasp" error.

Such deviations might not be obvious in large movements like running or dancing, but in fine manipulation, they directly determine success or failure. If the gripper is a few millimeters off from the cup handle, the robot might miss. If the plug's angle is slightly off, it will be hard to insert smoothly into the socket. For a robot, moving from "knowing how to do it" to "actually doing it" often gets stuck in those final few millimeters. If a world model simply smooths out these few millimeters, the smoother the future it presents, the more likely the robot is to overestimate the reliability of an action. As simulation time extends, new troubles arise. After a world model generates a segment of video, it must use it as the basis for the next prediction. A tiny initial error will continuously accumulate later on. Objects might slowly drift from their original positions, and the contact relationship between the gripper and the cup will also change. If only the most recent frames are retained, the model forgets what it did before; but if the entire history is kept, the computational load and memory usage will continuously grow. A realistic-looking video doesn't mean the policy selected by the model is actually better. Robot developers need to know whether the winner in virtual testing can maintain its lead on a real machine. World models need to ensure that different actions genuinely lead to different outcomes; a smooth video only accomplishes the most superficial step.**

How the Robot Moves is How the Future Should Change

When WALL-SS predicts the next few seconds of video, it first builds a "low-resolution preview."

In this preview, the model first determines the general direction: where the robotic arm goes, and roughly how the objects will move. Subsequently, it sharpens the same video frame by frame, filling in object contours, gripper aperture, and contact positions.

The paper refers to this hierarchy, from broad outlines to fine details, as "scales." The already generated frames and past actions become the history for the next prediction—this is "next-scale autoregression."

As the video becomes clearer, the actions once again influence the future.

Whether the arm moves left or right will first change the direction of motion in the low-resolution video. When the gripper closes will continue to affect whether the cup is touched in the high-resolution video.

Once a short segment of the future is generated, it enters memory along with the actions that have occurred, becoming the basis for the model to continue...