XPeng's New VLA Arrives in September: Remembering 30 Seconds Past, Predicting 6 Seconds Future, Driving in 'Time'
In the movie *Doctor Strange*, Doctor Strange can rotate the Eye of Agamotto to make time flow forward or backward. By *Avengers: Infinity War*, he could even foresee multiple possible futures before his showdown with Thanos.
This is magic in the movies, but it is the reality offered by XPeng.
On August 27, XPeng held a Physical AI sharing session, where it also showcased a brand-new version of its second-generation VLA. The biggest keyword of this briefing was "time." XPeng added an "Eye of Agamotto" to the in-car model: turning it backward allows the system to retrace the world of the past 30 seconds, while turning it forward enables it to predict what will happen in the next 6 seconds.
Cars do not have a Time Stone, but they are powered by large models. Based on the information accumulated over a past period, the model calculates several subsequent states with higher probabilities and then selects an action for the current moment. For intelligent driving, having a "memory of the timeline" can bring about more human-like operations.
When driving, a person never just looks at the current frame. You judge the road conditions behind you based on memories from the previous few seconds, and you also anticipate an e-bike that might appear in a few seconds based on the deceleration of the car in front and the intersection on the side of the road. In a sense, driving happens not just on the road, but also in time.
The latest brand-new version 6.3.0 of XPeng's second-generation VLA is a model that drives within the flow of time.
The new VLA will begin rolling out to all Ultra and Ultra SE models in September, with the XPeng G9L Ultra and Ultra SE being the first to be equipped with it. Meanwhile, single-Turing chip Max models will also receive the second-generation VLA Lite distilled version in September.
Driving Has No Single-Frame Answer
XPeng's core concept for the new VLA this time is to connect the past, the present, and the future.
Infini-VLA is responsible for remembering the past. Its underlying architecture can support longer temporal sequences. Considering driving tasks and edge-side computational load, XPeng has set the historical memory of the mass-production version to 30 seconds.
In other words, it can remember the road conditions of the past 30 seconds and bring them into its current judgment.
XPeng showed a demo video at the briefing. When the car in front made a U-turn in the middle of the road, a gap briefly opened up behind its body that could be passed through. The test car did not accelerate immediately; instead, taking into account the previous movement process, it waited for the other car to complete the U-turn before moving forward.
What the model understands is the continuous behavior of "the car in front is making a U-turn," rather than just the single frame of "a gap has appeared ahead.
This is a bit like the short-term memory humans use when conversing with others. People don't forget the previous sentence every time they hear a new one, and a driver won't forget the preparatory U-turn maneuver the car in front made earlier just because that car is currently sideways across the road.
If Infini-VLA ensures the vehicle doesn't forget the context, then Streaming Inference is responsible for shortening the time it takes for the vehicle to understand the present. According to XPeng's explanation, the model can acquire input, perform inference, and output a trajectory simultaneously, without having to wait to fully understand a segment of history before taking action. Data provided by XPeng shows that this streaming inference has increased decision-making speed by 300%.
Similarly, this is akin to how humans drive: the eyes watch the road, the brain makes judgments, and the hands and feet operate simultaneously. When a target suddenly appears from an occluded area, such parallel processing eliminates the blank pause of waiting for inference to complete.
Remembering the past and focusing on the present naturally require looking to the future. Therefore, X-Foresight in the new VLA version will deduce the next 6 seconds based on historical states, and Flow Matching will fit the distributions of multiple possible trajectories to help the model choose a more appropriate action.
The terminology is obscure, but the resulting effect is very clear. Through continuous observation of current road conditions, the new VLA anticipates situations that have not yet occurred but are possible, and makes its next decision. The "defensive driving" often mentioned by experienced drivers in real life is essentially a form of anticipation.
In the on-site video, an oncoming white car occupied the opposing lane. The test car needed to determine whether the other party was going to wait, continue forward, or yield some space, and then deduce whether moving forward at that moment would block the paths of both cars. Seeing where the other party is located is only the first step; understanding what the other party is going to do determines one's next move.
Memories of the past and deductions of the future ultimately serve every acceleration, deceleration, and turn, allowing the second-generation VLA version 6.3.0 to "flow in time."
Entering the Real World
There are many terms in the architecture diagram, but returning to the product, what VLA faces is still the various troubles of real-world roads. XPeng's answer is to use a larger model. This time, XPeng has expanded the model's parameter count by 3.5 times.
In a video played at the briefing, the test car arrived at a ferry terminal. There were no clear lane lines at the dock, and a terminal is not a conventional road. However, relying on the model's own thinking, the vehicle still found the entrance on its own, drove onto the ferry, and stopped behind the car in front.
After the ferry docked, it drove out with the flow of traffic.
Another video took place on a narrow mountain road. After a sharp bend, a hanging branch appeared on the road ahead. The vehicle did not treat it as an ordinary road obstacle and stop in place, nor did it drive straight toward a gap that seemed to have space.
Instead, it proactively decelerated, observing the relationship between the branch and the vehicle's body while looking for a passable space.
It is worth mentioning that a few days ago, a Tesla running FSD v14.3.7 drove all the way from the second floor to the sixth floor of an apartment parking garage in the United States, and attempted to continue searching for a path toward the metal wire mesh on the top floor.
Fortunately, the owner took over near the wire mesh, and the vehicle did not fall. Faced with similar extreme "corner cases," XPeng's VLA and FSD made different judgments. The model must not only see the space but also understand whether or not it is a road.
Ferry terminals and tree branches belong to rarely seen scenarios, which is exactly why XPeng insists on expanding the model. A system relying on fixed rules can be tailored for docks, flocks of sheep, heavy rain, or every new...