Embodied AI and vision-language-action models, explained

Embodied AI is intelligence that acts through a body. This guide explains how vision-language-action models turn images and a request into motion, which models shaped 2026, and where they still fall short.

By Synthra RoboticsPublished 11 min read

What is embodied AI?

Embodied AI is artificial intelligence that perceives and acts through a body, such as a robot, in a physical or simulated environment. A chatbot reads and writes text. An embodied system takes in camera images, sound, and its own motion, decides what to do, moves, and then sees what its movement changed.

Researchers treat perception, action, memory, and learning as the working parts of an embodied agent (Paolo et al., ICML 2024). What changed in 2023 is that large models pretrained on internet data began to drive real robots directly.

How embodied AI differs from a chatbot

  • Inputs. Not a prompt, but a continuous stream of camera frames, audio, and joint or wheel positions.
  • Outputs. Not text, but motion: joint targets, velocities, gripper commands, or a path.
  • Timing. A robot has to keep acting while the scene changes, many times per second.
  • Consequences. A wrong sentence can be edited. A wrong movement can spill a drink or bump into a person.
  • Data. Robot action data is not online; it has to be collected, often by people teleoperating robots.

What is a vision-language-action (VLA) model?

A vision-language-action (VLA) model is a neural network that takes camera images and a natural-language instruction and outputs robot actions. Most VLAs start from a vision-language model pretrained on web images and text, then are trained on robot demonstrations so that the same network that describes a scene can also produce motor commands.

Google DeepMind’s RT-2 paper named the category in July 2023 (arXiv). The web pretraining is the point: RT-2 could pick a rock as an improvised hammer, drawing on knowledge from web data rather than robot data (Google DeepMind).

How a VLA model works

  1. Vision encoder. Camera images become embeddings, vectors that describe what is in view. OpenVLA fuses two pretrained encoders, DINOv2 and SigLIP (OpenVLA).
  2. Language backbone. A pretrained language model reads the instruction with the image embeddings, bringing web knowledge of objects, places, and what things are for.
  3. Action output. Some VLAs bin each motion dimension and emit actions as tokens; RT-2 used 256 bins per dimension. Others use a continuous head based on diffusion or flow matching, which refines random noise into a smooth trajectory. π0 attaches a 300-million-parameter flow-matching “action expert” to a 3-billion-parameter vision-language model (π0).
  4. Closed loop. The model runs again on each new observation. π0 predicts chunks of 50 actions and controls robots at up to 50 Hz; the 55-billion-parameter RT-2 ran at 1 to 3 Hz from cloud TPUs.

Why many VLAs have two systems

Large models are slow; robots need fast reflexes. Figure’s Helix runs a 7-billion-parameter vision-language model at 7 to 9 Hz for understanding and an 80-million-parameter policy at 200 Hz for control, on embedded GPUs in the robot (Figure). NVIDIA calls this split System 2 and System 1. Google DeepMind pairs a reasoning model that plans with a VLA that acts.

What is the history of VLA models?

VLA models grew out of Google’s robotics research. RT-1 (December 2022) learned hundreds of real-robot tasks with one transformer, and RT-2 (July 2023) introduced the term VLA by training a web-pretrained vision-language model to output actions. Open X-Embodiment, OpenVLA, and π0 then opened up the data and the models by early 2025.

  • December 2022, RT-1 (Google): about 130,000 episodes and 700+ tasks from 13 robots over 17 months (Google Research), with code and checkpoints released.
  • July 2023, RT-2 (Google DeepMind): actions as text tokens; success on unseen scenarios rose from RT-1’s 32% to 62%.
  • October 2023, Open X-Embodiment: over 1 million real robot trajectories from 22 robot types and 34 labs (project page).
  • June 2024, OpenVLA: a 7-billion-parameter open model that beat the closed, 55-billion-parameter RT-2-X by 16.5 percentage points across 29 tasks.
  • October 2024, π0 (Physical Intelligence): flow matching on 10,000+ hours of robot data; weights released in February 2025 (The Robot Report).
  • February to April 2025: Figure’s Helix, Google DeepMind’s Gemini Robotics, NVIDIA’s open GR00T N1, and π0.5, which cleaned kitchens and bedrooms in unseen homes (π0.5).
  • September 2025, Gemini Robotics 1.5: reasons in natural language before acting and transfers skills across robot bodies (Google DeepMind).

What are the notable VLA models in 2026?

The main robot foundation models released or announced in 2026 come from Google DeepMind (Gemini Robotics 2), NVIDIA (Isaac GR00T N1.7, with GR00T N2 previewed), Figure (Helix 02 and Helix 2.5), and Physical Intelligence (π0.7). Tesla describes an end-to-end approach that it says carries over from its cars to its Optimus humanoid.

Notable VLA and robot foundation models, by first public announcement.
ModelOrganizationFirst releasedOpen weights?What it does
RT-2Google DeepMindJuly 2023NoVision-language model that outputs actions as text tokens.
OpenVLAStanford, UC Berkeley, and collaboratorsJune 2024Yes7B open VLA trained on 970,000 robot demonstrations.
π0 (pi0)Physical IntelligenceOctober 2024Yes (February 2025)Flow-matching action expert on a 3B vision-language model; up to 50 Hz.
HelixFigure AIFebruary 2025Not released7B model at 7–9 Hz guides an 80M policy at 200 Hz, onboard.
Gemini RoboticsGoogle DeepMindMarch 2025Not releasedGemini 2.0 with physical actions added as an output.
Isaac GR00T N1NVIDIAMarch 2025YesHumanoid model: vision-language module plus diffusion-transformer action module.
π0.5 (pi0.5)Physical IntelligenceApril 2025Yes (September 2025)Co-trained on mixed data; cleans rooms in unseen homes.
Helix 02Figure AIJanuary 2026Not releasedAdds a 1 kHz learned whole-body controller for balance and locomotion.
Isaac GR00T N1.7NVIDIAMarch 2026 (early access)Yes (NVIDIA Open Model License)3B humanoid VLA on Cosmos-Reason2-2B; commercial use allowed.
π0.7 (pi0.7)Physical IntelligenceApril 2026Not releasedSteerable with metadata and subgoal images as well as instructions.
Gemini Robotics 2Google DeepMindJuly 2026No (early-access partners)Whole-body humanoid and bi-arm VLA, directed by Gemini Robotics ER 2.
Helix 2.5Figure AISeptember 2026Not releasedPretrained on human-behavior data; tested zero-shot in 30 unseen homes.

Google DeepMind: Gemini Robotics 2

On July 30, 2026, Google DeepMind announced Gemini Robotics 2, a VLA that controls whole humanoids, legs included, and bi-arm robots; Gemini Robotics ER 2, a reasoning model that plans tasks, talks with people, and directs the VLA through tool calls; and Gemini Robotics On-Device 2, which runs on the robot. ER 2 is in Google AI Studio; the VLAs go to early-access partners (Google DeepMind).

NVIDIA: Isaac GR00T N1.7 and N2

GR00T N1.7, in early access since March 2026, has downloadable weights licensed for commercial use (model card). NVIDIA also previewed GR00T N2 and said it would be available by the end of 2026 (NVIDIA). N2 is a world action model: a policy built on a video world model, which learns how scenes change as well as what to do (NVIDIA Technical Blog).

Figure: Helix 02 and Helix 2.5

Helix 02 (January 2026) adds System 0, a 10-million-parameter network running at 1 kHz for whole-body balance and motion. It unloaded and reloaded a dishwasher in a four-minute autonomous run (Figure). Helix 2.5 (September 2026), pretrained on Figure’s human-behavior dataset, succeeded on 56% of zero-shot trials in 30 unseen homes, against 9% when trained from scratch (Figure).

Physical Intelligence: π0.7

π0.7 (April 2026) can be prompted with performance metadata and subgoal images as well as instructions, so it can learn from imperfect data yet be steered toward expert behavior (π0.7). Its researchers told TechCrunch it still needs step-by-step coaching on unfamiliar multi-step tasks (TechCrunch).

Tesla

Tesla said in 2024 that FSD v12 replaced over 300,000 lines of C++ with a single end-to-end neural network trained on video (Electrek). In 2025 its vice president of AI software, Ashok Elluswamy, said the approach carries over to Optimus (Humanoids Daily). More in Tesla Optimus vs. service robots and humanoid robots in 2026.

What do VLA models still struggle with?

VLA models still struggle with scarce robot data, unreliable generalization to new places and objects, latency when large models must drive fast motion, safety around people, and evaluation, since no standard benchmark exists. Even strong 2026 results report success rates well below the reliability that businesses expect from machines.

  • Data scarcity. The internet holds almost no robot action data (IEEE Spectrum), so labs collect it through teleoperation, human video, handheld devices, and simulation. Xiaomi reported training on more than 100,000 hours of real-world manipulation data in 2026 (arXiv).
  • Generalization. RT-2’s physical skills stayed limited to those in its robot data, and Google DeepMind’s On-Device 2 model card notes limited generalization to unfamiliar tasks. Gemini Robotics 2 still finds multi-finger manipulation hard, with 32% to 92% success across such tasks, and Figure’s 56% zero-shot result means 44% of trials failed.
  • Latency. RT-2’s authors warned that real-time inference could become a bottleneck. Figures cited on NVIDIA’s technical blog put world action models at 590 to 800 milliseconds per action chunk, against about 190 for π0.5.
  • Safety. Google DeepMind’s ASIMOV benchmarks test whether models recognize unsafe actions, and its model card recommends layers: a reasoning model for semantic safety, low-level controllers for collision-free motion, and hardware safeguards (model card).
  • Evaluation. Physical Intelligence’s researchers conceded that standard robotics benchmarks barely exist. RoboArena ranks policies by double-blind pairwise comparisons (arXiv), and Ai2’s MolmoSpaces, a simulation platform with 230,000+ indoor scenes, varies one factor at a time (Ai2).

What does physical AI mean?

Physical AI is AI that perceives, reasons, and acts in the physical world through machines such as robots, autonomous vehicles, and camera systems. It overlaps with embodied AI. The industry term, used heavily by NVIDIA, puts extra weight on the simulation, synthetic data, and world models used to train these systems.

NVIDIA applies the term to cameras, robots, and self-driving cars (NVIDIA); IBM defines it as AI that operates in the physical world rather than only in software (IBM). A VLA is one kind of physical AI model. Because robot data is scarce, the field leans on world foundation models, generative models that simulate how scenes evolve, such as NVIDIA’s Cosmos, launched at CES 2025 (NVIDIA).

Physical AI is also a theme of the Synthra research program: learning from interaction with the physical world, and the data, simulation, and evaluation that such learning requires.

How does embodied AI work in a restaurant service robot?

In a restaurant, embodied AI turns a request into a delivered order. The robot perceives guests, staff, tables, and free space; grounds words like “bring me the pasta” in a specific dish, guest, and table; decides a path; acts; and re-plans when someone steps into the aisle, instead of following a stored route.

Most VLA results above involve manipulation, such as folding laundry or loading dishwashers. A restaurant robot mainly moves through people, talks with guests, and hands over orders, in a room that changes by the minute. The loop still applies.

From “bring me the pasta” to a delivered plate

  1. See. Track guests, staff, tables, chairs, and free space, and note where the request came from: a guest at table 04.
  2. Understand. Ground the words. “Me” is that guest and table; “the pasta” is a dish on the table’s order, waiting at the kitchen pass. If two pastas were ordered, a short clarifying question is better than a guess.
  3. Decide. Turn the request into a task (collect at the pass, deliver to table 04) and choose a path through the room as it is now.
  4. Act. Follow the path at a pace suited to the load, keep clear of people, and stop within the guest’s reach.
  5. Adapt. When a server steps into the aisle, keep the goal and re-decide the path.

Every Synthra U1 model runs this loop; the technology page walks through an illustrative table 04 delivery. The restaurant supplies context the robot cannot see: it uploads its menu, dietary and service information, and operating parameters, and sets a standby area and docking location. The robot builds the rest of its venue model by exploring. Unlike route-following robots with predefined waypoints, it decides the path at the moment of service.

Synthra’s intelligence platform combines multimodal perception, vision-language-action reasoning, spatial understanding, contextual navigation, conversational service, and automatic docking across U1e, U1, and U1 Max. See U1 and the restaurant robot guide.

Frequently asked questions

Is ChatGPT an example of embodied AI?

No. Chatbots such as ChatGPT work in software and do not act in the physical world. The same kind of model can become part of an embodied system, though: VLAs such as RT-2 and Gemini Robotics start from a vision-language model and add robot actions as an output, which makes them embodied.

What is the difference between a VLM and a VLA?

A vision-language model (VLM) takes images and text and produces text, such as a caption or an answer. A vision-language-action model (VLA) also produces actions, so it can drive a robot. Many VLAs are VLMs trained further on robot demonstrations: π0 starts from Google’s 3-billion-parameter PaliGemma, and Gemini Robotics was built on Gemini 2.0.

Are there open-source VLA models?

Yes. OpenVLA released a 7-billion-parameter model and training code in 2024. Physical Intelligence publishes π0, π0-FAST, and π0.5 weights in its openpi repository, and NVIDIA publishes Isaac GR00T weights on Hugging Face, with N1.7 licensed for commercial use. As of September 2026, the newest models from Google DeepMind, Figure, and Physical Intelligence are not released as open weights.

How fast does a VLA control a robot?

It depends on the layer. RT-2’s 55-billion-parameter model ran at 1 to 3 Hz, and π0 at up to 50 Hz. Figure’s Helix runs its vision-language model at 7 to 9 Hz and its motor policy at 200 Hz; Helix 02 adds a 1 kHz whole-body controller. Fast inner loops keep balance; slower outer loops handle understanding.

Does a service robot need a humanoid body to use embodied AI?

No. Embodied AI describes how a robot perceives, reasons, and acts, not the shape of its body. VLAs have been trained on single arms, bi-arm systems, mobile manipulators, and humanoids, and Open X-Embodiment spans 22 robot types. Synthra’s U1 Series runs one intelligence platform on three physical configurations that differ in size, payload, and endurance, never in intelligence.

Sources

  1. A call for embodied AI, arXiv (ICML 2024 position paper), 2024-02-06
  2. RT-1: Robotics Transformer for real-world control at scale, Google Research, 2022-12-13
  3. robotics_transformer (RT-1 code and trained checkpoints), Google Research on GitHub, 2022-12
  4. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, arXiv (Google DeepMind), 2023-07-28
  5. RT-2: New model translates vision and language into action, Google DeepMind, 2023-07-28
  6. Open X-Embodiment: Robotic Learning Datasets and RT-X Models, Open X-Embodiment Collaboration, 2023-10-13
  7. OpenVLA: An Open-Source Vision-Language-Action Model, arXiv, 2024-06-13
  8. π0: A Vision-Language-Action Flow Model for General Robot Control, arXiv (Physical Intelligence), 2024-10-31
  9. Physical Intelligence open-sources Pi0 robotics foundation model, The Robot Report, 2025-02-07
  10. openpi, Physical Intelligence on GitHub, 2025-09
  11. π0.5: a Vision-Language-Action Model with Open-World Generalization, arXiv (Physical Intelligence), 2025-04-22
  12. π0.7: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities, arXiv (Physical Intelligence), 2026-04-16
  13. Physical Intelligence, a hot robotics startup, says its new robot brain can figure out tasks it was never taught, TechCrunch, 2026-04-16
  14. Helix: A Vision-Language-Action Model for Generalist Humanoid Control, Figure AI, 2025-02-20
  15. Introducing Helix 02: Full-Body Autonomy, Figure AI, 2026-01-27
  16. Helix 2.5: Zero-Shot 30-Home Generalization, Figure AI, 2026-09-17
  17. Gemini Robotics brings AI into the physical world, Google DeepMind, 2025-03-12
  18. Gemini Robotics 1.5 brings AI agents into the physical world, Google DeepMind, 2025-09-25
  19. Gemini Robotics 2 brings whole body intelligence to robots, Google DeepMind, 2026-07-30
  20. Gemini Robotics On-Device 2 model card, Google DeepMind, 2026-07-30
  21. NVIDIA Announces Isaac GR00T N1, the World’s First Open Humanoid Robot Foundation Model, and Simulation Frameworks to Speed Robot Development, NVIDIA Newsroom, 2025-03-18
  22. nvidia/GR00T-N1-2B model card, Hugging Face (NVIDIA), 2025-03
  23. nvidia/GR00T-N1.7-3B model card, Hugging Face (NVIDIA), 2026-04
  24. NVIDIA Isaac GR00T N1.7: Open Reasoning VLA Model for Humanoid Robots, Hugging Face blog (NVIDIA), 2026-04-17
  25. NVIDIA and Global Robotics Leaders Take Physical AI to the Real World, NVIDIA Newsroom, 2026-03-16
  26. Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models, NVIDIA Technical Blog, 2026-06-15
  27. NVIDIA Launches Cosmos World Foundation Model Platform to Accelerate Physical AI Development, NVIDIA Newsroom, 2025-01-06
  28. What is Physical AI?, NVIDIA Glossary, 2026-09
  29. What is physical AI?, IBM Think, 2026-01-19
  30. Tesla finally releases FSD v12, its last hope for self-driving, Electrek, 2024-01-22
  31. Tesla AI Chief Details Unified ‘World Simulator’ for FSD and Optimus, Humanoids Daily, 2025-10-24
  32. Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories, arXiv (Xiaomi), 2026-07-16
  33. Will Scaling Solve Robotics?, IEEE Spectrum, 2024-05-28
  34. RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies, arXiv, 2025-06-22
  35. MolmoSpaces, an open ecosystem for embodied AI, Ai2, 2026-02-11

Product and company names mentioned in this article are trademarks of their respective owners and are used only to identify those products. Synthra Robotics is not affiliated with, sponsored by or endorsed by them.