Humanoid robots guided by foundation models, the large-scale artificial intelligence systems that power modern chatbots, have begun taking commercial shifts in warehouses across the United States, Germany, and Japan. During the second quarter of 2025, Figure AI, Agility Robotics, and 1X Technologies transitioned from pilot tests to paid, multi-unit deployments, placing robot workforces in logistics centers and automotive assembly support roles. The reason captured industrial attention is straightforward: transformer-based models trained on video and sensor data finally give machines the spatial and language judgment that decades of traditional automation could not.
Outside the lab, these robots now unpack pallets, sort returns, and reposition storage containers while adapting their own next action based on what they see in real time.
Context: Why Older Factory Robots, Called Cages, Lag Behind
Industrial robotics historically rely on thoroughly taught sequences. A traditional articulated arm repeats the same wrist and shoulder movements with high precision, but it cannot parse a messy shelf, an unknown box, or a room where items have shifted positions
This limitation created the “robotics paradox”: capable arms that are functionally dense. To overcome that, changers placed robots in fixed cells and choreographed every coordinate, which in turn raised integration costs, created closures, and limited tasks to strips where fluid bendable objects played by standing.
The Vision-Language-Action Breakthrough
The new approach starts with large vision-language models trained on gigabytes of image-caption pairs, then adds an “action” head—a control layer that converts embedded features into joint commands for a robot’s arm and legs.
Consumers call these “Vision-Language-Action” (VLA) models. When a robot reads a user prompt such as “pull that red-hold the clamp””… I’m repeating; fix in final.