Embodied intelligence has become the competitive core of the humanoid robot industry. Robot large models, serving as the central “brain,” determine a robot’s capabilities in perception, decision-making, and execution. Overseas, Google has completed technological iteration and ecosystem positioning through its Gemini series. Within China, companies such as Galaxy General, Agibot, and ByteDance have carved out differentiated technical approaches grounded in local scenarios. At the same time, AI chips — as the computing foundation for robots — are experiencing technological transformation and landscape reshaping driven by the demand for edge-side deployment.


Google’s Gemini Robot Large Model: Full-Chain Iteration, Building a General-Purpose Robot Ecosystem

Google has maintained long-term positioning in the robotics field, completing the evolution from hierarchical models to VLA (Vision-Language-Action) end-to-end models. Today, with Gemini at its core, Google is building an embodied intelligence system with cross-robot-type capabilities and strong reasoning ability, advancing both technology and commercialization on parallel tracks.

In the early period, Google’s robots adopted a hierarchical collaborative architecture: the SayCan large language model was responsible for outputting high-level motion instructions but could not operate independently, requiring the RT-1 execution model to implement physical actions. RT-1 achieved action output through the tokenization and compression of instructions and images, serving as the core underlying execution layer. Google subsequently launched PaLM-E, a multimodal foundation model integrating language and visual capabilities, and integrated RT-1 with PaLM-E to create RT-2 — the first VLA large model — achieving unified vision, language, and action capabilities. The follow-on product RT-2-X expanded multi-robot-type training data to further strengthen task generalization capability, laying the technical foundation for the Gemini series.

Entering the Gemini era, Google’s technology has fully pivoted toward productization and generalization. At the technical level, Gemini Robotics has been upgraded from a research version to an engineering version; Gemini 2.0 deeply integrates embodied capabilities, with built-in spatial, physical, and temporal reasoning modules. Gemini Robotics-ER achieves single-model adaptation across all categories of robots including humanoids, quadrupeds, robotic arms, and drones, and is regarded as “the Android system of the robotics field.” To address the scarcity of real-world data, Google uses Genie to generate millions of virtual scenarios, combined with SIMA to achieve bidirectional transfer between simulated and real data. Google is also pushing hard into edge-side deployment, developing the proprietary Gemini Control Hub SoC and NPU chip, and using model distillation technology to achieve edge-side real-time control at 100Hz, removing dependence on cloud infrastructure.

On the commercialization front, Google is deeply tied to Boston Dynamics. Its fully electric mass-production Atlas robot is equipped with Gemini Robotics 1.5 and Gemini Robotics-ER 1.5, responsible for end-to-end control and embodied reasoning respectively. Deployment scenarios are concentrated in industrial manufacturing — testing is scheduled at Hyundai Motor’s Meta-Factory in Georgia, USA in the first half of 2026, with Kia’s plant in the same facility planning small-batch deployment by the end of 2026, formally initiating large-scale industrial deployment.


China’s Leading Robot Large Models: Differentiated Breakthroughs, Solving the Industry’s Data Pain Points

Facing overseas technical advantages, leading Chinese enterprises have sidestepped homogeneous competition and, targeting the three major pain points of data barriers, general-purpose capability, and scenario adaptability, have each developed distinctively characterized VLA models with notable real-world deployment results.

Galaxy General’s Grasp VLA places synthetic data-driven development at its core, adopting a standard VLA architecture and pioneering a “synthetic data first” two-stage training mode. Relying on a high-fidelity simulation pipeline, it generates one billion frames of synthetic data within one week to complete pre-training, giving the model zero-shot generalization capability. Only a small amount of real-world data is subsequently needed for fine-tuning to adapt to scenarios such as supermarkets and industrial environments. The model specializes in upper-limb grasping, successfully overcoming the challenge of grasping transparent and reflective objects, with a real-robot grasping success rate exceeding 95% — a benchmark product for simulation data deployment.

Agibot’s GO-1 adopts a ViLLA hybrid architecture with a primary focus on mining massive video data. Targeting the industry pain point of low utilization of heterogeneous data, it introduces implicit action tokenization technology, abstracting complex actions into standardized “action vocabularies” to bridge the gap between visual understanding and physical execution — enabling robots to “learn skills by watching videos.” The model supports whole-body control, and following technical optimization, the average success rate across multi-complexity tasks has risen from 46% to 78%, with notable advantages in few-shot generalization capability.

ByteDance’s GR-2 is positioned as a generative video-language-action model focused on general-purpose cognitive capability. The model first completes video generation pre-training using 38 million internet video clips, learning physical laws and temporal logic, then fine-tunes with 5,000 real trajectory data points, substantially reducing training costs. In testing, the model’s average success rate across 105 multi-task scenarios reached 97.7%, with an 87% success rate still maintained in completely unfamiliar new environments — placing it at the forefront of locally developed models in generalization performance. Overall, China’s local manufacturers all take low-cost data strategies as their entry point, focusing deeply on segmented scenarios and steadily advancing the commercialization of their technology.


AI Chips for Humanoid Robots: Upgrading the Computing Foundation, Edge-Side Deployment Becomes the Core Trend

The real-world deployment of large models cannot proceed without computing power support. Humanoid robots’ stringent requirements for low latency, low power consumption, and high reliability are driving AI chips to transition from general-purpose computing toward robot-specific edge-side chips, making this a new competitive focal point in the industry.

Evolution of Chip Technical Architecture

Traditional cloud-side chips carry high power consumption and high latency, making them difficult to adapt to robot operating scenarios. The heterogeneous architecture of SoC plus independent NPU has now become the industry mainstream. The SoC, integrating multiple types of processing units, assumes the primary control function, realizing integrated perception, decision-making, and control, with the combined advantages of small form factor, low power consumption, and high real-time performance. The NPU specifically handles AI inference tasks such as VLA large models, visual recognition, and motion planning. At the same time, processing-in-memory technology is progressively being deployed, overcoming the shortcomings of the separated processing-and-memory paradigm and capable of reducing edge-side inference power consumption by more than 50%. Combined with software optimization means such as model distillation and lightweight compilation, software and hardware are deeply coordinated — for example, Google’s Gemini Control Hub can achieve local response within 10 milliseconds, fully meeting robots’ real-time control requirements.

Global Chip Competitive Landscape

Overseas vendors firmly control the high-end market. Google’s proprietary Gemini Control Hub builds a full-stack “model + chip” closed loop, serving the high-end Atlas humanoid robot. NVIDIA’s Jetson series offers strong computing power and a mature ecosystem; the next-generation Jetson Thor is optimized for physical AI and has become the computing platform of first choice for high-end general-purpose robot models.

Chinese chip enterprises are accelerating their catch-up. Rockchip (瑞芯微), Horizon Robotics (地平線), and Chipways Technology (芯驰科技), among others, have launched robot-specific chips, leveraging mature automotive-grade processes to balance reliability and cost. China’s domestic solutions generally adopt a “large brain plus small brain collaborative” architecture: the main-control SoC handles high-level decision-making, while the co-processor manages joint movement control. The focus is on industrial and service-oriented mid-to-low-end robots. Currently, the local chip market penetration rate in China has reached 32%, with domestic substitution advancing steadily.

Core Industry Development Trends

The first trend is that edge-cloud collaboration has become the standard architecture: edge-side deployment ensures low-latency control, while the cloud handles model training and iteration, balancing performance and efficiency. The second is that chips are moving toward specialization and ecosystem integration: general-purpose chips’ cost disadvantages are becoming increasingly apparent, specialized SoCs will become standard configurations, and software ecosystems and toolchains are progressively becoming core competitiveness. The third is that automotive-grade standards continue to penetrate the industry: their characteristics of high reliability and strong anti-interference are being widely adopted, and they have also become a key pathway for reducing chip costs and improving mass-production capability.

[Disclaimer]: The above content reflects analysis of publicly available information, expert insights, and BCC research. It does not constitute investment advice. BCC is not responsible for any losses resulting from reliance on the views expressed herein. Investors should exercise caution.