阿里巴巴发布旨在驱动下一代机器人的AI大脑
China's Alibaba Unveils AI Brains Designed To Power The Next Generation Of Robots

原始链接: https://www.zerohedge.com/technology/chinas-alibaba-unveils-ai-brains-designed-power-next-generation-robots

阿里巴巴发布了“通义机器人”(Qwen-Robot)系列,这是一套旨在连接大语言模型与物理机器人动作的具身智能模型。该系列由通义实验室开发,包含三个专用模型:用于导航的 **Qwen-RobotNav**、用于物体物理操作的 **Qwen-RobotManip**,以及用于环境建模与预测的 **Qwen-RobotWorld**。 通过整合这些模型,机器人能够解读复杂的视觉和语言指令,从而执行现实世界中的任务,例如在陌生环境中导航或操纵物体。在测试中,阿里巴巴展示了一款四足机器人,它无需预装地图即可在家庭环境中成功导航;其操控模型在 RoboChallenge 基准测试中也获得了高分。 此次发布标志着阿里巴巴向具身智能领域的战略转型,该领域目前正面临来自谷歌 DeepMind、英伟达及各类机器人初创公司的激烈全球竞争。通过将中国的制造业优势与先进的 AI 软件相结合,阿里巴巴旨在使机器能够有效地在物理世界中进行推理、感知和交互。该公司还开源了“Chat2Robot”,以促进此类具身智能交互的进一步开发与测试。

相关文章

原文

Authored by Jijo Malayil via Interesting Engineering,

Chinese firm Alibaba has launched its first embodied AI model family, which links large language models with real-world robotic actions.

The Qwen-Robot suite includes three distinct models, each targeting a different layer of physical intelligence.Unitree/YouTube

The Qwen-Robot suite was developed by Alibaba's Tongyi Lab and is undergoing pilot testing with selected Alibaba Cloud enterprise clients.

The suite comprises three models focused on navigation, manipulation, and world modeling for robots operating in physical environments.

Alibaba said the models enable machines to perceive, reason, and interact with the real world, joining a growing global push to advance embodied AI beyond traditional chatbot applications.

Robots meet reasoning

Alibaba says its Qwen family of AI models has become very good at understanding the physical world. These models can recognize objects, understand spatial relationships, follow complex visual instructions, and reason about real-world environments. For example, a model can understand a command such as, "Go to the kitchen, find the red cup, pick it up, and place it on the shelf."

However, understanding a task is different from actually performing it. While a vision-language model (VLM) can describe the steps needed to complete a task, it cannot directly control a robot's movements.

The challenge is connecting human language and visual understanding with the motor actions required to interact with the physical world.

This problem is difficult because robot training data is very different from internet data. Information collected from navigation systems, robotic arms, vehicles, and cameras comes in different formats and is expensive to gather. Simply combining all this data often creates conflicts rather than improving performance.

To address this, Alibaba developed the Qwen-Robot Suite, which includes three specialized models. Qwen-RobotNav focuses on movement and navigation. It helps robots follow instructions, navigate to locations, track targets, and support autonomous driving.

According to its website, Qwen-RobotManip focuses on physical interaction. It enables robots to grasp, move, and manipulate objects using a large training dataset collected from different robotic systems. Qwen-RobotWorld acts as a world model, predicting how environments may change and helping robots understand the likely outcomes of their actions.

Together, these models aim to enable robots to understand instructions, interact with objects, navigate environments, and make decisions in the real world.

Physical AI accelerates

Alibaba showcased Qwen-RobotNav on a Unitree Go2 quadruped powered by NVIDIA Jetson Thor hardware and a single low-resolution camera. The robot successfully navigated an unfamiliar apartment, following spoken instructions across multiple rooms without preloaded maps, while maintaining an inference latency of 196 milliseconds.

The company claims that Qwen-RobotManip, its robotic manipulation model, was trained on more than 38,000 hours of open-source data covering object handling and interaction tasks. According to Alibaba, the model recently achieved the highest score in the generalist category of the RoboChallenge real-world robotics benchmark, earning a process score of 59.83 and a task success rate of 45 percent.

The company also unveiled Qwen-RobotClaw, a robotics agent framework that enables Qwen models to use the Qwen-Robot suite as physical-world tools. In one demonstration, an agent searched for a restroom, identified an out-of-order sign, and independently rerouted to another location. Alibaba further open-sourced Chat2Robot, a browser-based platform for testing embodied AI interactions.

As competition in embodied AI intensifies worldwide, Alibaba has expanded its ambitions beyond language and multimodal software with the launch of its Qwen-Robot models. The move reflects a broader industry shift toward creating AI systems capable of understanding and interacting with the physical world.

Alibaba's move comes as competition in physical AI accelerates globally. In the US, Google DeepMind is advancing Gemini Robotics, while Nvidia is expanding its robotics ecosystem through Cosmos, Isaac, and GR00T. Start-ups, including Physical Intelligence, Skild AI, and Figure AI, are also developing general-purpose robotic intelligence, according to the South China Morning Post.

China is strengthening its position by pairing its manufacturing advantages with growing investments in AI software for autonomous decision-making. The sector now spans AI developers, robotics firms, and EV makers. Companies such as Alibaba, Tencent, Unitree, AgiBot, UBTech, Galbot, Spirit AI, GigaAI, Xpeng, and Xiaomi are actively pursuing embodied AI technologies.

联系我们 contact @ memedata.com