Gemini Robotics 2 为机器人带来全身智能
Gemini Robotics 2 brings whole body intelligence to robots

原始链接: https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/

Google 推出了 **Gemini Robotics 2**,这是一项机器人智能领域的重大进步,旨在从重复的预编程任务转向真正适应性的自主行为。 这一新一代平台利用三种专用模型,使机器人能够在不可预测的环境中“思考、行动和交互”: * **Gemini Robotics 2**:一种先进的视觉-语言-动作 (VLA) 模型,提供全身控制,涵盖从复杂的人形运动到精细、灵活的物体操控。 * **Gemini Robotics ER 2**:一种具身推理模型,允许机器人与人类交流、理解周围环境并规划复杂的多步骤任务。它还引入了多个机器人团队协作的能力。 * **Gemini Robotics On-Device 2**:VLA 模型的高效本地版本,使机器人只需几小时的训练即可适应新的物理形态。 通过整合这些模型,Google 使机器人能够以空前的灵活性和独立性执行复杂的现实任务,例如打扫房间或进行协作工作。这一演进标志着机器人向能够无缝适应我们世界的方向迈出了重大一步,无论其具体外形或设计如何。

Google DeepMind 宣布推出旨在为机器人赋予“全身智能”的“Gemini Robotics 2”,这在 Hacker News 上引发了激烈讨论。 乐观者将机器人技术的现状比作大语言模型(LLM)早期并不起眼的阶段,认为技术的快速进步可能在几年内使家用机器人成为现实。他们认为,即使是行动缓慢的机器人,只要能可靠地处理重复性任务,也具有实用价值。然而,怀疑论者强调了硬件方面的巨大局限性,指出灵活性和物理耐用性比纯软件 AI 难解决得多。 讨论的重点包括: * **经济障碍**:虽然机器人可能会迅速进入工业环境(替代多班制劳动力),但高昂的成本使得它们在短期内难以大规模进入家庭。 * **安全性与可靠性**:该领域的专业人士警告称,目前的人形机器人模型难以应对现实世界的模糊性,且存在重大的安全风险。 * **形态因素**:一些人认为专用自动化设备(如改进型洗碗机)优于人形机器人;另一些人则反驳称,由于我们的物理世界是专门为人体设计的,人形结构是必要的。 总体而言,社区认为这项技术前景广阔,但对于其是否已准备好投入实际的日常使用仍持谨慎态度。
相关文章

原文

From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks

For decades, we’ve dreamed of robots that can seamlessly step into our world and lend a hand. Now, that vision takes a significant stride forward.

Most robots are pre-programmed or teleoperated for narrow, repetitive task sequences. They lack the ability to truly learn for themselves or adapt to unpredictable environments. Moreover, transferring learned skills from one robot body to another remains incredibly difficult. To take on the hardest problems at scale, robots of every shape and size need AI models giving them the ability to think, act, and interact intelligently to safely complete tasks.

We demonstrated how Gemini's multimodal understanding could drive real-world action with Gemini Robotics. Today, we are introducing Gemini Robotics 2 - the intelligence layer powering the next generation of truly adaptable robots. As it takes its first literal steps, this major advance unlocks intelligent whole-body control, advanced dexterity, and multi-robot collaboration.

Gemini Robotics 2 enables robots to reason through every movement, unlocking a broad range of tasks. For example, it can enable a humanoid to walk, crouch, stretch, and manipulate objects to clean up a cluttered room. It can even team up with other robots to finish the job faster. And this profound intelligence can also run locally on-device while seamlessly adapting to entirely new robotic bodies in just a few hours.

We are making this possible through three highly capable models:

  • Gemini Robotics 2: Our most advanced vision-language-action model (VLA) that converts vision and language input into motor control, enabling a robot to take action. This model is capable of controlling full humanoids, from feet to fingertips, and other bi-arm robots. It also brings a new level of dexterous manipulation on both hands and grippers.
  • Gemini Robotics ER 2: Our most capable embodied reasoning (ER) model. It is a vision language model (VLM) that acts as our agent, enabling robots to communicate with humans, understand the physical world and plan multi-step tasks lasting several minutes. We are also introducing the ability for robots to work together as a team.
  • Gemini Robotics On-Device 2: Our most efficient vision-language-action model (VLA) optimized to run locally on robotic devices. This model can now achieve fast adaptation to completely new robot embodiments with a few hours of data.
联系我们 contact @ memedata.com