GPT-6 Astra 在机械臂上的应用
GPT-6 Astra on robot arms

原始链接: https://openai.robocurve.org/gpt-6-astra/

Robocurve 最近的一项研究评估了 OpenAI 的 **GPT-6 Astra** 在 YAM 机械臂上的表现,并将其与 Anthropic 的 Claude Fable 5 和 5.1 模型在两项任务中进行了对比:将积木放入碗中以及将拼图块插入凹槽。 在“积木入碗”任务中,GPT-6 Astra 表现远超竞争对手,成功率达到 95%(19/20),而 Fable 5.1 的成功率仅为 40%(8/20)。Astra 的效率也更高,平均每次试验耗时 2.5 分钟,成本为 0.94 美元,比 Fable 5.1 便宜约 2.3 倍。 然而,拼图任务对所有模型来说都极具挑战性。Astra 和 Fable 5.1 均表现不佳,成功率仅为 10%(2/20)。尽管 Astra 的成本效益有所提升(1.36 美元对比 2.18 美元),但两个模型都在最后的插入步骤卡住了。 虽然 Astra 在简单的操作任务中表现出更快的速度和可靠性,但结果表明,复杂的、高精度的组装对于当前基于大语言模型的机器人控制策略而言,仍然是一个重大的障碍。

Hacker News 上的一场讨论探讨了多模态人工智能(如 Astra/GPT-6)与机器人技术的融合,这一话题由“RoboCurve”项目引发。 参与者就大语言模型(LLM)是否是解决自动驾驶或折叠衣物等复杂物理任务的关键展开了辩论。尽管一些用户因纯 LLM 缺乏“视觉”能力或现有机器人技术的机械脆弱性而持怀疑态度,但另一些人则强调了近期的进展。Sunday Robotics 在衣物折叠方面的准确性以及 Figure 在操作演示上的表现,都表明“物理大模型”正在飞速进化。 讨论达成的一致意见是,主要的瓶颈正在发生转移:随着 AI 智能的成熟,挑战正转向硬件可靠性以及机器人系统在现实环境中的耐用性。用户表示乐观,认为正如汽车发展的早期历史一样,随着更多训练数据被应用于物理世界,这些机器人领域“枯燥”但变革性的突破正成为不可避免的现实。
相关文章

原文
GPT-6 Astra on robot arms | Robocurve

September 4, 2026

A follow‑up to our comparison of Claude Fable 5 and Fable 5.1. We gave OpenAI's GPT‑6 Astra control of the same YAM arms under the same Inspect Robots agent policy, on the same two tasks:

“Pick up the red block from the table and place it inside the bowl.”

“Pick up the round blue puzzle piece by the knob at its center and place it into the matching circular groove in the board.”

On the bowl task Astra placed the block in 19 of 20 trials, against Fable 5.1's 8 of 20 and Fable 5 in 1 of 20, in 2.5 minutes per trial to Fable 5.1's 6.8, at an estimated $0.94 per run to $2.12.

The puzzle task is a different story: Astra completed the insertion 2 times in 20 against Fable 5.1's 2 in 20. It reaches the groove and stalls at the same final step Fable does, at $1.36 per run to $2.18.

联系我们 contact @ memedata.com