Flash News

Google DeepMind Announces Upgrade to Gemini Robotics-ER Version 1.6, Focusing on Multi-Perspective Reasoning for Physical World Understanding

Google DeepMind has announced the launch of the Gemini Robotics-ER version 1.6 upgrade, focusing on enabling robots to understand the physical world through multi-perspective reasoning. This model features enhanced visual and spatial perception capabilities, allowing it to identify target objects in cluttered environments, assess task completion, and autonomously decide whether to retry or switch processes. The system can integrate multiple camera feeds in real-time to construct a complete scene understanding, achieving "task-level decision-making."

DeepMind states that the new version strengthens industrial inspection and laboratory operation capabilities: it can correct common lens distortions, automatically read mechanical dial scales, and generate calculation code to calibrate measurement results. The upgrade also focuses on improving safety—automatically avoiding liquids or overweight objects during handling tasks and increasing damage risk detection accuracy by 10% compared to previous versions.

This version is viewed by several English tech media outlets as a key milestone for a "general robotic perception system," integrating visual understanding, world knowledge, and active reasoning into a single framework.

Source: Public Information

ABAB AI Insight

Gemini Robotics‑ER 1.6的重要性不在于单一性能提升,而在于它代表AI系统首次接近真正的“物理常识”。传统语言模型依赖文本逻辑,缺乏对空间、重量、材质的连贯理解,而ER 1.6通过多模态融合建立了“世界连续性”——AI开始理解自己行动在真实空间中的因果链条。

这意味着AI从“语言代理”向“物理代理”转变。它不再只是计算文字,而是参与现实操作系统;这对制造、仓储、能源巡检等领域是生产关系级的变革。AI不再是管理接口,而成为执行层自治力量。

更深层的趋势在于平台化——当感知模型与物理执行体系统一后,AI对现实世界的解释权与操作权将逐步收敛。Gemini‑ER象征的是一种工业权力重构:从“算法服务人”向“算法理解世界”跨了一步。谁能控制这类具备世界模型的AI,谁就控制了下一轮自动化文明的底层坐标系。

Google

Source

·ABAB News
·
3 min read
·120d ago
分享: