Back to news

Ethereum Co-founder Vitalik: AI Safety Should Learn from Governance Mechanisms

Ethereum co-founder Vitalik Buterin posted on X platform on September 13, suggesting that a significant application direction for adversarial governance mechanism design theory is likely AI safety. He believes that the long-standing "principal-agent" dilemma in governance systems has a structural "duality" relationship with the core challenges faced by AI safety.

He specifically compared two types of scenarios: one is a common situation in governance systems—relatively less savvy principals (usually static algorithms or a set of rigid rules) trying to manage more savvy human agents who can exploit system loopholes; the other is in the AI safety field—humans plus weaker LLMs acting as principals trying to supervise a much stronger LLM as an agent. He pointed out that both scenarios essentially involve "weaker principals trying to achieve ideal outcomes from stronger agents."

Vitalik emphasized that limiting the ability of agents to collude often leads to significantly better outcomes in governance systems, a judgment that extends his existing discussions on coordination issues from 2020. He directly applied this conclusion to the AI safety field, noting that the real "nightmare scenario" is not a single "malicious" model, but multiple powerful models coordinating actions in ways that are difficult for human supervisors to detect or understand.

In terms of specific mechanisms, he named several "anti-collusion" tools previously validated in governance, including quadratic voting, commit-reveal schemes, and identity verification layers, suggesting that these mechanisms could also be transplanted into the security design of AI systems.

He further proposed that future constraints on AI capabilities may resemble a complete institutional system that includes rules, permissions, adjudication, and record mechanisms, rather than merely setting up technical "sandbox" isolation. This idea is consistent with his earlier suggestion in February of this year about introducing AI stewards in DAO governance as an intermediary layer between token holders and complex decision-making.

Source: Public Information

ABAB AI Insight

Vitalik has long experimented with various anti-collusion mechanisms in the Ethereum governance system—he previously led the implementation of quadratic voting and quadratic funding in public goods funding scenarios like Gitcoin, and has proposed soulbound tokens for establishing non-transferable identity and reputation systems. These historical practices are direct sources for his current analogy to AI safety, rather than temporary concepts.

Unlike traditional AI safety research, which relies more on alignment techniques during the training phase (such as reinforcement learning from human feedback), Vitalik's proposed path essentially involves directly transferring the resource and experience of mechanism design invested in governance token distribution and public goods funding in the Ethereum ecosystem to the new but structurally similar issue of AI safety. This "resource transfer" does not involve financial investment but rather the reuse of knowledge and mechanism templates, motivated by the desire to avoid reinventing governance tools from scratch in the AI safety field.

This approach of "using governance mechanism design to solve AI safety" complements rather than replaces the mainstream paths of "alignment research" and "explainability research" in traditional AI safety academia. It is more akin to horizontally transferring the practical experience accumulated over more than a decade in blockchain governance on "how to prevent stakeholders from colluding to do harm" to a new and higher-risk field. Industry-wise, this suggestion emerged in the same week that CEOs from leading AI labs like Anthropic and OpenAI collectively called for "slowing down development and introducing third-party evaluations," providing a technical path from the crypto governance field that has been less incorporated into mainstream AI safety discussions.

Essentially, this represents a technical substitution at the level of regulatory toolboxes—if Vitalik's analogy holds, it implies that the mechanism design to address the risks of AI runaway may not necessarily rely on direct limitations on the capabilities of AI models (such as slowing down training or limiting computing power), but can, like blockchain governance, constrain the coordinated behavior of multiple AI agents through external institutional designs (permission layering, verifiable decision records, anti-collusion voting mechanisms). Mechanically, this redefines the AI safety issue from "how to make a single model safer" to "how to design an institutional environment that prevents multiple powerful agents from colluding to do harm," which presents two different levels of response strategies compared to the "slow down" approach advocated by Amodei and others—one is subtraction (slowing down capabilities), while the other is addition (adding constraint mechanisms).

ABAB News · Cognitive Laws

  1. If you can manage collusion, you can manage the system.
  2. Governance mechanisms are not limited to on-chain or off-chain; if they can prevent human cheating, they can prevent AI cheating.
  3. Constraining AI relies not on caging it, but on building a complete set of institutions.

Source

·ABAB News
·
5 min read
·4 hrs ago
分享: