Back to news

Meta's Chief AI Officer Alexandr Wang: A Group of Agents Can Easily Outperform 100 Engineers

Meta's Chief AI Officer Alexandr Wang stated during Y Combinator Startup School 2026 that there are internal cases where, if the agent loops, evaluation, and optimization metrics are correct, a group of agents can easily outperform 100 engineers.

He emphasized that it's not about smarter models. The loop includes micro-feedback loops, and agents must continuously operate and self-optimize within these loops. He identified three conditions: the agent loop must be correct, the evaluation must be correct, and the metrics must be correct. Missing any one of these means that the open window is just guessing together. He also mentioned that using continuous feedback loops to produce results still has significant room to spend 1,000 times or even 1 million times more tokens.

No specific internal cases were disclosed, including how many agents were used, how long they ran, or how the human baseline was calculated. He limited his statement to having "seen some cases," clarifying that not all engineering positions in the company have been replaced. Reports described this infrastructure as persistent memory stored in markdown files, scheduled with cron, which is a timed task running on Unix since the 1970s. Tasks are broken down into discrete loops, where agents review outputs, correct errors, and run again, without relying on a single model to handle everything.

Wang is the founder of Scale AI. Meta acquired about 49% of Scale for $14.3 billion in 2025, placing him in Meta's Superintelligence Lab as Chief AI Officer. Scale's original business involved providing data and evaluations for models. He transferred the same logic to agents at YC: without repeatable evaluations, neither numbers nor brainpower will automatically lead to completion.

The premise that 100 engineers could be outperformed by a swarm is based on the definition of completion. If handoffs between people cannot be written as metrics that machines can execute infinitely, even supervisory agents cannot compensate. Skill descriptions, markdown, and scheduled tasks must align with sufficiently dense real data, rather than just a one-time demonstration. If metrics are not clearly defined, the larger the swarm, the louder the guessing.

Buyers are companies that pay for agent hours based on results, while sellers are engineering teams that still charge by headcount and project duration. This is an expectation driven by internal case studies, not the layoff numbers already disclosed by Meta. Money is shifting from engineers' monthly salaries to recyclable tokens and evaluations. Teams that can define completion as automatically scored benefit, while positions that can only claim completion verbally and cannot be rerun by machines are under pressure.

Garry Tan later summarized the same approach in one sentence: you can write markdown skills with agents and then put them on cron, covering almost all useful knowledge work. Wang did not provide the cost of those 100 people or the billing for the swarm during this conversation.

Source: Public Information

ABAB AI Insight

Alexandr Wang founded Scale AI at the age of 19, focusing on the unseen layer of models: labeling, evaluation, and repeatable data pipelines. Meta spent $14.3 billion to acquire about 49% in 2025, placing him in the Superintelligence Lab. The path remains unchanged. Scale collects human judgment to form training sets, while Meta is now collecting human handoffs to create loops that agents can score themselves. In both cases, metrics come first, then models.

Money is now flowing towards tokens and scheduled tasks, not towards another supervisory model. He stated on-site that continuous feedback loops are worth spending 1,000 to 1 million times more tokens. Markdown stores memory, cron wakes up the loops, and evaluations determine whether a round is considered complete. This is the same accounting as Scale's collection of verifiable data from annotators: without acceptance, increasing manpower only increases divergence. The 100 engineers at Meta were compared to see if they could be rerun overnight, not based on qualifications.

This is akin to factories changing quality control from visual inspections by experienced workers to using measuring tools for sampling. Toyota's rhythm relies not on a smarter foreman, but on each workstation having standards that can stop the line. The software industry is still in the expansion phase of agent orchestration: YC presented methods as documents and scheduled tasks, and Garry Tan later suggested that knowledge work could be structured similarly. Until task names and billing are publicly disclosed, it remains a case study, not an organizational chart.

The essence is the transfer of pricing power. Engineering teams quote based on headcount and duration, while swarms quote based on whether each round of evaluation passes. The mechanism is that once the definition of completion can be executed infinitely by machines, headcount is no longer a measure of capacity. Missing any of the loops, evaluations, or metrics means that opening more windows only leads to more noise, and the more tokens spent, the more expensive the noise.

ABAB News · Cognitive Laws

  1. Numbers cannot outperform swarms if the definition of completion is not clearly written.
  2. If the loops, evaluations, or metrics are incorrect, opening more windows only leads to more noise.
  3. No matter how smart the supervisor is, it cannot replace a standard that can be rerun overnight.

Source

·ABAB News
·
7 min read
·6 hrs ago
分享: