Google's Logan Kilpatrick: Spend a Quarter of Time on Benchmarks for AI Products
Logan Kilpatrick stated that when using artificial intelligence to create products, more than 25% of the time should be spent on benchmarking, and efforts should be made to make model labs care about these benchmarks, as this is the easiest path for companies to accelerate progress.
He currently works at Google, responsible for Gemini, Google AI Studio, and interface-related tasks, previously at OpenAI. Months ago, he publicly stated that every company using models to create products should build their own evaluation sets, allowing model iterations to be "disproportionately beneficial to themselves," and pointed out companies like Zapier and Sierra that are already doing this. Public rankings measure general capabilities, but what contract review companies really need is accurate clause extraction, which cannot be measured by rankings.
Labs will adjust models to enter evaluation sets. Those who can write business tasks as reproducible and comparable problems will have the opportunity to be optimized in the next release. The public benchmark alpha, he said, is still significant because good problems are scarce, and labs lack pressure testing in vertical scenarios.
Buyers aim to reclaim evaluation rights from general rankings for their specific application scenarios; sellers continue to focus only on teams selecting models based on Arena and MMLU. The event-driven aspect is that insiders in labs have fixed the time allocation ratio. Beneficiaries are product teams that can provide private evaluation feedback to Gemini, OpenAI, and Anthropic; those under pressure are shell products that are only tied to the public number one.
The market trades in the definition rights of evaluations, not just another slogan about overtime.
Source: Public Information
ABAB AI Insight
Kilpatrick has seen how application companies select models and how labs schedule from the model distribution side. The 25% he mentioned is not just engineering motivation; it acknowledges that the misalignment between general benchmarks and vertical products has become significant enough to require separate projects. Labs optimize their objective functions according to publishable and comparable evaluations; if application companies cannot provide problems, they can only wait for the next version to coincidentally improve. Companies like Zapier and customer service AI firms building their own evaluations translate "our workflows" into scores that labs can run.
The capital path is to shift some product personnel from feature stacking to problem formulation, labeling, leak prevention, and lobbying. The cost is dataset maintenance; the benefit is automatically gaining targeted improvements during each base upgrade and using their own scores to negotiate prices instead of relying on public rankings. Whoever can get labs to include a specific evaluation in internal regressions will lock in the direction for the next generation of models.
Comparative objects include ImageNet reshaping vision, GLUE reshaping language, and LMSYS Arena reshaping conversational experience. If vertical benchmarks can become industry factual standards, they will be close to Bloomberg's position on financial NLP and SAE's position on automotive safety. The current stage is shifting the application layer from "selecting the strongest model" to "defining what is strong."
Structural judgment belongs to the transfer of pricing power. The pricing power of model quality shifts from public rankings to scenario-specific problems. The mechanism is that lab training and releases are driven by quantifiable scores; whoever formulates the problems dictates where the computational power will be used in the next round. The 25% time is an entry fee, exchanged for having others' R&D budgets work for your product.
ABAB News · Cognitive Laws
- Whoever formulates the problems dictates where the next round of models will optimize.
- Public rankings measure general capabilities, with revenue coming from the segments that cannot be measured.
- Making labs care about your benchmarks is cheaper than pushing them to produce larger models.