a16z Closed-Door Debate: From Existential Risk (X-Risk) Regulatory Traps, Internal Network Agent Protection to New Paradigms in Probabilistic Programming
a16z
a16z
Original Statement
1. Core Controversy: "Pacing" or Regulatory Capture?
• The open letter from Anthropic founder Dario Amodei has sparked intense debate in the industry:
• The frontier large model safety and sandbox isolation initiative proposed by Dario and others is reasonable at the pure technical engineering level;
• However, the focus of the controversy lies in the narrative packaging of "pacing" and the so-called "existential risk (X-Risk)".
• The logical flaw of the concept of "pacing":
• Lack of a reference frame: There has never been a publicly agreed benchmark speed for technological evolution. Announcing "we have decided to slow down" when no one knows the original completion node is essentially similar to the media claiming "the unannounced Apple car has been delayed," which lacks measurable standards for a self-consistent narrative.
• A middle ground that pleases neither side: Attempting to walk a compromise between the internal extreme pause faction (Pause/Doomer) and external regulatory bodies has resulted in radicals accusing it of "just slowing down instead of stopping," while regulators deem it as "acknowledging the existence of harm yet still racing ahead."
• Reflecting on the moral coercion of "existential risk (X-Risk)":
• Historically, nuclear weapons development teams were well aware of their absolute lethality on a physical level, thus establishing the highest level of national control mechanisms;
• If the heads of large model laboratories genuinely believe there is a 10% probability of human extinction, the only ethical response would be a complete halt or total nationalization;
• The reality is contradictory: If one side portrays extinction risk while continuing to accelerate financing and pushing into the commercial market, it can easily be exploited by politicians, leading to excessive regulatory capture that could strangle startups and the open-source ecosystem.
2. Historical Reflection: From Nuclear Bombs, Early Internet Viruses to Regulatory Lag
• Prerequisites of the "Fact Pattern" in policy-making:
• The automotive industry was already widespread in the early 20th century, and it wasn't until Ralph Nader published "Unsafe at Any Speed" in the 1960s that stringent safety regulations were formed;
• The aviation industry experienced a 40-year period of technical trial and error from the Wright brothers' first flight to the establishment of a complete FAA airworthiness certification system;
• Traditional regulation must be based on specific damages that have occurred and a clear causal chain; attempting to implement "predictive legislation" before technology has matured is almost certain to stifle innovation.
• If we were to apply today's fear paradigm to the Internet of the 1990s, it would never have emerged:
• In the era of Windows 95 and early PCs, connected computers faced widespread risks from worm viruses and network interruptions (for example, the Morris worm paralyzed 10% of backbone networks, leading to frequent disconnections and business losses);
• At that time, Congress passed the Computer Fraud and Abuse Act (CFAA) targeting specific system intrusion cases; if the Internet's underlying infrastructure had been completely locked down in 1994 out of fear of hackers, the modern digital economy would not exist.
• The potential spread risk of the European GDPR model:
• Europe, lacking local underlying tech giants, tends to legislate aggressively in compliance and antitrust areas;
• Beware of Europe refining regulation into a "digital airbag reminder" for AI Agents—where every time an Agent calls an external API or performs file read/write, a GDPR-like disclaimer confirmation box pops up, ultimately leading to widespread user cognitive numbness.
3. The Real New Security Front: Engineering Threats from the Emergence of Internal Network Agents
• The complete failure of traditional internal network security assumptions:
• Past IT and enterprise information security were based on an implicit assumption: 95%-99% of internal employees comply with regulations most of the time, with malicious insiders being a very low probability event.
• Internal GitHub, Slack, approval flows, and financial reimbursement systems often lack strict concurrency throttling, relying solely on single sign-on (SSO) and coarse-grained identity authentication.
• Agent swarms as "internal network distributed denial-of-service attacks (DDoS)":
• When internal employees start to batch schedule thousands of autonomous Agents to write code, call APIs, and run automated tests, the software behavior is structurally identical to high-frequency DDoS attacks.
• Agents possess rapid retry and tireless characteristics, easily exhausting internal microservice resources in a dead loop, and may even trigger dangerous data overwrites due to misunderstanding instructions.
• Next-generation operating systems and permission granularity innovation:
• The existing operating system permission system is too crude (either fully open or frequently prompting for confirmation);
• There is an urgent need to reconstruct the underlying security stack to support fine-grained dynamic permission isolation (e.g., instantaneous read/write control for specific folders, adaptive API rate monitoring, and audit tracking), shifting the security focus from the metaphysical level of species survival back to concrete engineering defenses.
4. Paradigm Shift: The Rise of the Jev Model and the Return of Probabilistic Programming
• The essential conflict between natural language interaction and traditional software systems:
• The industry has previously focused on a "text input, text output" chat model, but forcing structured software to parse unstructured long text is extremely costly and error-prone;
• Relying on complex prompt constraints and schema validation cannot guarantee 100% deterministic output and wastes enormous computational costs and response delays on meaningless question-and-answer redundancy.
• The core breakthrough of the Jev model: from "generating text" to "semantic decision-making":
• The input remains semantic context, but the output completely abandons generative text;
• Among a given set of discrete options or routing conditions, it directly returns the optimal decision and precise probability distribution with extremely high throughput and low latency;
• This completely overturns inefficient chat box interactions, allowing traditional software code to directly use large model outputs as the basis for logical branching decisions.
• The revival of half a century of computer science legacy:
• The core proposition of programming languages in the early 1960s-1970s was simulation and probabilistic modeling (e.g., handling ballistic trajectories, wind fluctuations, and other non-deterministic physical systems);
• Modern software code has long been limited to Boolean algebra (absolute certainty of if/else);
• With the implementation of Jev-like architectures, the core programming paradigm is shifting entirely to condition branches based on probabilistic confidence (if x% then ...), seamlessly integrating probabilistic language models with classical deterministic software engineering.
5. The Shift of Innovation Focus: From the Base Model Layer to the Model Periphery
• The "critical mass dilemma" at the giant platform layer:
• Leading large model laboratories are currently mired in maintaining large infrastructure operations, resolving alignment disputes, and managing massive user compatibility, severely dispersing their focus, making it difficult to optimize all vertical scenarios from the platform layer.
• The autonomous evolution of application layers and peripheral tools:
• The history of the software industry shows that core innovations often emerge outside the platform (analogous to the "Sherlocking" process where third-party independent tools are integrated into systems);
• Large models are no longer merely revered as ultimate deities but are evolving into standard underlying components like databases and compilers; the core technological moat of future software is shifting entirely to how to reconstruct contextual systems, state persistence, and high-reliability business flows around the model's periphery.
ABAB AI Insight
a16z's latest debate: The next war in AI is not about large models, but about regulation, safety, and the reconstruction of software fundamentals.
For the past three years, the entire AI industry has been dominated by one question:
Who can train smarter large models?
OpenAI.
Anthropic.
Google.
Meta.
xAI.
Everyone's attention has been focused on:
Parameters;
Benchmarks;
Training compute;
Reasoning;
AGI.
But by the second half of 2026, a change that is becoming increasingly evident is occurring:
The biggest bottleneck in AI is shifting from "Is the model smart enough?" to "Can the software world truly support this intelligence?"
This is also the common issue pointed out in several recent discussions by a16z.
Today's large models can already:
Write code;
Conduct research;
Analyze legal documents;
Automatically complete complex tasks;
Control tools.
But have enterprises really been automated?
Many times:
No.
Diogo Almeida even directly asked at the latest a16z discussion:
"Where the fuck is all the automation?"
AI is already so smart,
Why is real-world software still so dumb?
This may be the most important question to understand the next phase of AI.
────────────────
1. The AI industry is transitioning from a "capability revolution" to a "system revolution"
The main issues addressed in the past three years have been:
Can the model do it?
Can the model write code?
Can it reason?
Can it see images?
Can it call tools?
The answers are increasingly becoming:
Yes.
The next phase of questions becomes:
Can the system reliably let the model do it?
Are enterprises willing to let it:
Access databases;
Change code;
Call payments;
Modify production systems;
Read employee files;
Control infrastructure?
This is a completely different question.
It no longer primarily belongs to:
Model Science.
But increasingly belongs to:
Software Engineering.
────────────────
2. The stronger the model, the more thoroughly the problems of traditional software infrastructure are exposed
A person operates:
Dozens of internal APIs daily.
No problem.
An Agent:
May call:
Thousands of times in a minute.
An employee simultaneously opens:
20 Agents.
The company has:
10,000 employees.
What happens to the original enterprise backend?
The answer may be:
It gets directly overwhelmed.
a16z even used a very direct description in its 2026 infrastructure predictions:
When traditional systems see the behavior of Agents,
They will think they are:
Undergoing a DDoS attack.
This is not an exaggeration.
Because from a network perspective:
It is indeed very similar.
────────────────
3. The biggest difference between human software and Agent software is the difference in "speed scale"
Traditional enterprise systems assume:
A person clicks a button.
The system responds.
Then clicks the next one.
This is called:
Human-speed Workload.
Agents are different.
One goal:
"Upgrade the entire codebase."
May instantly unfold:
5,000 sub-tasks;
Thousands of database queries;
A large number of API calls;
Parallel tool execution.
This is called:
Machine-speed Workload.
Software infrastructure has never been designed for this kind of behavior.
────────────────
4. This means the AI revolution is truly beginning to challenge the underlying assumptions of computer systems
In the past, enterprise IT had an implicit premise:
Users are human.
Humans:
Are slow.
Easily fatigued.
Do not click buttons ten thousand times per second.
So many systems do not even have:
Particularly strict rate limits.
Today:
Usernames are still:
"Alice."
Credentials are also:
Alice's SSO.
But the ones truly operating the system may be:
200 Alice Agents.
The identity is still:
One person.
But the behavior has already become:
A machine cluster.
The entire security model changes instantly.
────────────────
5. This is why "successful authentication" no longer equals "trusted behavior"
Traditional security:
You logged in.
This means:
You are an employee.
Thus:
You can access.
This is called:
Identity-centric Security.
In the Agent era:
Alice's Agent obtains Alice's permissions.
It may:
Work normally;
Get stuck in a loop;
Misunderstand;
Be subject to prompt injection;
Call the wrong tools.
So future security systems must ask not just:
Who are you?
But also:
What are you doing, at what speed, and why?
────────────────
6. Security is transitioning from Authentication to Behavior Governance
The core of future Agent Security will include:
Identity;
Permissions;
Behavior;
Speed;
Scope;
Context;
Audit.
For example:
Allow Agents:
To read /reports.
But:
Cannot access /payroll.
Allow:
To call APIs 100 times per minute.
Cannot:
100,000 times.
Allow:
To generate code.
But deploying to Production:
Must be manually approved.
This is called:
Fine-grained Agent Control.
────────────────
7. This is closer to the real issues enterprises are facing today than "Will AI suddenly destroy the world"
This is a very clear stance in recent a16z discussions.
Databricks CEO Ali Ghodsi, along with a16z's Martin Casado and Sarah Wang, recently discussed and deliberately distinguished:
Long-term recursive self-improvement risks
From:
The very real issues of cyber attacks, enterprise automation, and contextual problems today.
This does not mean that X-Risk is not worth studying.
Rather:
Risks have a time scale.
Today, engineering teams must first address:
The real existing attack surface.
────────────────
8. What exactly is Anthropic's so-called "Pacing"?
A recently debated term is:
Pacing the Frontier.
Anthropic's claim is not simply:
Stop AI.
It advocates that:
The speed of capability growth should not long-term exceed the speed of safety, monitoring, and governance capability growth.
Anthropic has proposed:
Measuring the degree of AI automation in AI R&D;
Agent behavior oversight coverage;
How compute is allocated;
Even allowing external third parties to enter labs for independent assessments.
So simply translating Pacing as:
"Pause AI"
Is not accurate.
────────────────
9. However, "Pacing" indeed presents a very difficult policy question: What is a normal speed?
The questions raised by critics here are important.
Cars:
Have speedometers.
Planes:
Have speeds.
AI technology advancement:
Does not have a natural unit.
How do you know:
This year should increase by 20%,
Or 50%?
Anthropic itself is aware of this issue, which is why it has begun to try to build:
R&D Automation Index;
Agent Oversight Metrics;
Compute Allocation Metrics.
This indicates that:
"Slowing down" itself is not a complete policy.
It must first:
Measure the pace.
────────────────
10. Any "deceleration policy" without verifiable metrics is likely to become political language
For example:
"AI is developing too fast."
What does that actually mean?
Capability Benchmark?
Training FLOPs?
Agent autonomy?
Research automation?
Revenue?
Model release time?
Different definitions will lead to:
Completely different policies.
So truly high-quality AI governance must start from:
Measurement
Otherwise:
Pacing
Can easily become:
An unfalsifiable concept.
────────────────
11. Anthropic has begun to try to quantify this abstract debate
As of August 2026, Anthropic reports that:
Claude has reached the "AI Leads" automation level in about 26% of internal AI R&D work;
Over 90% of related work has reached at least "AI Collaborates."
Its main internal Agent platform operates about:
30,000 research and engineering Agents at any given time.
This is actually very noteworthy.
Because it indicates that:
Agent Swarm
Is no longer:
A future concept.
Frontier labs have already begun to operate it in reality.
────────────────
12. Anthropic also released a more interesting piece of data: Agent misbehavior is rare, but scale changes the risk
Its online security monitoring covers:
100% of Agent Actions.
In August 2026, out of over 1 billion Agent decisions,
About 0.002% were blocked by real-time monitors,
Approximately:
Once every 47,000 times.
It sounds:
Extremely low.
But what if in the future there are:
100 billion Agent Actions?
Even if the error rate remains unchanged:
The absolute number of incidents would still be very large.
────────────────
13. This is the most important math of Agent risk: low probability × massive scale
An employee makes a mistake once a year:
The company can handle it.
One million Agents:
Make a large number of actions every second.
Even if the single error rate is:
One in a million,
It could still lead to:
A large number of incidents daily.
This is called:
Rare-event Scaling.
Many future AI Safety issues,
Are actually:
Statistics.
Fourteen, therefore Agent Safety is not about pursuing "never being wrong".
This is unrealistic.
The correct goal is closer to:
Financial systems.
Aircraft systems.
Cloud computing.
The core is not:
Zero errors.
But rather:
Detect → Limit → Recover.
Discover.
Limit impact.
Recover.
This is:
Resilience Engineering.
────────────────
Fifteen, the biggest safety metric for future enterprise AI may not be Model Alignment, but rather Blast Radius.
After an Agent makes a mistake:
How large of a range can it affect?
Can it only:
Modify its own files?
Or:
Delete the entire database?
Can it cause:
$100?
Or:
$10 million?
This is called:
Blast Radius.
A truly mature Agent system must strive to:
Minimize the Blast Radius.
────────────────
Sixteen, therefore the "principle of least privilege" will become extremely important again.
The traditional security field has long had:
Principle of Least Privilege.
Only give the minimum permissions necessary to complete a task.
In the past, companies often neglected to implement it.
Because:
It was too troublesome.
In the Agent era:
It must be done.
Because:
An Agent's operation frequency
Is far higher than that of a human.
Coarse-grained permissions suddenly become:
Extremely dangerous.
────────────────
Seventeen, one of the biggest opportunities for the next generation of operating systems is to redesign Agent Permission.
Today permissions are usually:
Allow.
Deny.
Or:
Pop-ups.
In the future, it may be:
Only allow access to this Folder;
Only for the next 20 minutes;
Only Read;
Only for this task;
At most call 50 times;
Automatically escalate to human if exceeding risk threshold.
This is called:
Ephemeral Capability-based Security.
Short-term, fine-grained, task-bound permissions.
────────────────
Eighteen, this is a huge entrepreneurial market.
In the future, there may emerge:
Agent Identity;
Agent Permission;
Agent Firewall;
Agent Observability;
Agent Runtime;
Agent Rate Limiting;
Agent Audit.
a16z has already continuously invested in and discussed:
Keycard;
Neo
These types of Agent-native Security companies.
This indicates:
AI Security has begun to form a new:
Infrastructure Category from model security.
────────────────
Nineteen, Agent Security is likely to repeat the history of Cloud Security.
When AWS first appeared:
Everyone first asked:
Is the Cloud secure?
Then emerged:
CrowdStrike;
Wiz;
Zscaler;
Okta;
A large number of Cloud Security giants.
The same goes for Agents.
First stage:
Can Agents be used?
Second stage:
How to manage Agents?
Third stage:
Emergence of:
Agent Security Stack.
This is almost certain to become an industry.
────────────────
Twenty, why is the regulatory debate so easily hijacked by "X-Risk"?
Because "human extinction" belongs to:
The maximum possible loss.
As long as someone says:
There is a 10% probability of extinction,
Traditional cost-benefit analysis is almost ineffective.
Because:
The loss is:
Infinite.
Thus:
Any regulation can be rationalized.
This is also a major concern for critics.
────────────────
Twenty-one, but here we must avoid a fallacy: "the existence of serious risks" does not mean we can only shut everything down.
For example:
Nuclear power has significant accident risks.
We can:
Regulate.
Aviation can kill people.
We still:
Fly.
Drugs may have side effects.
We still:
Research and develop.
The way society has long dealt with high-risk technologies is usually not:
A binary choice.
But rather:
Risk Gradient.
Different risks:
Different regulatory intensities.
────────────────
Twenty-two, what AI policy really needs to address is not "regulate or not regulate" but rather:
What to regulate?
Model training?
Model release?
Dangerous capabilities?
Biological applications?
Critical infrastructure?
Autonomous weapons?
Ordinary enterprise customer service?
These risks are completely different.
If all are regulated with the same set of rules:
It will definitely create huge inefficiencies.
So the truly reasonable direction in the future should be:
Capability-based / Risk-based Regulation.
────────────────
Twenty-three, why is regulatory capture a risk that the AI field must seriously consider?
The basic mechanism of Regulatory Capture is:
Large companies have:
Lawyers;
Policy teams;
Compliance.
Startups do not.
If regulatory rules are exceptionally complex:
Large companies may complain,
But they can afford it.
Small companies:
Directly die.
Thus:
Regulation actually raises:
Entry Barriers.
This has occurred in banking, healthcare, telecommunications, and other industries.
────────────────
Twenty-four, so a very dangerous scenario is: safety standards become a "scale moat".
Assuming that only companies with:
A 1000-person Safety Team
Can legally train models.
Then the market will ultimately be left with:
A few companies.
This may reduce:
Innovation;
Open-source competition;
Price competition.
Therefore, regulators must solve a real contradiction:
Raise safety thresholds
But do not:
Artifically create oligopolies.
────────────────
Twenty-five, Anthropic itself does not actually advocate a complete ban on Open Weight.
This needs to be specifically corrected.
Dario Amodei stated in July 2026:
Anthropic has never advocated a complete ban on Open-weight Models;
Open-weight models without dangerous capabilities are:
A public good.
So simply describing Anthropic as:
"Wanting to kill open source through regulation"
Does not align with its public stance.
Of course, the outside world can still question whether its policy proposals objectively raise entry barriers.
But this belongs to:
Consequential judgment.
It cannot replace:
Facts.
────────────────
Twenty-six, the real lesson from historical regulation is not "we must wait until people die before we can regulate".
The history of automobiles and aviation is much more complex than this simple narrative.
The U.S. passed the Air Commerce Act in 1926, beginning:
Aircraft airworthiness certification;
Pilot licensing;
Aviation safety regulation.
In 1938, the Civil Aeronautics Authority was established.
The FAA, in its current sense, was not established until 1958.
So aviation regulation is not:
"First develop freely for 40 years, then suddenly regulate."
But rather:
As factual patterns continue to emerge, regulation is layered and upgraded.
────────────────
Twenty-seven, this is the truly normal evolution of technology regulation.
First stage:
New technology.
Very little regulation.
Second stage:
Specific risks emerge.
Establish:
Industry rules.
Third stage:
Scale expands.
Establish:
Formal institutions.
Fourth stage:
Accidents and new technologies continue to drive:
Standard upgrades.
This is called:
Adaptive Regulation.
Not:
One-time design of laws that are always correct.
────────────────
Twenty-eight, the history of automobiles is the same.
Ralph Nader's "Unsafe at Any Speed" was published in:
1965.
In 1966, the U.S. passed significant automobile and highway safety legislation,
Strengthening the federal vehicle safety standards system.
But the real lesson is not:
"The government completely ignored cars before 1965."
But rather:
As the scale of automobiles grew,
Society's demands for safety standards continuously increased.
AI may very well be the same.
────────────────
Twenty-nine, what regulators fear most is "the technical structure is not yet stable, and today's architecture is written into law".
This is where predictive regulation is truly dangerous.
Assuming that in 1994 the law stipulated:
The Internet must:
Operate according to a specific network architecture.
Today, that could be very absurd.
Because:
Technology is constantly changing.
So good regulation should focus more on:
Outcomes.
For example:
Must not leak specific data;
Must report significant accidents;
High-risk systems must pass testing.
And not:
Force stipulations on:
Every specific implementation.
────────────────
Thirty, the history of the Morris Worm is particularly suitable to illustrate "the real risks of new technology will gradually emerge".
In 1988, the Morris Worm infected:
About 6000 machines
In about 24 hours,
At that time about:
60000 networked computers.
Close to:
One-tenth.
This was an extremely serious security incident in the early Internet.
It illustrates:
The dangers of open networks:
Are real.
But society ultimately did not:
Shut down the Internet.
Instead, it developed:
Cybersecurity.
────────────────
Thirty-one, the truly correct experience from the history of technology is: risks often create another industry.
Internet:
Creates hacking risks.
Simultaneously creates:
Cybersecurity.
Cloud:
Creates Cloud Risk.
Simultaneously creates:
Cloud Security.
Crypto:
Creates Custody Risk.
Simultaneously creates:
Crypto Custody.
AI Agents:
Create Agent Risk.
Will also create:
Agent Security Industry.
This is a pattern that capital should pay close attention to.
────────────────
32. The era of AI Agents may redefine the total market for Cybersecurity
Past security protections:
Devices.
Networks.
Users.
In the future, we also need to protect:
Autonomous Actors.
Each enterprise may run:
Tens of thousands of Agents.
Each Agent:
Has an identity;
Permissions;
Memory;
Tasks.
This is equivalent to a sudden increase in:
A large number of new:
"Digital Employees."
And these employees:
Never sleep.
The Security TAM will therefore:
Expand significantly.
────────────────
33. Why is the emergence of Jev particularly noteworthy at this point in time?
Because it tackles a completely different problem.
LLMs excel at:
Generation.
Jev does not generate text.
What it does is:
Decision.
TypeSafe AI defines Jev as:
System One Model.
Input:
State.
Then provide it:
Limited options.
Output:
Choice + Probability + Confidence.
This is a very different AI Primitive.
────────────────
34. A simple example helps to understand Jev
Traditional LLM:
Input customer service ticket.
Prompt:
"Determine whether it should be forwarded to Billing, Fraud, or Support, please output JSON."
The model generates:
Text.
Then the software:
Parses JSON.
Validates Schema.
Handles errors.
Jev:
Directly inputs the ticket.
Gives three options:
Billing;
Fraud;
Support.
Output:
Fraud = 0.83.
Billing = 0.12.
Support = 0.05.
Software:
Uses it directly.
This is:
Language → Decision.
────────────────
35. Why might this seemingly simple change be very significant?
Because the traditional software world does not like:
Text.
Programs prefer:
Types.
Boolean.
Integer.
Enum.
Probability.
LLM's output is naturally:
Strings.
So developers have been doing something very absurd:
Making models write natural language.
Then parsing natural language:
Back into a structure that machines can understand.
The process is full of:
Failure points.
────────────────
36. What Jev is really trying to eliminate is the "text intermediary layer"
Traditional chain:
State
→ LLM
→ Text/JSON
→ Parser
→ Validator
→ Program Logic.
Jev:
State
→ Probabilistic Decision
→ Program Logic.
Fewer:
Steps.
Means:
Faster;
Cheaper;
More reliable.
This is the core thesis of TypeSafe.
────────────────
37. FT reports that Jev's positioning can even reach a speed and cost advantage magnitude of 100 times that of traditional LLMs
TypeSafe positions Jev as:
A model for high-frequency, programmatic decisions.
For example:
Request Approval;
Risk Assessment;
Routing.
And emphasizes that it is not:
A chatbot replacement.
But rather:
An internal judgment component of software.
This distinction is very important.
────────────────
38. Jev is not trying to kill LLMs, but to drive LLMs out of areas where they are not good
Open-ended:
Writing articles.
Research.
Complex reasoning.
Use:
LLMs.
High-frequency small judgments:
Is it fraud?
Should it escalate to human?
Which tool to call next?
Use:
Jev-like systems.
This is actually:
Division of Cognitive Labor.
Different models do:
Different tasks.
────────────────
39. Future software will not only have "one type of AI"
Just like computers:
CPU.
GPU.
TPU.
ASIC.
Different chips:
Different tasks.
Future AI will also have:
Reasoning Model;
Decision Model;
Embedding;
Vision;
Speech;
Small Local Model.
Ultimately, applications will have:
Model Heterogeneity.
Multi-model systems.
────────────────
40. This indicates that "bigger models getting bigger" is not the only direction for AI
In the past, the industry almost formed a subconscious belief:
Bigger = Better.
But the most important thing for real software systems is:
Intelligence per Dollar.
If a task:
Is completed by a 100B model
And:
A small Decision Model completes it
With the same result,
Enterprises will definitely choose:
The latter.
This is an economic law.
────────────────
41. The greatest strategic significance of Jev may be moving AI from "UI" to "program logic"
ChatGPT:
AI is a product.
Jev:
AI is:
An if statement.
Not directly facing the user.
But rather:
Existing within the software.
This is a very important paradigm shift.
────────────────
42. In the past, software logic was deterministic
if amount > 10000:
require_approval()
Clear.
Certain.
But many rules in the real world are not:
Boolean.
For example:
"Does this transaction look suspicious?"
"Is this customer service ticket urgent?"
"Is this user likely to churn?"
In the past, programmers found it difficult to write.
Because:
There was no deterministic formula.
────────────────
43. AI allows software to directly handle "fuzzy concepts" for the first time
In the past:
Humans understand meaning.
Software understands rules.
Now AI can:
Convert:
Meaning
Into:
Probability.
For example:
"The probability that this behavior is Fraud is 87%."
This forms a new Primitive:
Semantic Computation.
This is where Jev is truly interesting.
────────────────
44. Therefore, future software may exist in two worlds simultaneously
First layer:
Deterministic Logic.
How much money is there?
What is the user ID?
Is there a payment?
Second layer:
Probabilistic Logic.
Is it suspicious?
Is it important?
What does the user really want to do?
Future programs:
Combine the two layers.
This may be more realistic than:
"All software becoming chatbots"
────────────────
45. "Probabilistic Programming" did not just emerge in 2026
This needs a historical correction.
Probabilistic computation, Monte Carlo, and random simulation existed in large numbers in the early days of computing;
Languages like SIMULA in the 1960s promoted computer simulation.
Later, Probabilistic Programming developed into a more defined language and research field.
So Jev is not:
"Inventing probabilistic programming."
More accurately:
Turning modern neural models into probabilistic decision Primitives that ordinary software developers can directly call for the first time.
────────────────
46. "if 73% then" is more suitable as a metaphor rather than a future real syntax
Future developers might write:
if fraud_probability > 0.92:
freeze()
elif fraud_probability > 0.65:
human_review()
The core is not:
Syntax.
But rather:
Software begins to treat:
Uncertainty
As a first-class citizen.
In the past, software was very afraid of uncertainty.
In the future, software will:
Directly manage uncertainty.
This is a deeper change.
────────────────
47. This creates a huge new concept: Confidence-aware Software
Traditional software:
Success.
Failure.
In the future:
90% Confidence.
60%.
20%.
Based on:
Confidence
Choose:
Auto-execute;
Escalate to human;
Continue querying.
This is actually very suitable for:
Agent Systems.
Because Agents are inherently:
Uncertain.
────────────────
48. The real answer to reliable AI may not be "making models always correct"
But rather:
Letting software know when models might be wrong.
This is much more realistic than pursuing:
100% Accuracy
For example:
Confidence > 99%:
Automatic.
90%:
Human review.
<70%:
Change model.
This is:
Uncertainty Routing.
A very important architecture for future enterprise AI.
────────────────
49. Why did Jev suddenly ignite developer interest in San Francisco?
It was released on September 15,
And within days, it formed a very high level of developer discussion;
Business Insider reported on its Hackathon and early community expansion, and FT also reported that investors began discussing this company at extremely high valuations.
The real reason may not be:
Benchmark.
But rather:
Developers finally see:
Besides chatbots,
AI can also become:
A Programming Primitive.
────────────────
50. This may be similar to the moment when databases just entered software development
Today, no one thinks:
Databases are an "App."
Databases are:
Fundamental components.
In the future:
AI Decision Engine
May also be like this.
Developers never say:
"I am using AI."
Just like today, they wouldn't say:
"My software uses SQL, so it is an AI Startup."
The most mature form of AI:
Might just be:
Invisible Intelligence.
────────────────
51. If the Jev route holds, SaaS might actually become one of the biggest beneficiaries.
The market has long talked about:
SaaSpocalypse.
AI is going to destroy Salesforce.
Destroy Workday.
Destroy ServiceNow.
TypeSafe's counterpoint is:
Old SaaS has:
Data;
Users;
Workflows;
State.
If we add:
Probabilistic Intelligence,
they might suddenly:
Automate more work.
This is:
Inverse SaaSpocalypse.
────────────────
52. The biggest problem with traditional SaaS is not product differentiation, but rather rigid logic.
For example, CRM:
If a customer doesn't respond for three days:
Send an email.
This is very mechanical.
Intelligent CRM:
Understands:
Is the customer cold?
Are they negotiating?
Are they just on vacation?
Such judgments used to require:
Human intervention.
In the future:
It can be directly integrated into the program.
This will make old SaaS:
Suddenly become:
Context-aware.
────────────────
53. The real SaaS revolution may not be "Agent replacing SaaS," but rather "software gaining judgment on its own."
This statement is very important.
In the past:
Human + Software.
In the future:
Intelligent software.
The software itself:
Understands semantics.
Makes judgments on its own.
Agent is just:
One type of UI.
The deeper change actually happens inside:
The program.
This is what Diogo means by:
"putting intelligence inside software itself."
────────────────
54. This could open up a whole new software economy.
Many software functions in the past:
Simply couldn't be done.
Because it was impossible to convert:
Natural language
Into:
Definite rules.
For example:
Automatically approving expense reimbursements.
Judging:
"Is this a reasonable business expense?"
The rules are endless.
After the emergence of AI judgment models:
Suddenly possible.
This means:
Software Addressable Market Expands.
The things software can automate:
Increase.
────────────────
55. AI Coding Agent and Jev solve two completely different problems.
Coding Agent:
Helps programmers:
Write software faster.
Jev:
Enables software itself:
To do things it couldn't do before.
The former:
Production Efficiency.
The latter:
Capability Expansion.
In the long run:
The second may be larger.
────────────────
56. The real breakthrough of the Industrial Revolution was never about "craftsmen working a bit faster,"
But rather:
Machines being able to accomplish:
Things that human labor could never scale.
If AI only increases:
Programmer efficiency by 30%,
Of course it has value.
If it:
Allows software to understand:
Ambiguous business rules,
Then that is:
A new computing paradigm.
This could potentially:
Rewrite the entire software industry.
────────────────
57. The real big opportunity is shifting from "inside the model" to "around the model."
Including:
Context;
Memory;
State;
Permissions;
Routing;
Observability;
Security;
Decision Models;
Workflows.
All of these are:
Model-adjacent Infrastructure.
The next generation of huge companies,
May emerge significantly from here.
────────────────
58. Why might the giants of foundational models not necessarily consume all these markets themselves?
Because platforms can never:
Optimize all vertical demands.
AWS does not do:
All SaaS itself.
Windows does not do:
All software itself.
iPhone does not do:
All apps itself.
Successful platforms,
Instead will give rise to:
A larger application ecosystem.
AI is likely the same.
────────────────
59. So "Model is the new OS" is only half right.
More accurately:
Model is:
A new computing primitive.
Real software systems still need:
An Operating Layer.
For example:
Which model?
When to call it?
What permissions?
Where to save state?
What to do on failure?
Who approves?
How to audit?
This is:
Agent Operating System.
────────────────
60. The future AI Stack is likely to form a four-layer structure.
First layer: Compute.
GPU, ASIC, Cloud.
Second layer: Models.
Reasoning, Decision, Vision, Speech.
Third layer: Orchestration.
Memory, Permissions, Routing, Security.
Fourth layer: Applications.
Finance, Legal, Healthcare, Robotics.
Today, the vast majority of media attention is still on:
The second layer.
But the real entrepreneurial opportunities are migrating significantly to:
The third and fourth layers.
────────────────
61. This is the repeating pattern in the history of software.
Mainframes appeared.
Then:
Software.
PCs appeared.
Then:
Applications.
The Internet appeared.
Then:
Web Companies.
Cloud appeared.
Then:
SaaS.
After the emergence of the Foundation Model:
The real big Application Cycle:
May have just begun.
────────────────
62. This is also why excessive focus on AGI discussions may obscure many real business opportunities.
The models that already exist today:
Are sufficient to automate:
Many jobs.
The real barrier is not:
Insufficient intelligence.
But rather:
Companies not knowing how to:
Provide Context;
Provide Permission;
Secure access;
Design Workflows.
One of Ali Ghodsi's core judgments in a recent a16z talk is:
Today's model capabilities have already surpassed the actual utilization levels by enterprises,
The bigger bottleneck lies in organizational context and institutional knowledge.
This is very important for entrepreneurs to understand.
────────────────
63. The biggest mistake entrepreneurs might make is waiting for "smarter models."
Many founders say:
Wait for the next generation of models.
In fact, it may not be necessary at all.
Today, Claude / GPT / Gemini:
Are already sufficient to solve problems.
What really needs to be done is:
Systems Engineering.
Correctly embedding the models into work.
This is more important than:
Waiting for 20 more IQ.
────────────────
64. One of the biggest wars for AI companies in the next five years: who owns Context.
General models know:
The world.
What enterprises really need:
Is to know:
This company.
Who has made decisions in the past.
Who the customers are.
Internal rules.
Historical anomalies.
Organizational relationships.
This is:
Organizational Memory.
Whoever controls:
The Context Layer,
May control:
The enterprise Agent.
────────────────
65. This connects back to security.
The more context there is:
The smarter the Agent.
At the same time:
The greater the risk.
So future Agent products will always face a:
Core contradiction:
Utility vs. Permission.
The more you give:
The more useful it is.
The more you give:
The more dangerous it is.
Truly excellent systems must solve:
How to dynamically balance between the two.
────────────────
66. Therefore, the truly great AI Infra companies of the future are likely not the "coolest," but the "most boring."
Permissions.
Auditing.
Rate limiting.
Queues.
State.
Identity.
Logs.
These sound:
Not sexy at all.
But every time there is a computing platform revolution,
One of the ultimate sources of commercial value:
Is precisely built on such:
Boring infrastructure.
────────────────
67. Stripe is very boring: payment API.
Datadog is very boring:
Logs.
Okta:
Login.
Cloudflare:
Networking.
But these companies have immense value.
The Agent era will similarly produce:
Boring AI Infrastructure Giants.
Real capital should not only focus on:
Robot demos.
────────────────
68. One thing that regulators should really learn from software engineering is: do not merge all risks into one term.
"AI Risk"
Is too broad.
Just like:
"Internet Risk"
Is meaningless.
Fraud risk.
Cyber risk.
Privacy risk.
National security.
Youth safety.
Completely different.
If regulation is all:
One-size-fits-all,
It will definitely lead to inefficiency.
So future AI governance needs:
Risk Decomposition.
────────────────
69. Truly mature AI policies should gradually become engineered like cybersecurity.
Not:
"AI is very dangerous."
But rather:
This model:
What is its Cyber Capability?
How much Bio Capability?
How high is Agent Autonomy?
How high is Monitor Coverage?
What is the Incident Rate?
Anthropic has recently started to publicly disclose these metrics, which itself is a step in the regulatory discussion from:
Philosophy
to:
Engineering.
This is a commendable direction, regardless of whether one agrees with all of its policy claims.
────────────────
Seventy, what role should X-Risk truly play?
It should:
Promote:
Research;
Testing;
Tail-risk Planning.
But it cannot automatically become:
A pass for all regulatory policies.
Because:
Once extreme risks lack falsifiability,
it is easy to:
Infinitely expand government powers.
The correct approach should be:
To continuously break down abstract disaster risks into testable capabilities.
For example:
Autonomous Cyber Attacks.
Dangerous biological designs.
Autonomous model replication.
Instead of stopping at:
"The model could lead to human extinction."
────────────────
Seventy-one, this is the difference between scientific governance and fear-based governance.
Fear-based governance:
First assumes disaster.
Then:
Controls everything.
Scientific governance:
Defines:
Risk variables;
Measures;
Tests;
Adjusts rules.
Whether future AI governance can move towards the second model
will greatly impact:
Innovation speed
and:
Safety levels.
────────────────
Seventy-two, for entrepreneurs, the most important takeaway from this entire debate is: do not focus solely on Frontier Models.
The truly massive entrepreneurial opportunities include:
Agent Security;
Agent Permission;
Context Infrastructure;
Model Routing;
Probabilistic Decision;
Enterprise Memory;
Observability;
Evaluation;
Human Escalation.
These markets:
Are still very early today.
────────────────
Seventy-three, do not think "OpenAI will do it" so you cannot start a business.
Google has search.
Shopify still emerged.
AWS has Cloud.
Snowflake still emerged.
The larger the platform:
The larger the ecosystem.
The real entrepreneurial question is:
Is your layer:
Sufficiently vertical;
Sufficiently complex;
And does the platform lack the motivation to optimize itself?
If:
Yes.
Then there is an opportunity.
────────────────
Seventy-four, Jev also gives AI entrepreneurs a very important new perspective: do not assume AI products must be chat-based.
In the past three years:
All AI products:
Input boxes.
Conversations.
This is merely:
The path dependency of ChatGPT.
The real product should ask:
Do users really need text?
Or:
Decision?
Action?
Score?
Alert?
In many scenarios:
There is no need to:
Generate a sentence.
This can save huge:
Latency
and:
Cost.
────────────────
Seventy-five, in the future, AI will ultimately "disappear" into software.
Today:
AI is the selling point.
In the future:
Users will not even know:
Whether there is AI behind it.
Just like:
Today no one asks:
Whether Uber's backend is PostgreSQL.
Mature technology will ultimately:
Become Invisible.
Jev's model precisely represents:
AI moving towards intangible infrastructure.
────────────────
Seventy-six, this could mark the transition from Generative AI to Computational AI.
First stage:
AI generates:
Text;
Images;
Videos.
Second stage:
AI begins:
To become the logic of program execution.
Generation is:
Content Revolution.
The latter is:
Computing Revolution.
If the latter succeeds,
The impact could be deeper.
────────────────
Seventy-seven, the metrics that investors should truly observe will also change.
In the past, they looked at:
Benchmark.
In the future, they will look at:
Cost per Correct Decision.
Latency.
Calibration.
Reliability.
Human Escalation Rate.
Agent Error Rate.
Infrastructure Load.
These are the real production environment metrics.
If Jev's value ultimately holds,
It is not:
How smart it looks.
But rather:
How much per million decisions:
How much does it cost.
How many mistakes are made.
────────────────
Seventy-eight, this actually pulls AI back to the most traditional core of computer science: Engineering.
In the past few years, AI has been like:
Alchemy.
Training larger models.
Producing magical capabilities.
The next stage:
Must enter:
Engineering discipline.
Reliability.
Cost.
Permissions.
Testing.
Recovery.
This is true industrialization.
────────────────
Seventy-nine, every time revolutionary technology moves from the laboratory to industry, it goes through this stage.
Cars:
Initially could run.
Later:
Reliable.
Planes:
Initially could fly.
Later:
Safe.
Internet:
Initially could connect.
Later:
Security + Cloud.
AI:
Initially could think.
The next step:
Can reliably work.
This is the real change happening now.
────────────────
Eighty, the future AI's greatest value will not come from "models resembling humans," but from "software finally being able to understand human intentions."
In the past, computers required:
Humans to adapt to computers.
Buttons.
Menus.
Rules.
SQL.
In the future:
Humans will say:
"This customer is important; alert me if there is danger."
The system understands:
"Important"
and:
"Danger."
This is:
Intent-native Computing.
It may be much more important than the chatbots themselves.
────────────────
Eighty-one, Jev's ultimate vision can be condensed into one sentence:
Not:
Tell the computer exactly what to do.
But rather:
Tell the computer what you mean.
This is also the ultimate goal repeatedly emphasized by Diogo Almeida in the latest a16z program:
To enable technology to reliably:
"do what I mean."
If this step is truly achieved,
The abstraction level of software development will rise again.
────────────────
Eighty-two, every major leap in software history has essentially been about raising the level of abstraction.
Machine code.
Assembly.
High-level languages.
Databases.
Web Frameworks.
Cloud.
Each layer allows programmers to:
Describe less about "how to do it."
And more about:
"What I want."
AI is the next step in this history.
Jev's system:
Further transforms:
Semantics
into:
Executable software Primitives.
────────────────
Eighty-three, this may allow a large number of jobs that cannot be automated today to finally enter software.
In the past, many industry processes could not be automated,
Not because:
There were no computers.
But because:
The rules could not be articulated.
For example:
"Does this contract look unusual?"
"Is this customer worth prioritizing?"
"Is this content close to the brand tone?"
These are:
Fuzzy Logic in Human Language.
AI can finally:
Scale to handle.
────────────────
Eighty-four, so the largest TAM for AI may not be today's software market at all,
But rather:
All the jobs that had to rely on human judgment because the rules could not be clearly articulated in the past.
This includes:
Management;
Auditing;
Customer service;
Compliance;
Risk;
Sales;
Operations.
In other words:
AI ultimately attacks:
The Judgment Labor Market.
This is much larger than:
The SaaS market.
────────────────
Eighty-five, but the deeper we go into Judgment, the more important the issue of responsibility becomes.
If the model judges:
80% fraud.
Freezes accounts.
Customer losses.
Who is responsible?
The programmer?
The model company?
The software platform?
The enterprise?
This is the extremely difficult question of the future AI era:
Probabilistic Decision + Legal Accountability.
Technology runs faster than law.
────────────────
Eighty-six, this is also why Agents and Probabilistic Software will ultimately need Human Escalation.
High Confidence:
Automatic.
Medium Confidence:
Review.
Low Confidence:
Manual.
A truly mature system will not:
Pursue the complete disappearance of humans.
But rather:
Optimally allocate human judgment.
Humans will only handle:
The most ambiguous;
Highest risk;
Most expensive
judgments.
AI handles:
The remaining 95%.
────────────────
Eighty-seven, this could become the greatest source of productivity for enterprise AI.
Not:
Everyone unemployed.
But rather:
A Risk Analyst
Previously handled:
100 Cases.
In the future:
AI automatically handles:
900.
Humans only look at:
100 boundary cases.
The same employee:
Produces:
10 times more.
This is:
Human-AI Division of Labor.
────────────────
Eighty-eight, so the real winners may be those companies that have the best Escalation Architecture.
When:
Does AI do it?
When:
Do humans do it?
This boundary is very important.
It is not as simple as:
"Human-in-the-loop."
But rather:
Human-where-it-matters.
Truly excellent software will dynamically decide:
When people should appear.
────────────────
Eighty-nine, from a capital perspective, the most important point of this round of changes: the model value chain is disintegrating into multiple independent markets.
In the past, everyone thought:
After AGI appears:
One model
Will capture all value.
Now it is increasingly likely that:
Model;
Decision;
Agent;
Security;
Memory;
Context;
Vertical Workflow
Will each form a company.
This is:
AI Value Chain Specialization.
Specialization.
This is good news for the entrepreneurial ecosystem.
────────────────
Ninety, the largest AI company in the future may not be the "smartest" company, but the one that can turn intelligence into reliable cash flow.
Model intelligence:
Is just raw material.
Real commercial value comes from:
Automation;
Results;
ROI.
So in the future, we will increasingly ask less:
How high is this model's IQ?
And more:
How many dollars of economic output can each dollar of intelligence generate?
This is the ultimate commercial metric.
────────────────
Ninety-one, the recent discussions at a16z really point to a huge change.
First stage AI:
Intelligence Scarcity.
Who has smart models?
Second stage:
Execution Scarcity.
Who can really make models:
Enter systems;
Run securely;
Make reliable decisions;
Act continuously?
Today we are entering:
The second stage.
────────────────
Ninety-two, this also means that entrepreneurial opportunities are shifting from "building brains" to "building nervous systems."
Models:
Brains.
But a person also has:
Nerves;
Muscles;
Immune systems;
Memory.
AI systems are the same.
Beyond models, we need:
Tooling;
Memory;
Permissions;
Security;
Infrastructure.
In the future, dozens or even hundreds of large companies:
May emerge from:
"Around the brain."
────────────────
Ninety-three, this is why "the next era of software is not inside large models" is a good main line.
A more accurate understanding is:
Base models are still important.
But:
Marginal entrepreneurial opportunities
Are increasingly likely to appear at:
The periphery of models.
Just like:
Intel is important.
But the greatest wealth of the PC era was not all taken by:
CPU companies.
Microsoft;
Adobe;
Oracle;
Salesforce
All built on:
Chips.
AI will be the same.
────────────────
Ninety-four, a truly mature software industry will certainly commoditize models.
Today everyone talks about:
GPT-5.6.
Claude 5.5.
A few years later:
Ordinary developers may not care at all.
Systems will automatically route.
Just like today’s cloud.
No one discusses daily:
Which specific server model.
Underlying intelligence will gradually:
Become Infrastructure Commodity.
The application layer:
Creates value upwards.
────────────────
Ninety-five, this is also the most interesting industrial signal from Jev: not all intelligence must come from Frontier LLM.
Many tasks:
Actually do not need:
Poetry;
Reasoning;
Creativity.
Only need:
Judgment.
Once specialized models can:
Be 100 times cheaper;
Be 100 times faster,
The entire market will naturally:
Layer.
This will drive:
AI Efficiency Revolution.
────────────────
Ninety-six, the next major battle in the AI industry may shift from Scaling to Allocation.
In the past:
Who has more Compute?
In the future:
How to apply the right Compute to the right tasks?
Complex problems:
Frontier Model.
Small decisions:
Jev.
Private tasks:
Local Model.
High-frequency tasks:
Small Model.
This is:
Intelligence Routing.
This will become:
The new cloud computing.
────────────────
Ninety-seven, from an investment perspective, this change is very important.
The first wave of money was made in:
GPU.
The second wave:
Foundation Models.
The third wave:
Is likely to be made in:
Intelligence Infrastructure.
Who can:
Manage;
Secure;
Route;
Monitor;
Memory;
Execute.
This is a layer worth watching closely in the coming years.
────────────────
Ninety-eight, and regulation should really follow this layering.
High-risk Frontier Training:
A set of rules.
Agent access to nuclear facilities:
Extremely high standards.
Ordinary customer service routers:
Do not need equal regulation.
Otherwise:
The entire AI ecosystem will be dragged down by:
Minimum common standards.
The keyword for mature regulation should be:
Proportionality.
How big the risk is:
How strong the regulation should be.
────────────────
Ninety-nine, truly good AI policy is not "less regulation" or "more regulation"
But:
More precise.
The goal of precise regulation:
At high-risk areas:
Strict.
At low-risk areas:
Free.
This is the only sustainable path to allow:
Innovation
And:
Safety
To coexist.
This is also much more valuable than labeling each other around:
Doomer vs. Accelerationist.
────────────────
One hundred, what should entrepreneurs really take away from this series of discussions?
First:
Do not wait for AGI.
Today’s models are already sufficient to create a lot of automated value.
────────────────
Second:
Do not just create Chat UIs.
See if you can integrate AI directly into:
Program logic.
────────────────
Third:
Security is not a later feature.
Once the agent has action capabilities:
Security must be designed from day one.
────────────────
Fourth:
Treat probability as part of the product.
Admit when you don’t know.
Then:
Upgrade manually.
────────────────
Fifth:
The real moat may be outside the model.
Context;
Workflow;
Permission;
Data;
Memory.
────────────────
One hundred one, what should enterprises really do?
Do not ask:
"Which large model should we procure?"
But should first draw:
Workflow Map.
Which steps:
Are deterministic.
Which:
Require judgment.
Which:
Are high-risk.
Which:
Can be fully automated.
Then:
Choose the appropriate intelligence for each step.
This is:
True AI Transformation.
────────────────
One hundred two, what should security teams really do?
Do not just study:
"Did employees download ChatGPT?"
But prepare for:
Agent-scale IT.
In the future, one employee is not:
One User.
But may be:
One Human Principal
Plus:
100 Agent Workers.
Enterprises must design in advance:
Identity;
Rate Limit;
Permission;
Logging;
Kill Switch.
────────────────
One hundred three, what should investors really observe?
Do not just look for:
The next OpenAI.
Also look for:
The next:
Okta for Agents;
Datadog for Agents;
Cloudflare for Agents;
Snowflake for Agent Memory;
Stripe for Agent Payments;
And:
Jev as a new Computing Primitive.
The infrastructure market:
May be much larger than imagined.
────────────────
One hundred four, whether Jev ultimately becomes a large company is still far from proven.
It was just made public on September 15, 2026.
Independent large-scale benchmarks are still insufficient.
Many details of the training methods:
Have not yet been disclosed.
Early community interest:
Is very high.
But:
Developer Excitement ≠ Durable Platform.
The most reasonable attitude now is:
To pay close attention.
But do not prematurely equate:
New paradigms
With:
Proven industry standards.
This is the most important discipline in technology investment.
────────────────
One hundred five, to truly judge Jev, five things should be considered.
First:
Calibration.
Is the probability really accurate?
Second:
Reliability.
Is it stable in real production environments?
Third:
Cost.
Is it still cheap after scaling?
Fourth:
Developer Adoption.
Has it entered real applications?
Fifth:
Replacement Resistance.
After large LLMs add similar capabilities, does Jev still have an advantage?
These determine:
Whether it is:
A Feature.
Or:
A Platform.
────────────────
One hundred six, the truly interesting part of AI history is: everyone may have initially misunderstood the interaction mode.
Early PCs:
Everyone imitated:
Paper.
Web:
Imitated:
Newspapers.
Mobile phones:
Initially shrank:
Desktop web pages.
When new platforms first appeared:
Humans always brought over:
Old paradigms.
After ChatGPT's success:
All AI products:
Everyone has started to create chat boxes.
A few years later, we might find:
Chat is the command line of AI.
Important.
But not:
The final interface.
────────────────
107. In the future, AI may not have a unified interface at all.
AI:
In emails.
In IDEs.
In banking systems.
In robots.
In databases.
In risk control engines.
Users:
May not even see it.
This is:
Ubiquitous Intelligence.
Intelligence like electricity:
Exists in all systems.
Rather than:
Everyone opening a dedicated:
AI App every day.
────────────────
108. This is the true larger AI endgame beyond ChatGPT.
ChatGPT proves:
Machines can understand language.
The next stage is to prove:
The world's software systems can understand language.
Banks.
ERP.
CRM.
Operating systems.
Robots.
All can:
Understand "what I mean."
This could redefine:
Computers.
────────────────
109. If this judgment is correct, the future of software engineering will see a very deep division of labor.
Programmers will no longer be responsible for:
Writing every real-world rule as:
if / else.
But will be responsible for:
Defining:
Boundaries;
States;
Permissions;
Uncertainty;
Outcomes.
Machines will be responsible for:
Filling in:
The ambiguous middle space.
The role of programmers will shift from:
Rule Author
to:
System Designer.
────────────────
110. This may be the most profound impact of AI on software engineering.
Today, people discuss:
Will AI replace programmers?
This question is too shallow.
The real question is:
What exactly is a program?
If programs shift from:
Completely deterministic logic
to:
Deterministic logic + probabilistic intelligence,
then:
Programming itself
is changing.
This is far more important than:
"Copilot helps me write 40% more code."
────────────────
The most memorable statement:
a16z's recent seemingly completely different discussions:
Dario Amodei's Pacing;
AI X-Risk;
Agent Cybersecurity;
Internal system DDoS;
Jev;
Probabilistic Programming;
SaaS;
Application layer innovation,
actually point to the same big trend:
The AI revolution is moving from Demo to systems.
In the Demo phase,
the most important thing is:
How smart the model is.
Once it enters systems,
the questions change completely:
Is it reliable?
Is it safe?
What is the cost?
When should we trust it?
When should we let humans take over?
How to grant permissions?
How to make it understand the enterprise?
How to convert natural language into truly executable software behavior?
None of these questions can be solved automatically by:
"Training a larger model."
So the biggest opportunity for AI in the next phase may not lie within:
The model itself.
But rather in:
The relationship between the model and the real world.
This layer requires:
Security;
Permission;
Context;
State;
Memory;
Routing;
Probabilistic Decisions;
Human Escalation.
In the past three years:
Everyone has been building:
Brains.
The next phase:
The entire industry needs to start building:
Nervous Systems.
And truly great software companies often emerge at such moments:
A new foundational capability exists,
But the world has not yet learned:
How to use it reliably.
The internet was like this.
Cloud was like this.
Smartphones were like this.
AI will be no exception.
From a capital perspective, this may even mean:
The Foundation Model Boom
is not the peak of AI commercialization.
Rather, it is just:
The infrastructure laying phase.
The truly massive applications and software reconstruction cycle
may just be beginning now.
The three core insights worth retaining from this are:
First, the biggest bottleneck in the next phase of AI has shifted from "not enough intelligence" to "the system cannot support intelligence." Agent-native Security, permissions, throttling, and state management will become a whole new set of infrastructure.
Second, what Jev truly deserves research is not "another model," but how it pulls AI back from the chat box into the program itself—allowing software for the first time to directly convert ambiguous human semantics into probabilistic, executable decisions.
Third, the regulatory debate should truly upgrade from the abstract "Is AI dangerous?" to measurable engineering questions: what capabilities, what risks, what frequencies, what Blast Radius, and then decide what level of regulation is needed.
A