OpenAI CEO Sam Altman: Next Generation Can Watch Screens All Day
OpenAI CEO Sam Altman told Cory Levy, founder of Z Fellows, that in about six months, the next generation of ChatGPT could continuously watch computer screens, listen to every meeting, and record every phone call, having complete context of users' lives and experiences. He stated, "We might be just one model away from being truly useful," and hopes this will change how he works.
He elaborated: "Approaching a world where the next generation of ChatGPT watches screens all day, attends every meeting, records every phone call, and has perfect context of entire lives and experiences. Users can choose the scope of authorization, allowing access to texts, emails, documents, Slack, etc. It does not make decisions for you; however, startup CEOs always have endless tasks and context to remember, and cannot read all customer feedback daily. When writing sales pitches or strategic memos, it will suggest another idea, point out possible mistakes, or proactively take on tasks."
The timeline is his prediction, not an announced product release date. The Mac version of ChatGPT has launched on Computer History, allowing tracking of selected applications and website activities. The research preview of Codex for Mac, Chronicle, will periodically take screenshots, send them to OpenAI servers for OCR and visual analysis, and then store text summaries in local Markdown for future prompts. Greg Brockman stated this allows the assistant to automatically obtain a complete update on what you are doing.
He described the interaction as evolving from repeatedly providing background to interjecting while you type. The boundaries of authorization are still set by the user: it only sees the channels you allow. He emphasized that the system acts as a parallel working brain, not a decision-maker. The industry has spent the past two years focusing on prompt compression; he is rewriting the bottleneck into continuous memory and screen-layer perception. The six-month window starts from the date of the interview, landing around early 2027, in line with the company's public rhythm of "not launching in 2026, first solving safety and social adaptation."
This advances the assistant from an independent webpage to a transparent layer that can be integrated into operating systems. Whoever controls the screen flow, call flow, and email authorization will control the next generation of work entry points, not just the share of dialogue boxes. Privacy and corporate compliance will simultaneously become product switches: without authorization, there is no context; with authorization, there is no need to start from scratch.
In market mechanisms, this is a battle for entry points, not a new parameter leaderboard. Buyers are individuals and startup teams willing to exchange screens and meetings for less typing; sellers are OpenAI, turning continuous context into subscription premiums. Beneficiaries are platforms with existing desktop clients, screenshot pipelines, and connectors; those under pressure are chatbots that require users to paste background for each conversation, and middle layers that rely on prompt engineering for income. Funding is shifting from per-interaction billing to pricing based on "whether allowed to watch continuously." Real transactions occur in authorization pop-ups, not at model launch events.
Source: Public Information
ABAB AI Insight
Altman redefines capability leaps from IQ to memory. The notion of a generation model ties the yet-to-be-delivered persistent screen function to the existing scaling curve. Chronicle is already uploading screenshots, indicating that the engineering path is not a fantasy: first watch, then summarize, then write into local memory files. He uses the scenario of startup CEOs unable to read all feedback because they are high-value, highly fragmented, and willing to pay users. Not making decisions for you is to pass compliance and responsibility checks; always watching is the product.
The capital path is turning operating system permissions into a model moat. Once email, Slack, and screens are connected, the switching cost is no longer just changing a URL, but replacing entire work memories. OpenAI first captures screenshots on the desktop, then uses the next generation model to digest this messy context. The motivation is to lock in human-computer interfaces before API prices are diluted by open-source. Resources are shifting from prompt tools to permission management, local summary storage, and server-side visual pipelines. Anthropic's computer usage and other assistants' meeting records are competitors in the same race.
A similar structure is Google turning Gmail into a search entry point, and Microsoft tying Office to the operating system. Search once relied on web crawling; the next generation of assistants relies on you feeding your life into it. The industry phase belongs to control: model IQ continues to rise, the watershed shifts to who is allowed to forget to zero. The prompt competition will depreciate because background no longer requires manual transport.
Structural judgment indicates a transfer of pricing power. The mechanism is that context shifts from user labor to system attachment. Whoever can make "watching you" a default layer that can be turned off within six months will collect workflow rent. The psychological barrier is the last dam: whether one is willing to let a non-forgetting system read an entire screen. Computing power determines how deep the model can think, while authorization determines how much it can remember.
ABAB News · Cognitive Laws
- The next generation product does not increase IQ, but increases observation rights.
- Once background grows automatically, prompts will depreciate.
- Memory is closer to the true switching cost than computing power.