It's been a week since ChatGPT Agent launched, and the internet is losing its mind. It’s an AI that doesn't just chat, but acts, and completes tasks.
If you’ve been online, you’ve probably seen the mindblowing demos of the agent doing wedding shopping👇
But for product founders and leaders building the next wave of tech, the real question isn't "What cool things can it do?"
It's "How does this fundamentally change everything we build?"
This Week in Products, we move past the initial hype to think critically about this new era of task-capable AI agents. We'll explore what this means for your roadmap, your UX, and your entire product strategy.
Before we dive in…
Let's cut through the noise. In the simplest terms, OpenAI has given ChatGPT a virtual body. The Agent can operate its own "virtual computer," equipped with a suite of tools to take real actions on your behalf.
As OpenAI puts it, they’ve bridged "research and action."
They have levelled-up from just a smarter chatbot that can research the internet to a "unified agentic system" that combines the capabilities of two previous tools:
Operator: Excelled at browsing and interacting with websites (clicking, scrolling, typing).
Deep Research: Excelled at analyzing and synthesizing vast amounts of information.
By combining them and adding new tools, the ChatGPT Agent can now:
Browse the web like a human, using a visual browser to click, scroll, transact, or fill out forms.
Run code in its own terminal to manipulate files and analyze data.
Use your apps like Gmail and Google Drive through "Connectors" to access your data (with your permission, of course).
Build artifacts like full reports and editable PowerPoint or Excel files.
This ability to fluidly shift between reasoning and action, using the right tool for the job, is the game-changer.
Zooming In
For those of us who appreciate the engineering, the performance metrics OpenAI released are genuinely impressive.
On benchmarks like DSBench (data science) and SpreadsheetBench, the agent significantly surpasses human performance. On SpreadsheetBench, it scored 45.5% accuracy compared to Copilot in Excel’s 20.0%. It also set a new state-of-the-art score on BrowseComp, a benchmark for finding hard-to-find information on the web.
What do these benchmarks tell us? This isn't a brittle RPA script. It's a powerful reasoning engine capable of tackling complex, multi-step, asynchronous workflows.
But as with any 1.0 product, especially one this ambitious, the reality on the ground is more nuanced.
Early tests from ZDNet found only one in eight multi-step jobs completed without a single hallucination. A writer for WIRED watched it take 26 minutes to generate a mediocre pitch deck.
OpenAI’s own team admits they are "optimizing for hard tasks," not speed. You're meant to "kick something off in the background and then come back to it."
This frames the agent not as a real-time assistant, but as a powerful, asynchronous digital worker. And that distinction is crucial as we think about how to build for it.
The big picture
My true intent with this newsletter is to remove the wow factor of ChatGPT Agents and other task-capable AI Agents that are soon to follow. I want to reframe what these developments really mean FOR the builders… NOT for everyday casual users.
On that note, I want to start by asking the right question.
This is where product thinking comes in. We need to look past the slick demos and confront some hard problems while building task-capable AI agents.
Here are THREE Critical Questions We Should Be Asking Ourselves:
Unless you have a multi-million dollar budget for product development, the answer is to go Vertical. I think the race to build a general-purpose "do anything" agent is the game played by the behemoths like OpenAI and Google and xAI.
The Musks and Sam Altmans of the world are targeting the massive horizontal TAM that their budgets allow.
The opportunity for the rest of us is to build highly-specialized, reliable agents for specific, high-value vertical workflows.
For example:
Think in terms of "job roles," not "tasks." Instead of an agent that "creates a presentation," think of an agent that acts as a "Junior Marketing Analyst." It connects to your Google Analytics, Salesforce, and social media ad accounts to produce a weekly performance deck, complete with analysis and recommendations, formatted in your company's template.
Your Value Prop: The agent's value isn't its intelligence (that's a commodity from OpenAI). It's the reliability, domain-specific knowledge, and workflow orchestration you build around it.
Giving an AI the keys to your users’ digital life introduces novel risks. As we mentioned earlier, ZDNet benchmarks found only one in eight multi-step jobs were completed without hallucination using ChatGPT Agent.
Sam Altman himself urged caution, saying this is "...not something I’d yet use for high-stakes uses or with a lot of personal information until we have a chance to study and improve it in the wild."
OpenAI flags higher risk and slower runtimes when the agent juggles multiple tools. OpenAI knows this and has built-in safeguards like requiring permission for "irreversible" actions and a "Watch Mode" for sensitive sites.
As Box CEO Aaron Levie warns, “The reason why you probably won't see a complete compression of software where it's just a an agent and a database is because there's a lot of logic in the workflow and in the the company's specific sort of business process that needs to be built into that and around that database... The agent will make a mistake 1% of the time and it will share the wrong thing with somebody or open up an access privilege to the wrong person."
In business, that 1% can be catastrophic.
This is perhaps the most critical question.
The most astute leaders whom I’ve spoken to have told me that agentic UX is fundamentally unsuited for discovery and choice (even today).
It excels at transactional tasks where the outcome is well-defined ("Find flights from Chennai to Sri Lanka next Tuesday and book the cheapest one").
But what about ambient tasks? For example, browsing for a new pair of shoes when you’re not sure what you want, or idly scrolling through Myntra. The joy is in the exploration and serendipity.
An agent that asks you for specific criteria completely misses the point.
In an article I read in Wired magazine, the author says this about the new ChatGPT Agent. "…the decision arc is heavily skewed towards a range of choices, and agentic UX sucks at displaying choice."
This launch feels like a new chapter in software. The challenge for us is to move beyond the novelty, understand the deep-seated limitations, and identify the genuine opportunities to build valuable products.
How do we start doing that?
Think of this as a new layer emerging in the software stack: the Agentic Layer.
It sits between the user and the application layer. As product leaders, our job is to figure out how our products will interact with this layer.
Will we offer our own specialized agents?
Will our apps become "tools" that other agents can use via APIs?
These are the questions on my mind.
And I think the products that will win in the next decade will be those that learn to effectively build for these agents, not just with them.
📬I hope you enjoyed this week's curated stories and resources. Check your inbox again next week, or read previous editions of this newsletter for more insights. To get instant updates, connect with me on LinkedIn.
Cheers!
Khuze Siam
Founder: Siam Computing & ProdWrks

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.