THE VERIFICATION LAYER FOR COMMERCECHECKED DAILY · SEPTEMBER 23, 2026
Product.ai/Research/My warning to other AI-native operators: AI amplifies both your strengths and weaknesses
EssayThe Operator's Codex · No. 18
My warning to other AI-native operators: AI amplifies both your strengths and weaknesses
Research suggests AI tends to amplify both your strengths and your weaknesses, and may offer different benefits and pitfalls depending on your level of experience in a domain. Product.ai founder Michael Quoc calls this the Amplification Paradox, and this article covers six ways to make it work for you.
In February, Anthropic’s CEO Dario Amodei told an interviewer that comparative advantage is more powerful than people expect. Even if you do only 5% of a task while AI does the other 95%, that 5% “gets super amplified and levered,” and you become 20 times more productive. 1
As CEO of Product.ai, a profitable AI-native company, I agree with him, and I think that AI can help most people learn faster and do better at their jobs.
So I was surprised to read about a study in which 640 small businesses were given an AI advisor, and its impact was basically nothing. 2
But when I took a closer look, I realized that AI actually did have a measurable effect. It improved the performance of the most successful businesses by 15% while the weakest businesses performed 8% worse.
This is a different result than what I’ve seen in other studies, like the customer-support research I cited in my recent article on AI amplification that showed the least experienced employees get the most benefit from AI. 3
So, which one is it? Is AI more helpful to experts or to novices?
The answer is both, and neither.
The idea that AI can offer both harm and benefits to the same person based on how they use it is what I call the Amplification Paradox; Amodei’s math only works if you choose the right 5% of a task to work on yourself.
Employers who understand this can design much more effective AI-enabled processes and systems that enhance their people’s capabilities without running into issues. At my company, we’ve built the lessons of the Amplification Paradox into our AI-enabled collaboration platform and the processes that support our work every day.
While we’re still learning as we go, we find that AI amplifies people and results, as long as you use it the right way.
The current research, and my own day to day experience running an AI-native team, make it clear that AI’s harms and benefits depend on who is using it and in what context.
The type of AI usage that is most helpful for experts can be ineffective or even harmful for everybody else.
Across the studies in this article, AI suggestions tend to help people new to a domain when they execute a documented task and to help experienced people when they make a judgment call.
AI-human partnerships succeed when the human is an expert and fail when they’re not
When people and AI work together in pairs, they are most effective when the person is knowledgeable and experienced and least effective when they’re not. In fact, an inexperienced person working closely with AI, and given minimal instruction, will often achieve worse results than the AI alone.
A meta-analysis of 106 studies covering combined AI-human performance showed that AI-human pairs did better than either one alone when the human was the stronger performer and worse than the AI alone when the AI was the stronger performer. 4 The studies covered a wide range of tasks, and the pattern held on average across them.
At my company, we build guardrails, training, and instructions into our AI-enabled collaboration platform, so nobody, not even our newest and most junior people, is led by AI or receives outputs that are not grounded in our internal best practices.
Structured, AI-generated instructions are most helpful to beginners.
Novices and weak performers get the biggest performance benefit from structured AI help, while experts’ performance can actually decline. This is likely because AI recommendations represent the average accepted industry practices and published training materials.
In a recent field trial at Alibaba covering 5,940 new service agents, about half of whom received AI-generated drafts to guide their interactions with customers, the lowest performers used the AI drafts 39% of the time and raised their customer ratings by about 0.8 points on a five-point scale. Meanwhile, the top agents used the AI drafts 18% of the time and their customer ratings fell by about 0.3 points. 5
A classic consulting industry study produced similar findings. When 758 consultants at Boston Consulting Group used GPT-4 on tasks the model handled well, the consultants who had scored in the bottom half beforehand improved by about 43%, and the top half improved by about 17%. 6
I believe you can get the benefits of AI guidance and support for new employees without stifling experts by simply allowing more experienced people to override or ignore AI task suggestions.
AI assistance can give people false confidence
People are poor judges of whether AI has made them more effective, and AI use often increases their confidence faster than it raises their actual performance.
In a 2026 study, 246 people solved 20 logic problems from the Law School Admission Test with ChatGPT’s help. They scored about 13 out of 20 and guessed they had scored about 17. People taking the same test without AI also overestimated themselves, but by a much smaller margin, and the more AI-literate the participants said they were, the larger their overestimate. 7
Similarly, at Norway’s labor and welfare agency, researchers followed 39 developers over two years and 26,317 commits. Most of the 25 who used GitHub Copilot said they felt more productive, yet their measured commit activity did not change after adoption, and the correlation between how productive they felt and what the commit data showed was close to zero. 8
I’ve seen this myself many times, and this is why my organization carefully tracks business outcomes related to the tokens we’re consuming.
Experts can be subject to deskilling when they use AI, or any assistive technology, too much
As experts like doctors and engineers begin to use AI every day, we’re starting to see a potential harm of AI use called deskilling. This is when your own unassisted ability, after months of leaning on the tool, begins to decline.
For example, after four Polish clinics adopted regular AI assistance, doctors performing colonoscopies without AI found precancerous polyps in 22.4% of procedures, down from 28.4% before the AI arrived, a drop of about a fifth among experienced endoscopists. 9
Of course, it’s also important to remember that deskilling isn’t unique to the introduction of AI. Drivers experienced the same thing when GPS was introduced. Among 50 regular drivers, heavier GPS use went with worse unaided spatial memory, and when 13 of them were retested three years later, the hours of GPS they had used in between predicted how far their spatial memory had fallen. 10
I believe the root cause of deskilling in every case is simply lack of practice.
Uncritically accepting AI answers is harmful for everyone
AI suggestions can be useful when they’re right, but destructive when they’re wrong and no one questions them. This is common sense, but it’s shocking to me how often experts accept AI recommendations uncritically.
For example, in one study, 27 radiologists were asked to read mammograms and given what they were told was a suggestion from an AI system; the researchers then gave them a mix of helpful and harmful suggestions. The incorrect AI suggestions caused them to read mammograms with a much lower degree of accuracy, and this effect varied by experience.
Correct readings dropped to 19.8% for the least experienced radiologists, 24.8% for the moderately experienced, and 45.5% for the most experienced. But the average accuracy was about 80% when the AI suggestion was right. 11
This is why actively questioning and even trying to disprove AI outputs is baked into my company’s culture. AI-native does not mean uncritical acceptance.
When 27 radiologists reading mammograms were given a wrong AI suggestion, correct readings fell to 19.8% for the least experienced, 24.8% for the moderately experienced, and 45.5% for the most experienced, against an average of about 80% when the suggestion was right. Source: Dratsch et al., Automation Bias in Mammography.
How to make AI work for you, not against you
If you want to be successful with AI, it’s more important to look at how you’re using it than to try to pick the best model.
My company’s switch to an AI-first operations model was transformational, but we didn’t arrive there by simply giving people access to AI. Instead, we thought carefully about which people and processes would benefit most from AI as well as the dangers inherent in AI’s ability to give persuasive, but wrong, answers.
The six practices below are supported both by research and drawn from how my team and I tackle outsized challenges every day:
1. Determine if you’re executing a task or making a decision.
Every time you use AI, ask yourself this question: “Am I executing something with known steps, or making a decision?” For execution, let the AI give you specific instructions or build a rough draft of your work.
But, if you’re making a decision, look at AI’s response as a single data point, evaluate it critically, and consider whether that decision is consistent with your organization’s existing rules and values. If you’re uncertain, consider testing an AI answer adversarially by asking another model (or even a new session with the same model) to tear it down.
How we do it ourselves
The split between tasks and decision-making described above is written into how Product.ai operates. Our AI agents execute work around the clock, but they operate within strict rules and guardrails designed and maintained by our people.
We capture them in three kinds of plain markdown documents, which reside in the same GitHub repositories as the work.
Kernels are the constitution. A kernel describes a rule the whole company runs on. For example, one of ours says that nothing moves from a person’s private workspace into the shared company space unless that person publishes it. An agent can draft a change to a kernel and open a pull request for it, but the change takes effect only when a person approves it.
Drivers are the specs under each kernel. A driver says how one part of the system has to work. The driver for a category page on one of our shopping sites, for example, lists the structured data the page must publish and requires the content to be in the page’s HTML rather than loaded by a script, so that search engines and AI agents can read it. This means a teammate can ship a driver without waiting on me, and every driver has to name the kernel that covers it, or the system flags it as drift.
Code is what the agents write. In the example above, that means the page itself and the script that builds it. An agent can build a prototype with no rule at all. When a pull request changes one of the rule documents, an automated check looks for a record that the change went through the rule process, or an explicit override stating the reason, the risk, and the fix.
We also use teams of AI agents to research questions for us, such as whether the claims in an article like this one hold up. When a team reaches a conclusion, a second agent that never saw its work is handed the conclusion and told to break it.
2. Recognize that AI is not one-size-fits-all.
Let your people access AI in the form that is easiest and most natural for them to work with. And use AI to make company knowledge and best practices immediately available to newer and less experienced employees while giving more senior people the ability to disregard AI instructions altogether.
How we do it ourselves
Everyone at Product.ai works within the same set of AI models and the same constitution. But different roles work with different tools, and our AI-based collaboration system includes automated prompts designed to help all of us, at every level, stay on mission.
The right tools for the job. Engineers tend to work in Claude Code, in the terminal, and people in non-technical roles more often work in Claude Cowork, which behaves more like a chat window. But both read the same company files and run the same shared skills.
Required grounding. New employees’ first sessions with AI are informed by a global instruction file that instructs AI to ground itself in the relevant company context before any work starts. And if a session starts substantive work with no grounding receipt on disk, the system reminds the person once, and only once.
The override button. Experienced operators can deviate from any default, but they’re also responsible for working with AI to document these deviations, including the reasoning and the potential risks behind them.
3. Formulate your own answer before you ask AI.
Encourage everyone in your organization to try answering important questions themselves, and define their own point of view, before asking AI. This exercise can help prevent deskilling and reinforce the mindset that AI answers should not be accepted without question.
How we do it ourselves
Before any of us ask the AI a question with real stakes, we write down what we think the answer will be first. Examples include:
Essays start with a position. Every piece under my byline starts with a position I have written down; my team and I use AI to test it. This one started as a thesis I recorded in July.
Agent runs start with a person’s definition of done. Before any agent team runs for any reason, a person writes what the finished result looks like, in a file the agents read at the start of every cycle and cannot edit.
Research assignments start by defining what would prove or disprove the claim. When we give a question to a team of AI research agents, my team first writes down three or four conditions the claim has to meet, and for each one, what evidence would confirm it and what would refute it. The agents’ final report is graded against that list, and the agent that did the work may not change its score.
4. Default to doubt, and budget time for verification.
Treat every output as wrong until you have looked for mistakes. The better the AI gets, the fewer errors it makes, and the more tempting it is to stop checking its work, which is when the remaining errors get through.
Also, commit to fact-checking in advance and budget it into your work; assume you’ll need to read every AI output carefully before it’s published or used as the basis for a decision.
How we do it ourselves
Agents check each other’s work first, and then people check their work:
Nothing reaches production without a human review. Agents work in sandboxes, and they are not allowed to make changes to our production environment without a person’s approval. Our longest-running agents write everything to a staging area. Only human team members can add work to our canonical library.
Every data point is double-checked. When we use AI for research, we run an automated review of our most promising results. And we always ensure the reviewing model is from a different family from the research model; that way, we don’t have Opus or ChatGPT judging itself.
Content is marked up before it ships. My content team always carefully reviews the articles, papers, and studies they publish, both using agents and through manual checks. We build this time into our production schedules.
5. Treat every AI recommendation as a hypothesis with a test and a date.
Write down what you expect to happen if you follow AI’s recommendation, run the smallest test that could prove you wrong, and set the date you will look at the result. Then revisit the prediction when you have an outcome.
This will help you understand how good your own judgement is and identify areas where you may be more likely to accept faulty AI answers.
Across 759 firms in four randomized trials, founders trained to work this way were about 10 percentage points more likely to shut down a project that was not working, against a sample average of 34%, and they made one decisive pivot rather than many. 12
How we do it ourselves
Our goals at Product.ai are documented with a test and a clock, and when the clock runs out we look at the evidence and determine if we met the goal as well as what happened along the way.
The Test. Every outcome includes an evidence line and an explanation of how it can be verified, such as three dated records from outside developers showing that each got a working key in under five minutes. My own outcome for the quarter is one goal tied to a number, and it’s the first thing I look at each day.
One number each. Several of the people who run parts of the company have signed on to one number each, and they report on it every week in their own words. “Not yet, blocked on this” is an acceptable report.
6. Let people be the final judge of your work.
AI has made it quick and easy to build a draft or a prototype, and you can even get feedback from AI judges. But I believe you get better, more relevant feedback from the people who will be reading your draft or using your product.
In my experience, it can be very tempting to keep iterating on a product with AI giving continuous praise, but it’s usually a waste of time. Other people are much more likely to tell you what’s wrong with your product so you can rethink or fix it.
How we do it ourselves
At Product.ai we run a monthly event we call Golden Hour, and in August 13 guests saw a prototype of our shopping confidence card for the first time. The card is what our product shows when it answers a shopping question. Each guest named a purchase they had second-guessed, viewed a chatbot’s answer first, and then saw our answer card. 13
Each answer card also came with a link that opened the evidence behind it. We expected everyone would be eager to click on that link.
But, surprisingly, no one did. This gave our product team something important to think about.
At Golden Hour in August, 13 guests saw a prototype of the shopping confidence card, each with a link that opened the evidence behind the answer, and not one of them clicked it. Source: Product.ai Golden Hour, August 2026.
AI amplification won’t work without people in the lead
AI amplification is not the same thing as asking people to “use AI.” If you want to amplify your people with AI, you need to offer it within a framework that includes instructions, process guidance, and guardrails.
You also need to build in flexibility, so your most experienced domain experts can reject automated guidance and guardrails when they conflict with their best judgement.
And I strongly recommend establishing a culture of skeptical curiosity in which AI recommendations are not treated as facts.
For my team, there is no contradiction between being an AI-first company and being skeptical of AI answers. We all understand the Amplification Paradox, and we’ve designed our work so it mostly works for us, and not against us, every day.
Are you thinking about how to amplify the people in your organization with AI while avoiding the potential harms? I’d love to hear your thoughts.
Notes
Dario Amodei, interviewed by Nikhil Kamath, “The AI Tsunami is Here & Society Isn’t Ready,” People by WTF, episode 18, recorded February 24, 2026. https://www.youtube.com/watch?v=68ylaeBbdsg
Nicholas Otis, Rowan Clarke, Solène Delecourt, David Holtz, and Rembrand Koning, “The Uneven Impact of Generative AI on Entrepreneurial Performance,” Management Science, published online July 2026. https://doi.org/10.1287/mnsc.2024.06909
Erik Brynjolfsson, Danielle Li, and Lindsey Raymond, “Generative AI at Work,” Quarterly Journal of Economics 140, no. 2 (2025): 889-942. 5,172 customer-support agents. https://academic.oup.com/qje/article/140/2/889/7990658
Michelle Vaccaro, Abdullah Almaatouq, and Thomas Malone, “When Combinations of Humans and AI Are Useful,” Nature Human Behaviour, 2024. https://arxiv.org/abs/2405.06087
Xiao Ni, Yiwei Wang, Tianjun Feng, Lauren Xiaoyuan Lu, Yitong Wang, and Congyi Zhou, “Generative AI in Action: Field Experimental Evidence from Alibaba’s Customer Service Operations,” working paper, July 2026. Rating changes are the effect of being given access, by pretreatment performance quintile. https://arxiv.org/abs/2603.29888
Fabrizio Dell’Acqua et al., “Navigating the Jagged Technological Frontier,” Harvard Business School Working Paper 24-013, 2023; Organization Science 37, no. 2 (2026). The 43% and 17% figures are as reported by the authors in the companion BCG publication. https://www.bcg.com/publications/2023/how-people-create-and-destroy-value-with-gen-ai
Daniela Fernandes et al., “AI Makes You Smarter But None the Wiser: The Disconnect Between Performance and Metacognition,” Computers in Human Behavior 175 (2026). https://doi.org/10.1016/j.chb.2025.108779
Viktoria Stray, Elias Goldmann Brandtzæg, Viggo Tellefsen Wivestad, Astri Barbala, and Nils Brede Moe, “Developer Productivity With and Without GitHub Copilot: A Longitudinal Mixed-Methods Case Study,” HICSS 2026. https://arxiv.org/abs/2509.20353
Louisa Dahmani and Véronique Bohbot, “Habitual Use of GPS Negatively Impacts Spatial Memory During Self-Guided Navigation,” Scientific Reports 10 (2020). https://doi.org/10.1038/s41598-020-62877-0
Thomas Dratsch et al., “Automation Bias in Mammography: The Impact of Artificial Intelligence BI-RADS Suggestions on Reader Performance,” Radiology 307, no. 4 (2023). https://doi.org/10.1148/radiol.222176
Arnaldo Camuffo, Alfonso Gambardella, Danilo Messinese, Elena Novelli, Emilio Paolucci, and Chiara Spina, “A Scientific Approach to Entrepreneurial Decision-Making: Large-Scale Replication and Extension,” Strategic Management Journal 45, no. 6 (2024). https://openaccess.city.ac.uk/id/eprint/32437/
Product.ai Golden Hour, West Los Angeles, August 5, 2026.
Cite This Research
BibTeX
@misc{quoc-amplification-paradox-2026,
title = {My warning to other AI-native operators: AI amplifies both your strengths and weaknesses},
author = {Quoc, Michael},
note = {Personal essay published by Product.ai Research},
howpublished = {\url{https://product.ai/research/amplification-paradox/}},
year = {2026},
month = {09}
}
APA
Quoc, M. (2026, September 23). My warning to other AI-native operators: AI amplifies both your strengths and weaknesses [Essay]. Product.ai Research. https://product.ai/research/amplification-paradox/
MLA
Quoc, Michael. "My warning to other AI-native operators: AI amplifies both your strengths and weaknesses." Product.ai Research, 23 Sept. 2026, product.ai/research/amplification-paradox/.
M
Michael Quoc
Founder & CEO
Founded Product.ai (formerly Demand.io) in 2009 with a conviction that commerce data should be verified, not assumed. Built SimplyCodes into the leading coupon verification platform in the US, bootstrapped, profitable, and competing against billion-dollar acquisitions. Leads the company’s AI strategy, its patent-pending verification method, and the transition from coupon verification to full-spectrum commerce intelligence. Maintains 100% ownership because sovereignty and truth both require no conflicts of interest.