What the research says
Published studies of generative AI at work, each shown with its sample, its task and its caveat, and always beside findings that cut the other way. None of these figures is an iSystematic result.
The rule
iSystematic publishes figures from published studies, never its own results. Each figure shows what was measured, on whom, in which task, and the caveat that has to sit beside it.
At least one counter-finding is always shown beside the gains. These figures describe generative AI in studies; they say nothing about the effect of an iSystematic workflow, solution or pack.
Every figure links to its original source. The register is rechecked against the originals each quarter, and a new study is added only after the original has been read.
Where studies found gains
Four studies that measured time, output or quality with generative AI, from controlled experiments to a survey of US workers.
453 college-educated professionals (marketers, grant writers, consultants, data analysts, HR professionals and managers) in a preregistered online experiment; half were randomly given ChatGPT. The tasks were paid, occupation-specific writing assignments of 20 to 30 minutes: press releases, short reports, analysis plans and delicate emails.
Short, self-contained tasks in an experiment, without context-specific knowledge; the authors say this may inflate the estimates. Quality was graded by evaluators.
Noy and Zhang, Science, 2023 ↗5,172 customer-support agents at one firm, a Fortune 500 company that sells business-process software. The assistant was introduced in stages, not in a randomised trial. Task: customer-support chat, measured as issues resolved per hour.
The most experienced and highest-skilled agents saw small gains in speed and small declines in quality. The authors say the findings apply to one AI tool, used in one firm, in one occupation.
Brynjolfsson, Li and Raymond, Quarterly Journal of Economics, 2025 ↗758 Boston Consulting Group consultants in total, in a preregistered experiment, randomly assigned to no AI, GPT-4, or GPT-4 with a prompt-engineering overview. These figures come from the 385 who worked on 18 realistic consulting tasks inside the AI's capability, such as developing new product ideas.
The gains held only for tasks inside the AI's capability. On a task outside it, consultants using AI were less likely to be correct (E3b). One firm, GPT-4 as it was in 2023, and a working paper.
Dell'Acqua et al., Harvard Business School working paper, 2023 ↗Real-Time Population Survey, November 2024 wave: 5,329 US adults aged 18 to 64 in an online panel, targeted and weighted to the Current Population Survey; the authors say it is not a random sample. The time-savings question went only to generative AI users (933 employed users, per the working paper).
Self-reported, from one week's recall: the extra hours a user said they would have needed without generative AI. The time saved was not measured directly.
Bick, Blandin and Deming, Federal Reserve Bank of St. Louis, 2025 ↗Published studies, not our results. Ours will come from documented pilots.
Where studies found the opposite
Two results that cut against the gains: experienced developers who were slower with AI than they believed, and consultants who were less often right when the task lay outside the AI's capability.
16 experienced developers, with about five years on their projects, working 246 real issues in their own mature repositories; each issue was randomly assigned to AI allowed or not allowed. The main tools were Cursor Pro with Claude 3.5 and 3.7 Sonnet, February to June 2025.
Early-2025 tools, and expert users on familiar, complex work. METR says the result does not show that AI fails to speed up most developers. In February 2026 METR gave the 95% confidence interval as 2% to 39% slower, said developers are likely more sped up now, and said its newer data give an unreliable signal because developers opted out of working without AI.
METR (Becker, Rush, Barnes and Rein), 2025 ↗373 of the same 758 consultants, on a business problem-solving task using data and interviews, with the same preregistered random assignment.
The group without AI was correct about 84.5% of the time, against 60% and 70% in the two AI conditions. One firm, GPT-4 as it was in 2023, and a working paper.
Dell'Acqua et al., Harvard Business School working paper, 2023 ↗Published studies, not our results. Ours will come from documented pilots.
In healthcare
One randomised trial of ambient AI scribes, which draft clinical notes. It measured physicians, not front desks.
238 outpatient physicians across 14 specialties, randomised 1:1:1 to DAX, Nabla or no scribe; 72,369 visits over two months at one academic institution, measured as time-in-note per note in the Epic health record.
Single centre; physicians, not front desks; English-only visits. Burnout and workload were secondary measures and were not statistically tested; the authors say any improvement needs confirmation in larger trials. Clinicians reported that notes occasionally contained clinically significant inaccuracies. The scribes were used in only 29.5% (Nabla) and 33.5% (DAX) of visits, and time-in-note leaves out editing inside the vendors' apps.
Lukac et al., NEJM AI, 2025 ↗Published studies, not our results. Ours will come from documented pilots.
In software engineering
One controlled experiment on a single coding task. It is not representative of small-business work, and iSystematic cites it only for enterprise engineering work.
95 professional developers recruited on Upwork and randomised: 45 with Copilot, 50 without. Time ran until the code passed all 12 tests.
A single coding task; code quality was not measured, and time was measured only for those who finished. The 95% confidence interval ran from 21% to 89%. A preprint, not peer-reviewed. The authors work at Microsoft Research and GitHub, which make Copilot, and at MIT Sloan. Not representative of small-business work: iSystematic cites it only for enterprise engineering work.
Peng, Kalliamvakou, Cihon and Demirer, arXiv preprint, 2023 ↗Published studies, not our results. Ours will come from documented pilots.
Adoption in the United States
How many US small businesses say they use generative AI. This measures adoption, not efficiency.
3,870 US businesses with fewer than 250 employees, surveyed online by Teneo Research for the US Chamber of Commerce's Technology Engagement Center, 6 to 26 June 2025.
Adoption, not efficiency. Use was self-reported and the report does not define it. The report gives no weighting, margin of error or response rate, and its publisher is a business advocacy body.
US Chamber of Commerce, Empowering Small Business, 2025 ↗Published studies, not our results. Ours will come from documented pilots.
Adoption in Canada
Statistics Canada's survey of how many Canadian businesses use AI, and what the businesses that use it say it changed.
Canadian Survey on Business Conditions, second quarter of 2025: a stratified random sample of 21,357 business establishments with employees; 9,103 responses, with calibrated weights; collected 1 April to 5 May 2025.
Use in producing goods or delivering services, not any use. It varied widely by industry: 35.6% in information and cultural industries, 31.7% in professional, scientific and technical services, 30.6% in finance and insurance, and 1.5% in accommodation and food services. Businesses with employees only.
Statistics Canada, 2025 ↗Same survey as E8. The base is businesses that used AI (12.2% of all businesses), not all businesses.
Self-reported, over a 12-month horizon. Tasks were reduced to a large extent for 5.3%, a moderate extent for 32.4%, a small extent for 47.2%, and not at all for 15.1%. Employment rose at 4.3%, fell at 6.3% and was unchanged at 89.4%.
Statistics Canada, 2025 ↗Published studies, not our results. Ours will come from documented pilots.
How to read a figure
A figure from a study is only as useful as the three things printed beside it.
The sample
Who was measured, how many, and how they were chosen. A randomised trial, a staggered rollout and a self-reported survey carry different weight.
The task
What people were doing when they were measured. A result on a short writing task says little about long, unfamiliar work, and a coding experiment says little about a front desk.
The caveat
What limits the result, as the source itself says: one firm, one tool, self-reported time, a preprint. Read it before quoting the number.
How iSystematic's own figures will come
iSystematic's own figures will come only from documented pilots, measured before and after: task time including human review, draft acceptance, corrections, escalations, duplicates, failures and cost. This page shows none yet.
We will publish what the measurement shows, including no change or a worse result, and never select only the good results. Each result will state the number of organisations, the task and the period, and will sit beside the published studies, never merged with them. A result is published only with the client's permission.
Where to go next
Pilot
iSystematic's Pilot puts one solution live for one team for six weeks, measured against a baseline from day one: the only route to a published result.
Automation Assessment
iSystematic's Automation Assessment: two weeks, a fixed fee, ten candidate tasks scored, three recommended solutions and a plan to measure the baseline.
How we build
How iSystematic builds: every solution uses the same deposited frameworks. See what each part decides, what the client keeps, and each specification's DOI.
Workflows
iSystematic workflows run in your own Claude, ChatGPT, Perplexity or Grok account: a free card first, then a pack, a setup call and monthly care if needed.
Insights
iSystematic's writing currently lives on nabeelkhan.com as dispatches, beside the books and frameworks it draws on. This page links there instead of copying.