A field guide for business owners

The business
of making
it work.

The demo is the beginning. The real opportunity is a system your business can depend on.

Explore the report
Two different levels of AI impact
80%

report improved
personal productivity

37%

report positive impact
on company earnings

Share of survey respondents. Different questions, not a conversion funnel. Earnings = EBIT.
McKinsey, August 2026. [1]

2026 RESEARCH EDITIONEvidence. Judgment. A useful next step.8 sources ↗

01 / The execution gap

More output.
Now make it
count.

An employee can finish a task faster while the business still waits on the same approval, missing document, or broken handoff.

That is the central distinction in this report: making an individual task easier and improving an entire business process are different achievements. McKinsey’s latest survey reports personal productivity gains much more often than positive company-level earnings impact. It also finds workflow redesign more common among high performers. These associations do not establish what caused their results. [1]

Recent worker-level research adds another useful caution: AI adoption varies substantially even among people doing similar work. A tool being available is not the same as a team routinely using it. [8]

XT3’s recommendation: start with one recurring business outcome, then follow it from the first input to the last handoff. Choose where AI belongs only after you can see where the work actually gets stuck.

A better first question

“Which part of the business should work better by next month?”

01Define the outcome.Measure completed work, not generated output.

02Own the whole flow.Include people, permissions, and exceptions.

03Keep improving.Leave someone responsible after launch.

02 / Read the evidence carefully

The gains are real.
The context matters.

“Does AI improve productivity?” is too broad to guide a buying decision. Ask whose work, which task, which tools, and what was measured.

Measured in customer support

+15%

More issues resolved per hour.

A published study of 5,172 support agents found an average productivity improvement with AI assistance. Effects varied across workers; people retained control of customer conversations. [2]

One firm, a specific tool and workflow. This is evidence that assistance can help, not a forecast for your business.

A finding has a shelf life.

JUL 2025

METR’s trial found experienced developers took 19% longer with early-2025 AI tools on tasks in familiar repositories. [3]

FEB 2026

Follow-up results suggested faster work, but selection and measurement problems made the size of the effect unreliable. [4]

MAY 2026

METR examined self-reported gains while highlighting response bias and the difficulty of translating perceptions into measured value. [5]

These studies use different settings and methods. They are not points on a single productivity trend.

What to do with this

Run a bounded trial on your own recurring work. Compare completion time, corrections, and accepted outcomes against a baseline. Include review time. Keep the version that improves the whole result, even if it uses less AI.

03 / Life after the demo

The first launch
is a starting
line.

Bubble’s analysis found far more deployment activity among apps with payment activity than those without it during their first six months. [7]

It is an association, not a recipe. Paying users might encourage improvements; stronger products might attract paying users; other differences may influence both. Deploying more often does not, by itself, prove value.

The useful operating habit: make a small change, check the customer outcome, and decide what to change next. Leave time and ownership for that loop in the original plan.

Median deploys in the first six months
Apps with payment activity45
Apps without payment activity2

One square = one deploy. Bubble’s platform data; payment activity does not mean profitability. [7]

ShipObserveImprove
The system around the tool

DORA’s 2025 research describes AI as amplifying an organization’s existing strengths and weaknesses. For an owner, that makes clear priorities, useful feedback, and dependable delivery part of the investment. [6]

04 / Put the research to work

Start with a friction.
Finish a workflow.

These are proposed starting points, not measured client results. Choose a recurring problem you can observe, with a useful outcome and a safe fallback.

01

The lead that waits too long.

Start with: a qualified inquiry sitting in an inbox while someone finds the context, assigns an owner, and drafts a reply.

  1. Capture the inquiry
  2. Check required details
  3. Route to an owner
  4. Review the reply
  5. Log the outcome

Use the simplest tool. Let ordinary rules handle routing and deadlines. AI may help summarize the inquiry or draft a response from approved information.

Keep the boundary. A person approves prices and commitments. Missing details go to a review queue; duplicate inquiries do not trigger duplicate messages.

Measure the result. Track time to the first useful reply, qualified inquiries handled, and corrections. Do not count automated acknowledgments as meaningful responses.

02

The report assembled by hand.

Start with: a recurring client update assembled by copying numbers between tools.

  1. Read approved sources
  2. Validate the totals
  3. Draft the narrative
  4. Review exceptions
  5. Publish the update

Use the simplest tool. Calculate totals with ordinary code or spreadsheet formulas. Use AI for a draft explanation only after the numbers reconcile.

Keep the boundary. Mark stale and missing inputs. Preserve source links. Client access must stay limited to the correct client’s information.

Measure the result. Compare preparation time including review, late reports, and corrected figures. Keep a manual version available during the pilot.

03

The answer only one person knows.

Start with: the same internal question repeatedly interrupting a senior employee.

  1. Find approved material
  2. Check access
  3. Answer with sources
  4. Escalate uncertainty
  5. Improve the material

Use the simplest tool. Improve the document or search first. Add an assistant when questions genuinely need synthesis across maintained sources.

Keep the boundary. Apply the reader’s permissions before retrieving content. An unsupported answer should become a question for the owner, not a confident invention.

Measure the result. Review correctness, successful self-service, and interruptions avoided. Give outdated guidance an owner and a review date.

Buy, connect, or build?

Buy when an existing product covers the important job and its limits are acceptable.

Connect when the tools work but information gets lost between them.

Build when a distinctive workflow justifies ongoing ownership, testing, and maintenance.

XT3 judgment: compare candidates by observed friction, expected benefit, reversibility, and total effort. Validate the strongest candidate first. No universal “80/20” split or fixed savings rate is assumed.

The economics of useful work

Time saved is
only the
first line.

A faster draft can still mean an expensive process if someone has to inspect, repair, and re-enter every result.

Agree on what “better” means before the pilot starts. Include review and exception handling in the comparison. Count cash savings only when spending actually falls; reclaimed capacity is valuable, but it is a different benefit.

Benefit
Accepted work completed, usable capacity released, or additional contribution earned
Less
Software, model usage, human review, errors, support, and maintenance
Then compare
The net recurring benefit with implementation effort and the next-best alternative

A decision framework, not a financial forecast. Use your actual costs and measured pilot results; avoid counting the same benefit twice.

05 / Before the business depends on it

Looks good.
Ready to run?

Use this as a conversation with the person building your system. Check an item when you can point to evidence.

A working checklist, not a certification. Checkmarks are temporary and are not saved or submitted.

Operational readiness checklist

A practical first month · suggested sequence

Week 1Observe the work. Set a baseline and choose one outcome.

Week 2Build or configure the smallest usable flow. Test awkward cases.

Week 3Run a limited pilot with a person reviewing consequential actions.

Week 4Compare results. Expand, revise, or stop. Name the ongoing owner.

Planning guidance, not a delivery promise. Access, integrations, data quality, and risk may require a longer schedule.

The XT3 point of view

Build something
you can depend on.

You run the business. I handle the technology. Bring the workflow that keeps getting stuck; I’ll help you question the solution, choose a useful first step, and work out what it takes to own it properly.

Ian Dikhtiar · Hands-on fractional CTO · XT3

06 / Sources & methodology

Follow the evidence.

This report synthesizes public research available on September 8, 2026. It is written for owners considering business software and automation. XT3 did not run a new survey, audit these datasets, or measure client outcomes for this report.

Survey findings, field research, and vendor observations answer different questions. They are not pooled into a single success rate. All chart values come from the named sources; the workflows, checklist, and operating sequence are XT3 recommendations. None predicts the reader’s return.

Bubble’s report informed the editorial format. Its data remains attributed to Bubble. McKinsey’s findings describe a global company sample; the support-agent study is one organization; METR’s experiments concern particular developers and tools. Findings may not transfer to a smaller business or newer system.

Download chart data (CSV) ↓

  1. [1]
    The state of AI in 2026: On the road to ROI ↗

    McKinsey · August 25, 2026

    Global, self-reported survey: 1,719 respondents in 97 nations, fielded May 4–June 8, 2026. GDP-weighted; 36% work at companies above $1B revenue. These are not representative small-business benchmarks.

  2. [2]
    Generative AI at Work ↗

    Erik Brynjolfsson, Danielle Li & Lindsey R. Raymond · The Quarterly Journal of Economics · May 2025 · 140(2), 889–942

    Published study of a staggered rollout to 5,172 customer-support agents at one firm. Uses the published 15% result, rather than the earlier working paper’s 14%. Historical technology and a specific work setting.

  3. [3]
    Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ↗

    Joel Becker, Nate Rush, Beth Barnes & David Rein · METR · July 10, 2025

    Randomized trial: 16 experienced developers, 246 tasks in familiar repositories, February–June 2025 tools. The result is bounded to that setting and period.

  4. [4]
    We are Changing our Developer Productivity Experiment Design ↗

    Joel Becker and colleagues · METR · February 24, 2026

    Follow-up evidence suggests improvement, but selection effects and time-measurement problems prevent a reliable estimate of the current productivity effect.

  5. [5]
    Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity ↗

    METR · May 11, 2026

    Self-reported productivity research. Low response rates and potential selection bias limit generalization; perceived gains are not independently measured business returns.

  6. [6]
    State of AI-assisted Software Development 2025 ↗

    DORA · Google Cloud · 2025

    Software-delivery research: AI can amplify the strengths and weaknesses of the surrounding organization. Used for operational context, not as causal proof of a specific intervention.

  7. [7]
    The Business of Building 2026 ↗

    Bubble · 2026 · data through mid-August

    Vendor analysis of 250,000+ live, deployed Bubble apps. Six-month comparisons use matched observation windows. “Monetizing” means running a payment workflow or calling a payment API; it does not establish profit. Observational, platform-specific evidence.

  8. [8]
    What Work Does Generative AI Do? ↗

    Alexander Bick, Adam Blandin, David J. Deming & Tyler R. Schumacher · NBER · August 2026 · Working Paper 35677

    Nationally representative worker survey linking adoption to occupations and tasks. Working paper; adoption measures do not establish productivity or returns.

Back to the beginning ↑

Notes on this report

Practical follow-ups