How to Choose an AI Automation Agency: 12 Questions That Reveal Everything

The complete buyer's guide to choosing an AI automation agency: the 12 questions that separate builders from demo artists, the five non-negotiable deliverables, red flags, and how to compare quotes fairly.

Here is the uncomfortable position you are in when choosing an AI automation agency: you are buying something you cannot fully evaluate, from sellers who all sound identical, in a market young enough that credentials barely exist.

Every agency's website says the same six words in a different font. Every demo is impressive, because demos are built to be. And the difference between the partner who quietly compounds your operations for years and the one who leaves you with an undocumented liability is entirely internal: their process, their governance habits, their honesty about maintenance. None of it visible from the sales deck.

The good news: it is extractable. The right twelve questions, asked in the first two meetings, reveal essentially everything, because good agencies answer them with specifics and relief (finally, a buyer who asks), while the other kind pause, generalise and change the subject. This guide is those questions, with what good and bad answers sound like, plus the five deliverables that belong in every contract, the red flags that end conversations early, and the fair way to compare quotes. It is, deliberately, the guide we would want our own prospects to read, including the parts that make our industry uncomfortable.

Before the Questions: Know What You Are Buying

One piece of homework makes every question sharper: know which tier of work you are shopping for. A no-code platform configuration, a grounded assistant, an acting agent with approval gates, and a scheduled loop are different products at different prices with different risks (the [pricing tiers] and the [four jobs of automation] map them). Walk in knowing your workflow, your two-week tally of its hours, and your rough tier, and you transform from a prospect to be pitched into a buyer to be answered. Agencies calibrate honesty to the buyer's literacy; raise theirs on arrival.

Download the 12 Questions Card (free PDF). Two printable pages: all 12 questions with good and bad answers, tick boxes for three vendors, and the five contract deliverables. Get the card →

The 12 Questions

On competence and fit

1. "Walk me through a project like mine, including what went wrong." Good answer: a specific build, its stumbles (a supplier's format change, a threshold tuned three times), and what the fix taught them. Bad answer: a highlight reel with no friction anywhere. Every real project has a "what went wrong" chapter; vendors without one have either done few projects or edited honestly out of the pitch.

2. "Which platforms and frameworks would you use for this, and why those?" You are not testing the answer's contents (that is [the frameworks guide]'s job); you are testing whether the reasoning starts from your workflow's stakes and volumes or from their comfort zone. "We do everything in X" is a hammer describing the world as nails; "for your volume and approval needs, probably X, but the audit decides" is a professional thinking.

3. "What would you NOT automate in my business?" The single most revealing question on this list. Good agencies name things readily (the judgment-heavy, the low-volume, the not-yet-ready) because scoping-out is half their craft. An agency that finds everything automatable has confused your operations with their revenue.

On safety and governance

4. "Show me the tool list and permissions from a past agent build." The [toolbox inventory]: which systems the agent touched, read versus write, with what limits. Its existence proves least-privilege thinking; fluent walking-through proves it is practice, not paperwork.

5. "Where do human approvals sit in your builds, and show me an approval card." Money, commitments and mass communication should pause for one-glance human decisions ([the gate map]). Agencies who design gates show them proudly; agencies who bolt them on describe them vaguely.

6. "Show me yesterday's logs from a live system: actions, checks, catches." The audit trail and the [catch log], live, not screenshotted. This one request separates production operators from demo artists more reliably than any credential, because you cannot fake a boring, populated log.

On honesty and economics

7. "What will this cost to RUN, monthly, including your maintenance, and what happens in month thirteen?" The [month-thirteen question]. Good answer: a running-cost estimate, a maintenance arrangement (10-20% of build annually is the honest range), and named responsibilities. Bad answer: a build price with silence after it, which is not a cheaper offer; it is an incomplete one.

8. "How will we measure whether this worked?" Good answer: a baseline captured before building, a success number agreed with you, a day-90 reckoning ([the ROI method]). An agency that resists measurement is planning to be unmeasurable.

9. "What is the smallest version of this project you would recommend?" Tests for scope discipline. The professionals suggest starting smaller than you asked (one workflow, shadow-run, then expand); the other kind returns a bigger quote than you imagined. Appetite for your money is inversely correlated with respect for it.

On your protection

10. "If we part ways, what do I own and can someone else take over?" The answer should include: repositories and configurations in your accounts, credentials in your names, and documentation good enough for a successor. This single clause, the [documentation clause], protects you from every other risk on this page, including the agency being excellent but getting acquired.

11. "Who exactly will work on this, and how much is subcontracted?" Not a purity test (good agencies use specialists), but you deserve to know whether the impressive person in the sales meeting will ever touch your build, and whose name is on the maintenance rota.

12. "Can we start with a paid audit instead of a build?" The finale, and a filter that costs nothing to deploy. Agencies confident in their value happily start small: an [audit] that maps your workflows, ranks the candidates and prices the options, deliverables you own either way. Agencies that push past the audit toward the big build are telling you their sales process matters more than your sequencing. Believe them.

The Five Non-Negotiable Deliverables

Whatever the project, these five belong in the contract, and their absence is a decision, not an oversight:

  1. Working automation measured against the pre-agreed baseline.
  2. Monitoring included: quality and cost visibility from day one, not an upsell.
  3. Least-privilege access: scoped permissions per system, documented.
  4. Documentation sufficient for a different developer to take over.
  5. A maintenance plan with named owners and the month-thirteen economics in writing.

Any quote carrying all five can be compared fairly against any other such quote. A cheaper quote missing two of them is not cheaper; it is deferred.

Red Flags That End Conversations Early

Guaranteed outcomes before seeing your data. "We'll save you 40%" prior to any audit is astrology with an invoice.

Everything-is-agentic vocabulary. If every product is an "AI agent" including the obvious chatbots, apply [the ten-second test] (does it decide its own steps?) and watch the pause.

Demo-first, workflow-later meetings. Professionals ask about your operations before showing anything; performers open the laptop first.

No maintenance conversation until you raise it. The month-thirteen silence. You now know to break it, and how the answer sounds.

Pressure against the audit-first path. Covered in question 12; worth repeating as a flag because it is the most common one in the wild.

Your industry's compliance handled with a wave. If you are in finance, health, legal or anywhere regulated and the agency's data-protection answer is "all encrypted, all fine," the depth stops there. Real answers name where data lives, who accesses it, and what the [governance] regime looks like.

Comparing Quotes Fairly: The Same-Scope Rule

Quotes are only comparable against identical scope, so write the one-pager first: the workflow, the systems it touches, the boundaries and exceptions, the success number. Send it to two or three candidates, require the five deliverables in each response, and compare three numbers: build, monthly running, and year-one total. Expect legitimate spread ([the who-builds-it factor]: freelancer, specialist, consultancy can quote 1x, 2.5x, 10x for real reasons), and interrogate outliers in both directions: the suspiciously cheap are usually missing deliverables; the startlingly expensive are sometimes selling process you do not need, and sometimes seeing complexity you missed. Ask which.

The Two Meetings: How the Questions Deploy in Practice

Twelve questions do not fit one conversation gracefully, and ambushing a vendor with a checklist reads as hostile when the goal is diagnostic. Here is the deployment that works.

Meeting one: the fit conversation (questions 1, 2, 3 and 12). You describe the workflow (your one-page scope does the talking); they respond with the case study, the tool reasoning, and, if you ask question 3 early, the leave-alone instincts. Close with question 12 (the audit-first offer) and watch the reaction more than the words: relief and a price means partner; a pivot back to the big build means performer. Most disqualifications happen here, cheaply, in under an hour.

Meeting two: the evidence session (questions 4 through 11). For the two or three who passed: a working session where they bring the artefacts: the tool lists, an approval card, live logs, documentation samples, the maintenance rota. This meeting has a tell of its own: good agencies arrive with this material half-volunteered, because buyers who ask for it are the buyers they want. You are not just evaluating answers; you are previewing what working with them feels like when you ask hard questions later, mid-project, when it matters more.

Between the meetings: the reference call. One past client, your questions not theirs: what broke, how fast was the response, what does month thirteen actually cost, would you rehire? Fifteen minutes that validates or vaporises everything the artefacts implied.

The cadence matters as much as the content: vendors calibrate to buyers, and a buyer running this structure gets the agency's A-team behaviour from the first hour, which is, quietly, the thirteenth thing the process buys you.

Frequently Asked Questions

The Takeaway

You cannot evaluate an agency's code, and you never needed to. You evaluate their habits: what they refuse to automate, what their logs look like on a Tuesday, whether month thirteen has a price, what you own at goodbye. Twelve questions extract all of it in two meetings, and the five deliverables hold whoever passes.

Ask them all, including of us. The agencies that flinch have answered; the ones that brighten are your shortlist, because the questions that protect you are the same ones good operators wish every buyer asked.

Bots and Brand Works publishes its answers to all twelve questions, starts every engagement audit-first, and puts the five deliverables in writing as standard. Bring this list to our first meeting, or to anyone's. That is what it is for.

Need Help Implementing AI Solutions?

Bring this list to our first meeting, or to anyone's.

Resources and Further Reading

  • How to Automate Your Business Operations
  • Related: What AI Automation Agencies Charge
  • Related: The AI Automation Audit
  • How Much Does AI Automation Cost? · [AI Agent Governance]( · [Why 40% of AI Projects Fail](
Previous
Previous

How to Automate Your Business Operations: The Complete 2026 Guide

Next
Next

The AI Automation Stack Explained: A Business Owner's Guide to AI Architecture