← Blog

What to Ask an AI Agency Before Signing Anything

Most AI agencies are easier to hire than they are to hold accountable. The scope is vague by design, the deliverables are defined after the money moves, and “pilot” covers a lot of ground when nobody agreed on what success looks like. These questions exist because we have been on the receiving end of that contract, and because we have written enough of our own to know what the difference looks like in writing.

The AI Agency Category Has a Structural Problem

Since 2024, thousands of digital shops rebranded as “AI agencies” without adding any engineering depth. The phrase means nothing on its own now. One firm might deploy a working tool integrated into your CRM, another might hand you a 40-page playbook and call it an AI strategy. Both call themselves AI agencies.

The gap shows up in the data. MIT research covering 300+ AI initiatives found 95% of organizations saw zero measurable return from generative AI pilots. Gartner forecasts 40%+ of agentic AI projects will be cancelled before 2027. The failure mode is almost always the same: vague scope, no defined success metrics, no accountability checkpoint between “we built a prototype” and “this is running in production.”

The Difference Between a Real Deliverable and a Demo

This is the line most agencies blur. A real deliverable is live, load-bearing, and survives your next software update. A demo, pilot, or proof-of-concept is not a deliverable, it’s a sales tool dressed up as one.

What counts as delivered

Ask the agency to define “delivery” in writing before any SOW is signed. A working deliverable means the tool runs in your actual environment, on real data, with a documented handoff. It means you can use it tomorrow without the agency on the phone.

What agencies call deliverables but aren’t

  • A demo built on synthetic data
  • A pilot that runs in a sandbox environment
  • A strategy document, roadmap, or playbook
  • A prototype the agency hosts on its own infrastructure

None of these are deliverables. If an agency’s milestone schedule ends with “pilot complete” rather than “live deployment with defined success metrics,” that contract will not protect you.

The Questions to Ask Before You Sign

These are direct, specific questions. An agency that hedges, redirects, or can’t answer them clearly is telling you something important.

”What does delivered mean, exactly?”

Push for a written definition. It should reference your production environment, your real data, and a specific date. If they say “delivery is when we hand off the prototype,” stop there.

”How will we measure whether this worked?”

46% of agencies don’t measure AI’s business impact at all. Before signing, get a written success metric, not “improved efficiency,” but a specific number. Reduction in manual processing time by X hours per week. Increase in qualified leads handled by the tool by X%. If they can’t name it before the project starts, they can’t be held to it after.

”Who owns the architecture when this engagement ends?”

Rolling SOWs and lock-in architectures are how agencies create dependency. Ask directly: if we stop working with you tomorrow, can we maintain, update, and extend this ourselves, or do we need you for every change? The answer tells you whether you’re buying a tool or renting access to one.

”Show me a project where you told a client not to use AI”

This is the single most reliable signal. An honest agency turns down work that isn’t the right fit. If they can’t name one, if every client need apparently has an AI solution, that’s not versatility, it’s a sales stance. We scope every AI engagement before any commitment specifically to identify where AI adds no value, get in touch if you want that read on your situation.

”What’s your data readiness requirement?”

83% of organizations plan to deploy autonomous AI agents, but only one in three say their infrastructure is ready. A competent AI agency won’t accept a project without first auditing your data quality, access controls, and integration requirements. If they’re ready to sign before asking about your data, they’re building on assumptions they’ll charge you to fix later.

”What happens when the AI gets it wrong?”

Failure modes are not edge cases, they’re predictable. A serious agency documents error conditions, defines human oversight checkpoints, and builds fallback handling before deployment. If the answer is “we’ll cross that bridge when we come to it,” that bridge will cost you extra.

Red Flags in AI Agency Proposals

Proposal language is where overpromising lives. These are the specific tells to look for.

Indefinite timelines

“We’ll have initial results within a few weeks of launch” is not a timeline. Research on AI capability announcements found 73% of claims used indefinite temporal language, no dates, no milestones, no accountability. A real project plan has specific dates tied to specific outputs.

”Pilot” framing for everything

Pilots are fine for genuinely experimental work. They’re not fine as the default delivery model for a $30,000 engagement. If every phase is framed as exploratory, the agency has pre-written its excuse for underdelivering.

No governance or data plan

Any agency proposal that skips data governance, access controls, and model selection rationale has not thought through the actual build. These aren’t optional, they’re the foundation. The absence of them in a proposal is not an oversight.

Metrics defined after deployment

“We’ll establish success metrics once we understand the system better” is a clause that prevents accountability. Success metrics belong in the contract, before the first invoice.

What Honest AI Agency Behavior Looks Like

The “we’re honest” claim is now so universal it means nothing. What matters is whether the agency’s business model forces honesty, or makes deception structurally convenient.

Fixed-price, defined-scope engagements force honesty because the agency absorbs cost overruns. Custom WordPress development at Designodin works this way, fixed scope, fixed price, no rolling invoices. For AI work, the equivalent is a properly scoped engagement where success metrics and deliverables are defined before any money moves, which is exactly how we approach it. Talk to us about what that looks like for your setup.

Model-agnostic advisory forces honesty because no specific vendor is paying the agency a referral fee. Ask which AI model they recommend and why. If they can’t explain the tradeoffs between Claude, GPT-4o, and open-source alternatives for your specific use case, they’ve picked a default, not a solution.

See how we scope and build this at designodin.com/ai.

FAQ

What is the failure rate for AI agency projects?

RAND Corporation analysis of AI projects found 80.3% fail to deliver intended business value. MIT research covering 300+ initiatives found 95% of generative AI pilots produced zero measurable return. These numbers reflect both enterprise and SMB deployments.

How do I know if an AI agency is overselling its capabilities?

The clearest signal is proposal language. Watch for indefinite timelines (“results within a few weeks”), pilot framing for all deliverables, metrics to be defined post-deployment, and no data readiness requirement. Ask them to define “delivered” in writing, agencies that hedge this answer don’t have a clear delivery model.

What contract language actually protects me?

Insist on written definitions of: what “delivered” means (production environment, real data, specific date), success metrics defined before work begins, ownership of architecture and codebase on engagement end, and milestone payments tied to live deployment rather than prototype completion. Have a lawyer review any SOW where those elements are absent or vague.

Can a small business afford honest AI advisory?

Yes, and the proportional cost of a failed AI project is worse for an SMB than for an enterprise. Large enterprises average $7.2M per failed AI initiative, but they can absorb it. A $40,000 AI engagement that delivers nothing is catastrophic for a 20-person business. Fixed-price scoped engagements with defined success metrics are available at the SMB level, they’re just less common because they require the agency to commit upfront.

What does “model agnostic” mean and why does it matter?

A model-agnostic agency selects AI tools based on your specific requirements, cost, latency, accuracy, data sensitivity, rather than defaulting to one vendor. It matters because many agencies are effectively resellers of a specific platform. If the agency can’t explain why they chose Claude over GPT-4o or a local model for your use case, the choice was made for commercial reasons, not technical ones.

What happens if my data isn’t ready?

A good agency tells you before starting. Data readiness, quality, completeness, access controls, integration points, is a prerequisite, not a Phase 2 problem. If an agency doesn’t audit your data before signing, expect a “data issues” excuse during delivery. The honest answer is to fix your data first and start the AI project after, even if that means a shorter initial engagement.

Before You Sign Anything

Get written answers to those six questions. If an agency can’t or won’t answer them, you’ve learned the most important thing you needed to know.

If you want to talk through what this looks like for your operation, start a conversation. We’ll be direct about whether we can help.