Framework AI Agency Case Studies: How to Evaluate the Proof
← All Resources
Framework

AI Agency Case Studies: How to Evaluate the Proof

How to read, pressure-test and request proof before you hire an AI agency

A credible AI agency case study names a verifiable client and states five things: the baseline period, the spend range, the creative volume tested, the time window and the metric that moved. If any of those are missing, treat the headline number as marketing, not evidence.

This guide shows you how to read a case study, which metrics to trust, which red flags to walk away from and how to verify results before you sign. It is written for buyers who already have a shortlist and need to separate proof from polish.

What makes an AI agency case study credible?

Case study credibility is the degree to which you can check a claimed result yourself. A credible case study gives you enough detail to ask a sharp question and get a straight answer.

Use this five-part checklist on every case study you read:

  1. Named or verifiable client. A logo and a reference contact, or at minimum a brand whose ads you can find in the Meta Ad Library.
  2. Baseline period. The CPA, ROAS or CTR before the agency started, over a stated date range.
  3. Spend range. Even a band such as "$30K-$50K per month" tells you whether the result holds at your budget.
  4. Creative volume. How many variants were tested, not just how many won.
  5. Time window. The start and end dates, so you can rule out seasonality and platform shifts.

Score each case study out of five. Anything below four should not carry weight in your decision.

How do you read the baseline behind a case study's results?

The baseline is the pre-engagement performance of the account, usually measured as CPA, ROAS or CTR over a fixed period. A percentage improvement means nothing without it, because "up 40%" is easy to reach from a bad starting point.

Here is an illustrative before and after. The numbers are hypothetical, built to show the method:

Metric Baseline (90 days before) After (90 days) Change
Monthly spend $40,000 $42,000 +5%
Creative variants tested 12 96 +700%
CPA $58 $41 -29%
Blended ROAS 2.4x 3.1x +29%

This one holds up. Spend is roughly flat, so the gain did not come from buying more volume. The baseline covers a full quarter, and the variant count explains the mechanism.

Now compare a weak version: "We cut CPA by 40% for a leading DTC brand." There is no baseline, no spend, no window and no name. You cannot tell whether the brand started at $200 CPA or $30.

Always ask what the account looked like in the 90 days before the agency touched it.

Which metrics matter most in AI creative case studies?

The metrics that matter are the ones tied to money and to learning speed: CPA, blended ROAS, creative win rate and cost per tested concept. Diagnostic metrics such as hook rate and thumb-stop rate explain why a result happened, but they are not results themselves.

Metric What it tells you Why it matters
CPA Cost to acquire one customer The core efficiency number
Blended ROAS Revenue per dollar of ad spend and production cost Captures the real return, including creative cost
Hook rate Share of viewers who watch the first 3 seconds Diagnoses opening-frame strength
Thumb-stop rate Share of impressions that halt the scroll Diagnoses visual pattern interrupt
Creative win rate Share of variants that beat the control Shows how repeatable the process is
Cost per tested concept Total creative spend divided by concepts tested Shows testing efficiency

Judge claimed numbers against a yardstick. In our AI ad creative benchmarks, AI video ads average a 1.2-1.8% CTR on Meta cold traffic, and DTC products under $100 see 2.5-4.5x ROAS. A case study claiming a 6x ROAS on a $60 product is not impossible, but it should come with unusually strong evidence.

For definitions and formulas behind ROAS and blended CPA, see AI ad creative ROI.

How many creative variants should a case study show?

A strong case study shows the volume tested, not only the winner. Creative velocity, the number of new concepts you can test per week, is the real AI-native differentiator. It is also the number vendors most often leave out.

Why it matters: a winner from 4 tests and a winner from 120 tests are very different claims. The first could be luck. The second reflects a process.

Look for these signals:

  • Total variants produced and total variants launched. The gap shows how much the agency filters.
  • Win rate. If 10 of 100 variants beat the control, that is a 10% win rate you can plan around.
  • Weekly cadence. How many new concepts entered testing each week.

Our data shows AI production is 5-10x faster than traditional production, and that AI ads fatigue 15-20% faster. Both facts point the same way: a good AI agency should be shipping fresh variants continuously, and its case studies should prove it.

The Social Briefing

A weekly briefing on what's working in social -- trends, frameworks, and real campaign data. Delivered to LinkedIn.

Subscribe

What red flags signal a weak or inflated case study?

A weak case study hides the details you would need to check it. Watch for these six red flags:

  1. A single cherry-picked ad. One hero result with no mention of what else was tested.
  2. No spend disclosed. A 5x ROAS on $2,000 is not the same as 5x on $200,000.
  3. Vanity metrics. Impressions, views and likes without CPA or ROAS.
  4. An unnamed "leading DTC brand." If the agency cannot name the client or offer a reference, ask why.
  5. No time window. Results without dates cannot be checked for seasonality or a promotional spike.
  6. Platform-driven gains claimed as agency wins. Meta and TikTok change their algorithms often. If CPA improved account-wide in the same month, the agency may not deserve the credit.

One red flag is a question to ask. Three or more is a reason to move on.

How do you verify an AI agency's results before signing?

Verification takes three steps: a reference call, an ad library check and read-only screenshots of the platform data. Do all three for your top two candidates.

Step 1: Run a reference call. Ask the agency for two clients, ideally one at your spend level. Ask what the baseline was, how many variants shipped per week and what the agency did when a test failed.

Step 2: Check the Meta Ad Library. Search the client's page in the Meta Ad Library and look at the creative the agency claims to have produced. You can see how many ads run, how long they have been live and how often new creative appears. Long-running ads usually indicate winners. A page with only a handful of ads contradicts a claim of high-volume testing.

Step 3: Request read-only screenshots. Ask for platform screenshots covering the full test window, with the date range visible. Compare CPA, spend and conversions against the case study numbers. If the agency will not share them, even redacted, that is informative.

For the wider evaluation process, including scorecards and pilot design, read how to choose an AI creative agency.

What should you ask an AI agency for beyond published case studies?

Ask for the raw operating data behind the polished story. Published case studies are selected. Working records are harder to curate.

Request this list:

  • Raw test logs. Every concept tested, launch date, spend and outcome.
  • Win/loss ratio. How many variants won, lost or were inconclusive.
  • Revision rounds. The average number of edits before an asset was approved.
  • Human-in-the-loop QA process. Who reviews each asset, what they check and how often they reject work.
  • Failed tests. One example of a concept that flopped and what the agency learned.

An agency with a real testing system will have these records ready. An agency without one will offer a fresh slide deck instead.

How do you compare AI agency case studies against in-house or traditional results?

Normalize every result to cost per winning concept. This puts agencies, freelancers and in-house teams on one scale, regardless of how much each spent or how many ads each shipped.

The formula is simple:

Cost per winning concept = total creative spend / number of variants that beat the control

Here is an illustrative comparison, using hypothetical numbers:

Team Creative spend Variants tested Winners Cost per winning concept
Traditional agency $12,000 8 1 $12,000
In-house team $9,000 15 2 $4,500
AI agency $6,000 60 6 $1,000

Production cost also shows the gap. Our benchmarks put human UGC video at $500-$2,000 per asset and AI UGC video at $50-$200. More tests per dollar means more chances to find a winner.

Use the same formula on each case study you collect. If an agency does not publish enough data to calculate it, that tells you what you need to know.

Case study scorecard

Use this table to grade any AI agency case study. Copy it into your evaluation doc and score each row as pass, partial or fail.

Criteria What good looks like Red flag
Client identity Named brand or verifiable reference "A leading DTC brand"
Baseline CPA, ROAS or CTR over a stated pre-engagement period Percentage gains with no starting point
Spend range Monthly spend disclosed as a figure or band No spend information
Creative volume Variants produced, launched and won One winning ad shown
Time window Start and end dates "Within weeks" or no dates
Primary metric CPA or blended ROAS Impressions, views or likes
Attribution Notes on external factors such as seasonality or platform changes Platform-wide gains claimed as agency results
Verifiability Reference call, ad library presence and screenshots available Refuses all verification

A case study that passes six or more rows is worth weighing. One that passes fewer than four is a brochure.

Want an audit built on this framework?

If you want to see how your own creative testing would score, book a Social Operator audit. We review your account against these same criteria, using the baselines in our benchmark data, and show you where the fastest gains are.

Frequently Asked Questions

What should an AI agency case study include?

A credible case study names or lets you verify the client, states the baseline period, discloses a spend range, shows how many creative variants were tested and gives the time window. Without all five, you cannot tell a real result from a cherry-picked one.

How do you verify an AI agency's results?

Speak to a reference client, check the agency's ads in the Meta Ad Library and ask for read-only platform screenshots covering the full test period. Match the numbers in the case study against what the platform shows.

What metrics matter most in an AI ad agency case study?

Cost per acquisition, blended ROAS, creative win rate and cost per tested concept matter most. Hook rate and thumb-stop rate explain why an ad worked, but impressions and views alone are vanity metrics.

How do you compare AI agency case studies with in-house results?

Normalize every result to cost per winning concept: total creative spend divided by the number of variants that beat the control. This puts agencies, freelancers and in-house teams on the same scale regardless of spend or volume.

The Social Briefing

A weekly briefing on what's working in social -- trends, frameworks, and real campaign data. Delivered to LinkedIn.

Subscribe

Published by Social Operator -- the AI creative agency for performance brands.

Ready to build your content engine?

See how Social Operator can scale your brand's social content and ad creatives.