FIELD NOTES — EVALUATION
How to Evaluate an AI Visibility Agency: The Fintech Buyer's Checklist
Jeff Jenkins · August 3, 2026 · 9 min read

Evaluating an AI visibility agency comes down to six questions covering proof, staffing, methodology, evidence, and measurement, scored the way procurement teams score proposals rather than the way marketing teams read portfolios. Most agency-evaluation advice assumes Google is the only place buyers discover vendors. That assumption is now wrong, and it makes most of the advice wrong with it.
The Advice Is Better Than the Evidence
We pulled the sources Google's AI Overview cites for the two core agency-evaluation queries, "choosing an SEO agency" and "questions to ask an SEO agency," on August 3, 2026 (Ahrefs, US). The AI-assembled answer draws on roughly 35 sources: agency-written listicles, Reddit threads, a Quora answer, Clutch, Forbes, Neil Patel, and ten YouTube videos. One cited source is a "Best SEO Agencies in UAE" roundup, surfaced for a US query.
Zero of those sources are specific to fintech. Zero mention answer engine optimization (AEO), generative engine optimization (GEO), citations, or any AI visibility criterion at all.
The answer engines tell the other half of the story. On August 3, 2026, we ran three prompts through ChatGPT, Gemini, Claude, and Perplexity: how to evaluate an agency for a fintech company's AI search visibility, a request to draft an RFP for those services, and the question buyers actually type: what to ask an SEO agency before hiring one. Three prompts, twelve answers. Here is what we observed.
| Prompt (all run August 3, 2026) | ChatGPT | Gemini | Claude | Perplexity |
|---|---|---|---|---|
| How to evaluate an AI visibility agency | No sources shown | No sources shown | No sources shown | Agency listicles and vendor blogs |
| Draft an RFP for AI visibility services | No sources shown | No sources shown | No sources shown | GEO vendor and agency guides |
| Questions to ask an SEO agency | No sources shown | No sources shown | No sources shown | Google documentation and an agency listicle |
| Named case study with verifiable numbers, in any answer | None | None | None | None |
| Original research cited, in any answer | None | None | None | None |
The advice itself varied by engine. On the familiar hiring question, one engine volunteered no AI search criteria at all, one raised them in a single conditional line, and two raised them unprompted, the same engine-to-engine inconsistency our benchmark found in fintech brand visibility itself. The RFP templates were fluent, but none included pass/fail minimum qualifications, and their scoring weights disagreed wildly.
Across all twelve responses, two things never changed.
- No named case study with verifiable numbers.
- No original research cited.
Where sources appeared at all, they came from agency listicles, vendor blogs, and guides published by generative engine optimization vendors, the same category of vendor the advice exists to evaluate. The evaluation standard has to come from evidence, not recommendations, at the exact moment fintech buyers are building vendor shortlists inside AI engines before they ever visit a website.
Existing advice tells you how to hire an SEO agency. What follows is how to evaluate an AI visibility agency.
What an AI Visibility Agency Actually Does
An AI visibility agency builds the content architecture that gets a company retrieved and cited when ChatGPT, Gemini, Claude, and Perplexity assemble answers about its category. The discipline combines answer engine optimization (AEO), which makes content extractable, and generative engine optimization (GEO), which makes it selected. AEO vs GEO covers the distinction in depth.
The work overlaps heavily with strong technical SEO and content strategy. The difference is the target: inclusion in AI-generated answers, which is binary. There is no page two.
Six Questions Every Agency Should Be Able to Answer
These are written as questions because they are meant to be asked out loud, in the interview, in this order.
1. Can You Show Me Named Case Studies With Numbers?
A verifiable case study identifies the client, the competitive context, the baseline, the measured outcome, and the timeframe. Many public-sector RFPs require measurable outcomes and budgets inside submitted case studies and warn vendors that references will be contacted. Hold agencies to the same standard.
For example, one fintech client of ours entered an engagement at Domain Rating 41, competing against an incumbent at DR 93. The case study names the query won (unified payments platform), the outcome (#1 above Stripe), and the AI visibility result (AI-referred sessions grew from 17 to 477 in a single quarter). Every element can be checked.
A weak answer is a logo wall, results described as "significant improvements," or an NDA that conveniently covers every number.
2. What Do the AI Engines Say About You?
Before the first call, ask ChatGPT, Gemini, Claude, and Perplexity about the agency's own category. If an agency specializes in AI visibility, you should expect evidence that it has applied those techniques to itself.
Apply this test with some calibration. A young specialist firm may not dominate broad prompts yet, and that alone is not disqualifying. What you are listening for is whether the agency can show you its own citation footprint, explain what is working, and explain the gaps. An agency that has never run the query on itself is telling you how it will treat yours.
3. Who Actually Does the Work?
This is the question almost no buyer asks, and in fintech it matters more than the methodology deck. Everyone can describe a process. Very few agencies will name the humans who execute it.
Ask directly: Who writes the content? Are the strategists also the authors? Is the research outsourced? Who interviews your subject matter experts? Who on the team has ever had a draft go through a compliance review?
It is common for public-sector RFPs to handle this with a bluntness worth borrowing: named key personnel with resumes, the percentage of time each person is committed to the account, and pre-approval of every subcontractor. Regulated content written by a generalist content mill fails legal review, which is why content for regulated verticals typically carries a 20 to 30 percent compliance premium. If the agency cannot name who does the work, the premium buys you nothing.
4. Can I Read Your Methodology Before I Buy?
Ask the agency to explain, in writing, how AI systems decide which sources to retrieve and cite, and how its process addresses each stage. Any legitimate methodology qualifies. What disqualifies an agency is having none in writing, or hiding it behind "proprietary process."
The reason is practical, not philosophical: evaluation committees score method as heavily as experience, and you cannot score what you cannot read. An agency confident in its methodology publishes it, because publishing it is itself an AI visibility strategy.
5. What Evidence Supports Your Approach?
Strong evidence takes several forms: original benchmarks, published research, repeatable testing, documented experiments, or longitudinal client data. Any of these is acceptable. None of them is optional.
The strongest agencies often publish original research, because publication lets prospects inspect the evidence and the methodology at the same time. Published studies in this category have produced findings a buyer can act on directly.
A 2026 benchmark of We measured 50 distinct fintech and payments companies in 51 vertical placements across ChatGPT, Gemini, Claude, and Perplexity. found only 16 percent consistently visible in AI-generated shortlists, with query phrasing alone moving one company's score by 39 points. A separate study of embedded finance citation patterns found that below roughly DR 80, domain authority stopped predicting AI recommendations entirely: Treasury Prime at DR 49 was recommended 19 times while Rapyd at DR 74 was recommended once.
An agency does not need to have produced these exact studies. It needs to show you evidence of comparable rigor, from its own work, that its approach produces the outcome it sells.
6. What Will You Report, and How Often?
A strong answer names the metrics: AI-referred sessions, citation rate across ChatGPT, Gemini, Claude, and Perplexity, and non-brand traffic share, on a defined cadence of lightweight weekly checks and full monthly audits. Require a sample report in the proposal, not a description of one.
A strong answer is also honest about the limits. Today, no equivalent of Google Search Console provides comprehensive AI search reporting. Structured multi-engine audits are the current measurement standard, and they are directional. Any agency promising clean single-post ROI attribution is selling a model that overfits. Measure the system, not the post.
The RFP Checklist
Copied into a document, the following is a working RFP for AI visibility services. It is modeled on how procurement teams structure agency evaluation, not on how agencies market themselves.
A. Minimum Qualifications
Pass/fail, evaluated before scoring
- At least one named case study in fintech or another regulated vertical, with measurable outcomes
- A written methodology available for review before contract
- Named key personnel with roles; disclosure of all subcontracted work
- Willingness to provide references who will be contacted
B. Proof Requirements
- Two case studies minimum, each documenting client name, scope, baseline metrics, measured outcomes, timeframe, and the personnel responsible, cross-referenced to the team proposed for your account
- Three references with contact information and engagement dates
C. Method Requirements
- Written methodology summary, including how AI engines select sources and how the process addresses it
- Sample monthly report
- Measurement plan naming AI-referred sessions and citation tracking across ChatGPT, Gemini, Claude, and Perplexity
- First-90-days plan with sequencing rationale
D. Commercial Terms
- Pricing by engagement structure, with ranges. State your budget in the RFP; procurement teams disclose budgets and receive sharper proposals for it
- Rate card if any work is hourly
- Contract term, exit terms, and ownership of content and data
How to Score the Answers
Publish the weights before you evaluate. That is common procurement practice, and it keeps the decision honest. Pass/fail gates come first; scoring applies only to agencies that clear them.
| Criterion | Weight |
|---|---|
| Demonstrated results: named case studies and references | 30 |
| Methodology and measurement plan | 25 |
| Staffing: who actually does the work | 15 |
| Evidence of efficacy: research, testing, longitudinal data | 10 |
| The agency's own AI visibility | 10 |
| Cost and commercial terms | 10 |
Notice that cost receives the smallest weight. That is intentional. The difference between a mediocre and an exceptional AI visibility program is typically far larger than the difference in agency fees, and the benchmark data above shows why: visibility outcomes vary enormously between companies competing in the same category at similar spend.
Score independently, then interview the top two or three. If written scores land within a few points of each other, the interview decides.
The One-Page Scorecard
Print this. Use it in the interviews. One row per question, one column per agency.
| Question | Evidence Provided | Score (1 to 5) |
|---|---|---|
| Named case studies with numbers | — | — |
| What the AI engines say about them | — | — |
| Who actually does the work | — | — |
| Written methodology | — | — |
| Evidence their approach works | — | — |
| Reporting: metrics, cadence, sample | — | — |
| References (contacted, not just listed) | — | — |
| Commercial terms and ownership | — | — |
Multiply by the weights above if you want a single comparable number. Most buyers find the unweighted scorecard is enough to make the ranking obvious.
What Matters Less Than You Think
| Matters less than you think | Matters more than you think |
|---|---|
| Agency size and headcount | Written methodology |
| Years in business | Measurement discipline |
| Awards and badges | Evidence of efficacy |
| Office location | Regulated-industry experience |
| Follower counts | Named, verifiable case studies |
| Polish of the pitch deck | Who actually does the work |
Everything in the left column is visible before the engagement starts. Everything in the right column predicts what happens after it starts.
The citation study cited above makes the same point with data: below the giant tier, authority signals like size and domain rating stopped predicting AI recommendations at all.
Frequently Asked Questions
What Should an AI Visibility Agency Cost?
Diagnostic-stage engagements typically run $2,500 to $3,500 for a fixed two-week audit. Ongoing build engagements run $5,000 to $10,000 per month. Content for regulated verticals like fintech carries a 20 to 30 percent premium for compliance review. Treat pricing far outside these ranges, in either direction, as a question to ask rather than a reason to reject.
How Long Until AI Visibility Work Shows Results?
First ranking signals typically appear in 60 to 90 days, non-brand movement in 90 to 180 days, and pipeline contribution in 6 to 9 months. The compounding matters more than the start: one global payouts client (PayQuicker) grew organic clicks from 85 to 330 and AI-referred traffic from 15 to 257 over 18 months of sustained work.
Is an AI Visibility Agency Different From an SEO Agency?
Yes, in target and in method. An SEO agency optimizes for rankings on a results page. An AI visibility agency optimizes for inclusion in AI-generated answers, where selection is binary, and citation behavior differs by engine. The full comparison is at AEO agency vs GEO agency vs SEO agency.
Can We Run This Evaluation Without Hiring Anyone?
Yes. The checklist and scorecard are designed for self-service, and the shortlist of AI visibility agencies serving fintech gives you the candidate pool to run them against. The Diagnostic exists for teams that want the visibility gap documented before they shortlist anyone.
The Test That Matters Most
A good agency should not be afraid of this checklist. They should encourage you to use it.
An agency that meets the standard will hand you named case studies, a readable methodology, the people who do the work, and a sample report without being asked twice. An agency that argues with the checklist is answering question six early.
Start from evidence, not proposals.
The AI Visibility Diagnostic documents your visibility gap across all four engines in two weeks, before any agency, ours included, asks you to commit to a build.
Book the AI Visibility Diagnostic