FIELD NOTES — METHODOLOGY

The Myth of a Single AI Visibility Score

What I found measuring five fintech platforms across every major AI engine.

Jeff Jenkins · August 5, 2026 · 7 min read

Comparison matrix of five fintech platforms (Stripe, Marqeta, Unit, Lithic, Galileo) across CARL, Ahrefs, ChatGPT, Claude, Gemini, and Perplexity showing that AI visibility scores disagree.

TL;DR

"AI visibility" is not one metric. It is at least six, each answering a different question: who gets recommended, who gets cited, who influences Google's AI answer, who gets the click, who gets seen, who is findable. A brand can lead one and be invisible on another. Before you read any dashboard, decide which decision you are trying to influence.

There is no single AI visibility score. There are many, and they are all real. CARL has one. Profound has one. AthenaHQ has one. The problem is not that these numbers are wrong. It is that none of them is the AI visibility score, and treating any one of them that way quietly answers a question you never asked.

I tested five card-issuing and payments platforms across four AI engines and two measurement systems. The same brand routinely looked like a category leader on one metric and nearly invisible on another. That is not a data problem. It is a definition problem.

How I Got Here

I recently got access to CARL, the AI visibility platform from Xponent21, and I lined it up against the metrics I already use, including Ahrefs citation data and traditional search reporting. I expected to learn which tool was more accurate.

Instead, I realized they were not measuring the same thing. One was telling me who gets recommended. Another was telling me who gets cited. A third was telling me who gets clicked. I had been treating those as one number. They are not.

Six Questions Wearing One Name

"AI visibility" is not a metric. It is a category of metrics, and each one answers a different question about a different moment in the buying journey. Understanding the distinction is foundational to both answer engine optimization (AEO) and generative engine optimization (GEO).

MetricThe question it answersWhere it can mislead you
Prompt-level recommendations (e.g. CARL)Who does AI recommend when a buyer asks?Swings hard on the exact prompt and the model
Domain citations (e.g. Ahrefs)Whose site does AI cite as a source?A brand can be cited constantly and still never be the recommendation
Google AI OverviewsWho influences Google's AI answer?Overweights informational content; presence is not endorsement
Referral traffic (GA4)Who actually gets the click from AI?Often mislabeled as "direct" or "other," so it undercounts
Search impressions (Search Console)Who gets seen in results?Impressions are not clicks, and clicks are not recommendations
Rankings (traditional SEO)Who is findable on Google?A number one ranking can still lose the click to the AI answer above it

Read that list again. "Who gets cited," "who gets recommended," "who gets clicked," and "who gets found" are four different questions about four different moments. A fintech can win one and lose the rest.

The Proof: Five Brands, Five Different Stories

We tested five fintech platforms across four AI models using two measurement systems. Here is the whole picture on one screen.

BrandCARL (recommended)Ahrefs (cited)ChatGPTClaudeGeminiPerplexity
Stripe95%5,151100%100%79%100%
Marqeta66%3771%71%64%57%
Unit57%3171%57%71%29%
Galileo46%2171%29%57%29%
Lithic39%20%57%57%43%

CARL is prompt-level recommendation share. Ahrefs is total domain citations across AI answers. The four model columns are CARL mention rates per engine. Snapshot, US, August 2026.

The columns do not line up, and that is the point. If every column told the same story, we would not need six metrics. This matters because teams are already making budget decisions, content decisions, and executive reports based on whichever AI metric happens to be in the dashboard. If that metric answers the wrong question, you can optimize successfully toward the wrong outcome.

The leaders were not surprising. Stripe topped every column, exactly as you would expect. The disagreements below it are where the story is.

Lithic is the clearest case. On CARL's prompt recommendations, it scored a respectable 39 percent, close to Galileo. On Ahrefs domain citations, it registered just 2, an order of magnitude below Galileo's 21.

One system says Lithic is a credible presence. The other says it is nearly invisible. Both are correct. They are answering different questions. Lithic appears far more often in CARL's prompt sample than its citation footprint would suggest, which is worth knowing, and you would miss it completely if you trusted a single number.

The Model Matters as Much as the Metric

The sharpest finding was not about the tools at all. It was about the models.

  • Lithic:0 percent in ChatGPT, 57 percent in both Claude and Gemini.
  • Galileo:71 percent in ChatGPT, 29 percent in Claude.
  • Unit:71 percent in ChatGPT and Gemini, 29 percent in Perplexity.

Same brand. Same day. Same prompt. The only thing that changed was which model answered, and the verdict flipped from invisible to dominant. This is not a measurement artifact you can tool your way around.

Different models retrieve, weigh, and recommend differently, so "AI visibility" is not even one number per tool. It is one number per model, per question. The engine matters as much as the metric you chose to read it with.

The Mistake Most Marketers Make

The mistake is assuming those questions are interchangeable. They are not.

A brand that is cited constantly as a source is not necessarily the brand the model recommends. A brand that ranks number one on Google is not necessarily the brand that appears when a buyer asks ChatGPT which platform to use. When a marketing leader asks "what is our AI visibility," they usually mean one specific thing, but the tool answers a different one, and nobody notices the swap.

Why the Tools Disagree, and Why That Is the Point

It would be easy to conclude that one of these tools is wrong. That is the wrong conclusion.

They diverge because they are observing different parts of the same buying journey. Citation data watches what AI reads. Recommendation data watches what AI says. Referral data watches what the buyer does next.

None of them is inaccurate. Each is a partial view, and the disagreement between them is information, not error. When two honest instruments disagree, the gap itself is telling you something about the thing you are measuring.

What This Means for Fintech Marketing Leaders

While this dataset comes from fintech, the problem is not unique to fintech. Any category measured across recommendation, citation, and referral data will hit the same wall.

You cannot manage what you have not defined. If your team reports a single "AI visibility" number, you are almost certainly optimizing for a question you did not choose on purpose. The fix is not a better tool. It is a better question.

Before you open any dashboard, decide which decision you are trying to influence. Are you trying to get cited as a source, get recommended on a shortlist, or get the click? Those three goals lead to three different bodies of work.

The next time someone on your team asks, "What is our AI visibility score?" do not answer with a number. Answer with a question. Which decision are we trying to influence?

That question is where a real diagnosis starts. It is the difference between a dashboard and a strategy, and it is the first thing we work out in a DIGI CONVO AI Visibility Diagnostic.

FAQs

Is there a single best AI visibility metric for fintech?

No. Each metric answers a different question. Citation data tells you who AI reads, recommendation data tells you who AI names, and referral data tells you who gets the click. The right metric depends on which decision you are trying to influence.

Why does my brand appear in Perplexity but not ChatGPT?

Because each model retrieves and weighs sources differently. In our five-brand test, the same platform scored 0 percent in ChatGPT and 57 percent in Claude and Gemini on the same prompt. Model-level differences are normal, which is why a single blended score hides more than it shows.

Does ranking number one on Google mean I am visible in AI?

Not necessarily. A number one organic ranking answers who is findable. It does not answer who gets recommended or who gets cited, and an AI Overview can absorb the click before a buyer ever reaches your page.

Data: CARL (Xponent21) prompt-level share of voice and Ahrefs Site Explorer AI citation data, US, point-in-time snapshot, August 2026. Figures are a snapshot and will shift as the underlying models change.

About the Author

Jeff Jenkins runs DIGI CONVO, where he works on AI visibility for fintech and payments companies. Connect on LinkedIn.

NEXT STEP

Which decision are you trying to influence?

The AI Visibility Diagnostic maps your citation rate, recommendation share, and referral traffic across ChatGPT, Gemini, Claude, and Perplexity so you are optimizing for the right question.

Book the AI Visibility Diagnostic