FIELD NOTES — METHODOLOGY
The Myth of a Single AI Visibility Score
What I found measuring five fintech platforms across every major AI engine.
Jeff Jenkins · August 5, 2026 · 7 min read

TL;DR
"AI visibility" is not one metric. It is at least six, each answering a different question: who gets recommended, who gets cited, who influences Google's AI answer, who gets the click, who gets seen, who is findable. A brand can lead one and be invisible on another. Before you read any dashboard, decide which decision you are trying to influence.
There is no single AI visibility score. There are many, and they are all real. CARL has one. Profound has one. AthenaHQ has one. The problem is not that these numbers are wrong. It is that none of them is the AI visibility score, and treating any one of them that way quietly answers a question you never asked.
I tested five card-issuing and payments platforms across four AI engines and two measurement systems. The same brand routinely looked like a category leader on one metric and nearly invisible on another. That is not a data problem. It is a definition problem.
How I Got Here
I recently got access to CARL, the AI visibility platform from Xponent21, and I lined it up against the metrics I already use, including Ahrefs citation data and traditional search reporting. I expected to learn which tool was more accurate.
Instead, I realized they were not measuring the same thing. One was telling me who gets recommended. Another was telling me who gets cited. A third was telling me who gets clicked. I had been treating those as one number. They are not.
Six Questions Wearing One Name
"AI visibility" is not a metric. It is a category of metrics, and each one answers a different question about a different moment in the buying journey. Understanding the distinction is foundational to both answer engine optimization (AEO) and generative engine optimization (GEO).
| Metric | The question it answers | Where it can mislead you |
|---|---|---|
| Prompt-level recommendations (e.g. CARL) | Who does AI recommend when a buyer asks? | Swings hard on the exact prompt and the model |
| Domain citations (e.g. Ahrefs) | Whose site does AI cite as a source? | A brand can be cited constantly and still never be the recommendation |
| Google AI Overviews | Who influences Google's AI answer? | Overweights informational content; presence is not endorsement |
| Referral traffic (GA4) | Who actually gets the click from AI? | Often mislabeled as "direct" or "other," so it undercounts |
| Search impressions (Search Console) | Who gets seen in results? | Impressions are not clicks, and clicks are not recommendations |
| Rankings (traditional SEO) | Who is findable on Google? | A number one ranking can still lose the click to the AI answer above it |
Read that list again. "Who gets cited," "who gets recommended," "who gets clicked," and "who gets found" are four different questions about four different moments. A fintech can win one and lose the rest.
The Proof: Five Brands, Five Different Stories
We tested five fintech platforms across four AI models using two measurement systems. Here is the whole picture on one screen.
| Brand | CARL (recommended) | Ahrefs (cited) | ChatGPT | Claude | Gemini | Perplexity |
|---|---|---|---|---|---|---|
| Stripe | 95% | 5,151 | 100% | 100% | 79% | 100% |
| Marqeta | 66% | 37 | 71% | 71% | 64% | 57% |
| Unit | 57% | 31 | 71% | 57% | 71% | 29% |
| Galileo | 46% | 21 | 71% | 29% | 57% | 29% |
| Lithic | 39% | 2 | 0% | 57% | 57% | 43% |
CARL is prompt-level recommendation share. Ahrefs is total domain citations across AI answers. The four model columns are CARL mention rates per engine. Snapshot, US, August 2026.
The columns do not line up, and that is the point. If every column told the same story, we would not need six metrics. This matters because teams are already making budget decisions, content decisions, and executive reports based on whichever AI metric happens to be in the dashboard. If that metric answers the wrong question, you can optimize successfully toward the wrong outcome.
The leaders were not surprising. Stripe topped every column, exactly as you would expect. The disagreements below it are where the story is.
Lithic is the clearest case. On CARL's prompt recommendations, it scored a respectable 39 percent, close to Galileo. On Ahrefs domain citations, it registered just 2, an order of magnitude below Galileo's 21.
One system says Lithic is a credible presence. The other says it is nearly invisible. Both are correct. They are answering different questions. Lithic appears far more often in CARL's prompt sample than its citation footprint would suggest, which is worth knowing, and you would miss it completely if you trusted a single number.
The Model Matters as Much as the Metric
The sharpest finding was not about the tools at all. It was about the models.
- Lithic:0 percent in ChatGPT, 57 percent in both Claude and Gemini.
- Galileo:71 percent in ChatGPT, 29 percent in Claude.
- Unit:71 percent in ChatGPT and Gemini, 29 percent in Perplexity.
Same brand. Same day. Same prompt. The only thing that changed was which model answered, and the verdict flipped from invisible to dominant. This is not a measurement artifact you can tool your way around.
Different models retrieve, weigh, and recommend differently, so "AI visibility" is not even one number per tool. It is one number per model, per question. The engine matters as much as the metric you chose to read it with.
The Mistake Most Marketers Make
The mistake is assuming those questions are interchangeable. They are not.
A brand that is cited constantly as a source is not necessarily the brand the model recommends. A brand that ranks number one on Google is not necessarily the brand that appears when a buyer asks ChatGPT which platform to use. When a marketing leader asks "what is our AI visibility," they usually mean one specific thing, but the tool answers a different one, and nobody notices the swap.
Why the Tools Disagree, and Why That Is the Point
It would be easy to conclude that one of these tools is wrong. That is the wrong conclusion.
They diverge because they are observing different parts of the same buying journey. Citation data watches what AI reads. Recommendation data watches what AI says. Referral data watches what the buyer does next.
None of them is inaccurate. Each is a partial view, and the disagreement between them is information, not error. When two honest instruments disagree, the gap itself is telling you something about the thing you are measuring.
What This Means for Fintech Marketing Leaders
While this dataset comes from fintech, the problem is not unique to fintech. Any category measured across recommendation, citation, and referral data will hit the same wall.
You cannot manage what you have not defined. If your team reports a single "AI visibility" number, you are almost certainly optimizing for a question you did not choose on purpose. The fix is not a better tool. It is a better question.
Before you open any dashboard, decide which decision you are trying to influence. Are you trying to get cited as a source, get recommended on a shortlist, or get the click? Those three goals lead to three different bodies of work.
The next time someone on your team asks, "What is our AI visibility score?" do not answer with a number. Answer with a question. Which decision are we trying to influence?
That question is where a real diagnosis starts. It is the difference between a dashboard and a strategy, and it is the first thing we work out in a DIGI CONVO AI Visibility Diagnostic.
FAQs
Is there a single best AI visibility metric for fintech?
No. Each metric answers a different question. Citation data tells you who AI reads, recommendation data tells you who AI names, and referral data tells you who gets the click. The right metric depends on which decision you are trying to influence.
Why does my brand appear in Perplexity but not ChatGPT?
Because each model retrieves and weighs sources differently. In our five-brand test, the same platform scored 0 percent in ChatGPT and 57 percent in Claude and Gemini on the same prompt. Model-level differences are normal, which is why a single blended score hides more than it shows.
Does ranking number one on Google mean I am visible in AI?
Not necessarily. A number one organic ranking answers who is findable. It does not answer who gets recommended or who gets cited, and an AI Overview can absorb the click before a buyer ever reaches your page.
Data: CARL (Xponent21) prompt-level share of voice and Ahrefs Site Explorer AI citation data, US, point-in-time snapshot, August 2026. Figures are a snapshot and will shift as the underlying models change.
About the Author
Jeff Jenkins runs DIGI CONVO, where he works on AI visibility for fintech and payments companies. Connect on LinkedIn.
