At one B2B client, referral tracking credited AI assistants with 6 customers over twelve months. In six months of asking customers how they heard about the company, 39 selected AI.
Those figures are hard to ignore. They're also easy to turn into a claim the analysis doesn't support. Different reporting windows and different definitions mean you can't divide one by the other and call the result an AI undercount.
What caught my attention was how little the two lists overlapped. Customers were telling us something the referral report couldn't explain. The useful question was what that should change about our reporting and spending.
Download the anonymized analysis (PDF)
Two ways of seeing AI
The client has a considered purchase and a long sales cycle. Referral tracking recorded signups arriving directly from AI tools. A “how did you hear about us” question appeared in a step every customer completes before closing.
| Measure | Window | Signups | Customers |
|---|---|---|---|
| Tracked AI referrals | Twelve months | 417 | 6 |
| Self-reported AI | Six months | 228 | 39 |
Counts are scaled by a fixed factor to protect the client. Ratios and rates are unchanged.
The lists overlapped by 24 accounts. The survey reached 11 to 15% of signups and every customer, so signup findings describe respondents. Customer findings cover the client's customers in the survey period.
Even there, the wording matters. The AI option didn't distinguish a ChatGPT recommendation from an AI summary in Google. I'll call this group self-reported AI, because labeling everyone “AI-sourced” would imply more certainty about their journey than we have.
Where those customers arrived
Among the 228 signups who selected AI, paid search was the largest tracked arrival channel, followed by direct and organic search. Direct arrivals from AI tools ranked fourth, at about one in ten.
A possible journey explains the mismatch. A buyer asks ChatGPT who to consider, sees a company named, then searches Google for that company and clicks an ad. Analytics records the search visit. Without a click from the assistant, there is no AI referral to record.
That example shows how the gap can happen. We didn't observe that exact sequence for every customer. Referral information can also be missing, and customers can interpret the survey differently. The finding is that paid search and self-reported AI frequently described the same accounts. It doesn't establish that most AI-related demand was hiding in brand ads.
Checking what the answers could support
Rand Fishkin's critique of “how did you hear about us” questions raises real problems. People forget exposures, interpret the question differently, and answer from an incomplete memory of the purchase. Asking only customers also leaves out everyone exposed to the marketing who never bought. That limits what the survey can tell you about marketing effectiveness.
I checked the answers against ad clicks and sales outreach records to see whether the patterns held together.
Among paid signups who answered, 29% of the brand group selected AI versus 9% of the non-brand group, a 3.2-times difference. Those groups contained 144 and 486 respondents, respectively, using the same scaled counts. Only 44% of the brand group named search, compared with 70% of the non-brand group.
That association makes the AI-to-brand-search journey worth investigating. It doesn't prove AI caused the brand search. Someone already familiar with the company could have used an assistant later to compare options.
The outreach history added context. Earlier SDR contact appeared for 7.0% of paid brand signups, compared with 2.0% of paid non-brand signups. That supports the broader point that brand search can follow earlier exposure. It doesn't independently validate the AI answers.
Kevin Indig's The usefulness of attribution for AI Search, with George Bonaci at Ramp, explores using measures with different blind spots. That's the value here. Each source helps narrow the questions we still need to answer.
What happened after purchase
The customer-quality analysis used a separate year-to-date cohort: 48 self-reported AI customers out of 270, using scaled counts. That is 17.8% of customers, contributing 17.5% of revenue. Their median deal was about a third larger than the median search deal. This is a different reporting window from the six-month comparison above.
Those figures gave me a reason to keep investigating. The AI answers were attached to paying customers contributing a similar share of revenue, rather than just a collection of low-value signups.
Other differences need more time to interpret. The group averaged 2.06 purchases per customer versus 1.97 for search, and took 16 days from signup to first purchase versus 11 for search. Repeat purchases depend on how long customers have been observed. Without matching cohort age, I'd treat that difference as an observation, not evidence of better retention. The underlying group is smaller than the scaled counts suggest, which makes precise comparisons fragile.
What I would change in the reporting
Start by joining the records at the account level. Put the self-reported answer beside the tracked channel, signup date, customer date, and revenue. Compare the same acquisition window and allow the same time for customers to convert. That would give us a cleaner view of the overlap than the initial twelve-month and six-month reports.
Keep the two attribution fields separate. One records the tracked arrival. The other records what the customer remembers. An account can belong in both views, but adding their customer totals together would count it twice.
I'd also improve the question. Ask every customer, retain an “other” answer, and distinguish AI assistants from Google AI results. An optional text field can capture the tool or experience someone remembers. Preserve those words alongside the category so you can inspect ambiguous answers later.
The next spending question is brand paid search. How many customers would still arrive if those ads weren't running? Neither report answers that.
Where volume supports it, I'd test a regional holdout with comparable control regions, checking pre-test trends and keeping other activity as stable as possible. Agree on a window long enough to capture the buying cycle. Measure total customers and revenue across channels, because a drop in paid conversions could simply reflect buyers moving to organic results. A small test with too few purchases won't settle the budget question.
The referral report alone made AI look small. The customer answers made it worth investigating. Together, they gave us a better next decision: improve the measurement and test what brand ads add before moving spend.
FAQ
Is “how did you hear about us” data reliable?
It's useful context, with imperfect recall and ambiguous wording. Reaching every customer improves coverage, but still excludes nonbuyers. Compare answers with recorded behavior before using them to make spending decisions.
Why doesn't Google Analytics show traffic from ChatGPT?
It can show identifiable referral visits. It can't record an AI referral when someone reads an answer and later arrives through Google or another route. Missing referral information can also limit visibility.
Does AI search show up as branded search?
It can. At this client, brand-ad signups were 3.2 times as likely to select AI as non-brand-ad signups. That association supports further investigation, not a causal claim about each purchase.
Should I cut brand paid search if AI is driving the demand?
Test first. Prior exposure doesn't tell you whether the ad helped secure the purchase. A well-designed holdout can help estimate its contribution if you have enough volume and observation time.
How should AI be reported alongside channel data?
Keep tracked channels and self-reported discovery in separate fields, linked to the same account. Compare matching windows and show overlap. Don't add the totals together.
