We took the pattern off the island. It held.
23-agency national cohort, 18 hash-locked questions, 3 usable runs on 2026-08-05 and 2026-08-06, 8,833 classified citations. Figures generated from the aggregate tooling. Individual agencies anonymized. Counts and distributions named.
Why this category matters as a measurement subject
Every teardown before this one measured a Hawaii local services market. Dentists, med spas, HVAC, real estate, banking, wealth management, law, CPA firms. A reasonable objection to the whole body of work is that Hawaii is small, the cohorts are local, and the patterns might be an artifact of measuring an island economy.
National marketing and SEO agencies are the cleanest available test of that objection. They are not geographically bounded, they sell to businesses rather than consumers, they are unusually good at their own search marketing, and they are the category most likely to have already optimized for AI answers. If the cross-category patterns were a Hawaii artifact, this is where they should break.
They did not break.
Methodology summary
- Cohort: 23 national marketing and SEO agency domains, surfaced from the citations of a discovery run and hard-audited to exclude directories, ranking sites, and SaaS vendors.
- Questions: 18, locked by hash
8732c92d, written as the questions a business owner actually asks when looking for an agency rather than as brand queries. - Tools: 7. Five that search the live web (Perplexity, ChatGPT search, Gemini grounded, Microsoft Copilot, Google AI Overviews) and two that answer from training data (Claude, Gemma).
- Repetition: 3 reps per question per tool per run, 378 calls per run.
- Runs: 3 usable, on 2026-08-05 and 2026-08-06. A fourth run was fired and excluded because it completed only 46 to 47 of 54 calls per engine. Partial runs are not small complete runs, because engines fail unevenly and the missing calls concentrate.
- Geography: none. The question set is deliberately geo-free, which is what makes the result comparable across the 127 markets the cohort spans.
Source-type distribution (cohort-wide)
All 8,833 classified citations across 3 runs, 7 tools, and 18 questions.
| Source type | % of citations | Count |
|---|---|---|
| Independent web (see the section below on what this bucket is) | 77% | 6,791 |
| Competitor (agency-owned websites in the cohort) | 12% | 1,055 |
| YouTube | 4% | 316 |
| 3% | 278 | |
| Social | 2% | 207 |
| Review directories | 1% | 97 |
| Forum | 1% | 45 |
| Wikipedia | 0% | 44 |
For comparison, the published Honolulu real estate teardown records 15% own-firm sites and 77% independent web. Agencies come in at 12% and 77%. The most aggregator-dominated local category we have measured and the national B2B category look close to identical.
Per-AI-tool breakdown
| AI tool | Agency-site share | Independent-web share | Total citations |
|---|---|---|---|
| Gemma (training data) | 39% | 61% | 489 |
| Gemini grounded | 16% | 81% | 1,717 |
| Perplexity | 14% | 68% | 3,184 |
| Google AI Overviews | 9% | 63% | 975 |
| Claude (training data) | 3% | 97% | 569 |
| Microsoft Copilot (Bing) | 2% | 87% | 782 |
| ChatGPT search | 1% | 98% | 1,117 |
The spread between the most and least agency-citing tool is roughly thirtyfold, on identical questions in the same week. This is the same shape the cross-category teardown documents across nine Hawaii categories, and it is the reason a visibility number taken from one tool is a measurement of that tool rather than of the business.
The prediction that held
Copilot, ten Hawaii categories, then one national B2B category
The cross-category teardown states that Microsoft Copilot cites businesses’ own websites “0% to 2% of the time in every measurement.” That range was derived entirely from Hawaii local services: banking 0%, wealth management 0%, dental 0%, med spas 0%, HVAC 0%, real estate 0%, law 1%, CPA 2%, plus Austin CPA at 0% and a Nashville CPA control at 1%.
National marketing agencies, measured months later, in a different country-wide market, in a business category rather than a consumer one, came in at 2%.
A published range predicting an out-of-sample category is worth more than any single number in this teardown. It is the difference between a collection of observations and something that generalizes.
It also carries a practical consequence. Copilot points at almost nobody’s own website in any category we have measured, which means the slot is decided by whatever ranks first in Bing organic rather than by anything on the business’s own site.
What “independent web” is, and what it is not
This section exists because we got it wrong internally before publishing, and the correction is more useful than pretending we did not.
Independent web is a residual bucket. The classifier sorts a citation into it when the host is not an agency in the cohort, not a review directory, not YouTube, Reddit, Wikipedia, a forum, or a social network. It is deliberately vague, because distinguishing a major publication from a personal blog is not honestly knowable from a hostname alone.
It is not a measure of how often AI cites agency websites. In this run its largest members are:
| Host | Citations | What it is |
|---|---|---|
| google.com | 591 | a search engine |
| semrush.com | 145 | an SEO software vendor |
| merriam-webster.com | 117 | a dictionary |
| dictionary.cambridge.org | 97 | a dictionary |
Anyone reading 77% as “agency websites get cited 77% of the time” would be reading two dictionaries and a search engine into that number. The defensible figure for agency-owned websites is the competitor row, 12%.
Top recurring agencies (anonymized)
The 5 cohort agencies AI cited most often across the 18 questions and 7 tools, by total citations across the 3 runs:
| Agency (anonymized) | Total citations | Questions cited on |
|---|---|---|
| Agency A | 101 | 8/18 |
| Agency B | 83 | 12/18 |
| Agency C | 72 | 6/18 |
| Agency D | 68 | 9/18 |
| Agency E | 68 | 8/18 |
Nobody owns this category. The most-cited agency website in the entire run holds 101 of 8,833 citations, which is 1%. The top five combined reach 4.4%. No agency was cited on more than 12 of the 18 questions, so there is no equivalent of the long-established brokerage that the training-data engines have learned in real estate. Fragmentation cuts both ways: nobody has locked the category up, and being cited once does not mean much.
What this teardown does and does not prove
What it supports:
- The Copilot 0% to 2% own-site pattern, previously derived only from Hawaii local services, held in a national B2B category measured months later.
- Agency-owned websites take 12% of citations in their own category, against 15% for local firms in Honolulu real estate. The categories behave alike.
- The category is fragmented. The most-cited agency holds 1% of all citations, and no agency appears on more than 12 of 18 questions.
- The per-tool spread runs from 1% to 39% on identical questions in the same week, consistent across all 3 runs.
What it does not support:
- That agencies are worse or better at AI visibility than the clients they serve. The numbers are close, and this measurement was not designed to rank the two against each other.
- That AI behavior on these questions stays stable over months. Models refresh training data and search indices on schedules outside our control. Re-measurement is the only honest answer.
- That changing an agency’s site or its third-party presence would cause AI to cite differently. We measured what AI cites. Causation requires pre-registered experiments against named sites with control for confounds. Different scope.
- Anything about what is inside the 77% independent web bucket beyond the hosts named above. It is a residual, and we do not characterize it.
Why this is anonymized
None of the 23 agencies in this cohort are paying NeverRanked customers. The non-customer anonymization rule applies: counts, distributions, and per-AI-tool numbers are public. Individual agency names are not. The pattern is what is informative on a public surface. An agency that becomes a customer gets a 1:1 deliverable that names every agency in the cohort, names the questions it is missing on, and ranks the closable conditions. That deliverable is private to the customer.
Measurement window: 3 usable runs on 2026-08-05 and 2026-08-06. A fourth run was excluded by the aggregate completeness gate at 46 to 47 of 54 calls per engine. Figures generated from the aggregate tooling. Pattern-readiness rule of 3 runs cleared. Refresh cadence is monthly or on customer request.
Substantiation: question set locked by hash 8732c92d..., documented method at /methodology/, named AI tools on named dates with the dated runs on the /claims/ ledger. Gemma is open-weight, so the model itself is independently inspectable.
Anonymization: the 23-agency cohort is kept anonymized at the agency level per the non-customer rule. Counts, distributions, and category-level source surfaces are public. Individual agency names are not.
Removal: any agency in this cohort can be removed on request. Email takedown@neverranked.com and it comes down within 24 hours.