Citation Index
Methodology
Everything needed to repeat this measurement and disagree with it.
What is being measured
Whether an AI answer engine returns a company’s site among its sources when someone asks the question that company would want to be found for. Not traffic, not rankings, not impressions: citations.
The honest limit, first
No answer engine exposes “was I cited” as an API. There is no ground truth to read, so this is sampling: I ask the questions myself, on a schedule, of engines that return their sources, and record what came back. Every number here is evidence, never a measurement of record.
The questions
Two shapes, both asked per category rather than per company, and neither ever naming a company. A probe that mentions a brand steers the engine into citing it, which would measure the prompt instead of the market.
- Discovery: “What are the best tools for [category]?”, the question from someone who does not know you yet, and the citation that actually matters.
- Comparative: “What are the best alternatives to [incumbent]?”, the other way a buyer arrives. Skipped entirely for categories with no obvious leader, rather than inventing a rival.
Repetition
Answer engines are not deterministic: the same question minutes apart can return different sources. Every question is therefore asked 5 times per engine, and a result is reported as a rate over those runs, never as a yes or a no from a single call.
What counts as a citation
- Cited: the company’s domain appears among the sources the engine returned. Subdomains count, because an engine citing
blog.company.comis citing that company. The match is on a dot boundary, sonotcompany.comnever matchescompany.com. - Mentioned: the company’s name appears in the answer text without its site being cited. Reported separately and never merged into the citation count.
- Ambiguous: the name matched but it is an ordinary word (Linear, Bolt, Arc), so the match may be coincidence. Flagged rather than resolved. The headline number counts citations only, so ambiguity can never move it.
The control group
Half the sample is companies that paid to sit on a pay-to-rank board; the other half is comparable companies in the same categories that did not. Without the control the number would be unreadable: a 12% citation rate is meaningless until you know what it is for everyone else. Both halves are judged against the identical source list from the identical probe, so the comparison is not confounded by having asked different questions.
Rates are per engine
Engines disagree, and averaging them produces a number that describes none of them. Every rate on the report carries the engine it came from. A rate over zero probes is reported as not measured, never as 0%: the first says I have not asked, the second says the engine never cites you, and confusing them invents a finding.
What this cannot tell you
- Causation. A company on a board being cited does not mean the board caused it. The control group and the repeated editions over time are the only honest answer, and they narrow the question rather than closing it.
- Latency. Answer engines update their sources over weeks. The first edition is a baseline taken days after the boards appeared, which is early by design: the point of a baseline is to exist before the thing you want to observe.
- Personalisation. Probes run through APIs with no session and no history. A logged-in human asking the same question may see something else.
Who publishes this
GetTrafficFast, which sells a product for getting cited by answer engines. That is a conflict of interest, which is precisely why the questions, the repetition count, the matching rules and the raw probes are all published: so the measurement can be repeated by someone who does not trust me.