How the AI Index is built.
The exact method behind every ranking: the questions we ask, the engines we ask them of, how the score is computed, and the rules that keep it honest.
Methodology v1.3.1 · current as of October 2026
The question we ask
For every category, each AI model is asked the question a real buyer asks: “What are the best {category}? Recommend the top brands or products that people actually use.” Product categories, whose boards rank individual models, ask for specific models instead: “Recommend the specific models people actually buy”. No brand names are supplied either way. Regional editions ask the same question the way a local buyer would, naming the country.
Every monthly refresh asks each AI model the same buyer question once, and the exact run count behind every edition is published in its JSON record. Every figure is therefore reproducible. One engine’s quirky answer cannot create a ranking on its own: the consensus gate below requires at least two AI models to agree before a brand is ranked at all.
The engines
All 9 leading AI models, every refresh: no engine is skipped because it is inconvenient to reach. The exact model version behind every captured answer is recorded with it.
| Engine | Consumer flagship | The model the Index asks |
|---|---|---|
| ChatGPT | ChatGPT | gpt-5.5 |
| Claude | Claude Sonnet 5 | claude-sonnet-5 |
| Gemini | Gemini 3.8 Flash | gemini-3.5-flash |
| Perplexity | Sonar | sonar |
| DeepSeek | DeepSeek V4 Pro | deepseek-v4-flash |
| Grok | Grok 4.5 | grok-4.3 |
| Copilot | Copilot | Read from the consumer surface, no API model to pin |
| Google AI | AI Overviews | Read from the consumer surface, no API model to pin |
| Google AI Mode | AI Mode | Read from the consumer surface, no API model to pin |
Where the two columns differ, the Index is pinned on purpose and the reason is dated in the changelog below. The workspace product keeps its own pins, which are a separate unit-economics decision and are not what this page describes.
The score: the full formula
Brands are ranked by their AI Recommendation Score (0 to 100). The complete calculation is published here so any position can be recomputed from the answers. There is no editorial weighting and no hidden step: every brand recommendation is extracted from every captured answer, a brand counts at most once per answer however many of its products that answer names, a consensus gate requires at least two AI models, and the score combines share of voice, mention rate and how early the AI models name the brand.
What counts as an answer: an AI model that returns nothing (no AI Overview shown, a provider outage) still counts as an answer, so it stays in the mention-rate denominator, and brand pages show it as “no answer” rather than as a brand the AI model declined to name. That is the empty-answer correction of 2 September 2026 below. An answer the extractor could not read is a different case: it leaves the mention-rate pool entirely under the v1.3 revision of the same date, and brand pages label it “unreadable” while still publishing its verbatim text.
Ties: when two brands land on the same score, the tie breaks on average position (earlier wins), then on total answers naming the brand, then on how often each is named first in an answer, and last on brand name. Every key is a property of the answer corpus, never of the order in which the answers were read, as the v1.3.1 entry in the changelog below records. Ranks are strictly sequential, so three brands tied on 67.5 rank 2, 3 and 4, never joint 2nd. Breadth comes before primacy in the score: a brand named first more often can rank below one named by more AI models.
Why there is no citation component: our July 2026 audit showed citations reward the wrong thing. Review sites that answers cite as sources outranked products that six of eight AI models actually recommended, so citations were removed from the score.
The gates
A brand must be recommended by at least two different AI models to be ranked at all. One engine’s quirk is not a market position.
Every ranking keeps its receipts: the verbatim answers the AI models gave are stored and shown, so any position can be checked against the answers that produced it.
Each monthly refresh is captured as an immutable snapshot. Movement arrows compare against the previous edition; past editions are never edited.
What never ranks: entity exclusions
A ranking must contain entities of the category, so two published exclusion rules run before scoring (added 7 August 2026, dated in the changelog below). First, in service-provider categories, the directories and marketplaces that rank providers are never themselves ranked: an entity is excluded when its domain sits on our curated source table as a directory or marketplace. Second, the per-category exclusion list below removes brands AI models name in passing that are not members of the category, such as an ad platform named inside an agency recommendation. Both rules require the entity’s name to identify it as the domain’s owner, so a real provider an answer happened to attribute to a directory’s domain always stays ranked, and a per-category exception list keeps legitimate incumbents in place (Amazon remains a third-party-logistics provider). Service boards rank providers, not product makers: a manufacturer whose products those providers install or resell is excluded from a service board even when AI models name it beside them, and it stays eligible on any product board of its own.
Excluded mentions leave the share-of-voice pool, so remaining scores reflect only category members. Nothing is removed from the record itself: every excluded mention stays verbatim in the published answer corpus and its extractedBrands, so any board remains recomputable from its own record. The list changes only with a dated changelog entry. An exclusion can apply to one country’s board alone when the same company is a member of the category in another market; those entries name the board region beside the category. Where AI models name a non-member without a web address, or under an address its name does not own, the entry is the brand’s name instead, marked “by name” below: it matches that name exactly, ignoring case and punctuation, and never a longer name that contains it.
Show the exclusion list (1000 entries across 84 categories)Hide the exclusion list
Tamper-evident records
Every refresh freezes its receipts twice. At capture time, the complete verbatim answer corpus is hashed with SHA-256 and the hash is stored on the immutable snapshot and published in the record’s JSON download, next to the answers themselves. Anyone can recompute the hash from the published corpus, using the spec shipped in the same file, and confirm the answers were not edited after publication.
Once a monthly sweep is verified complete, the per-category hashes are combined into a single run-level root hash, published here and inside that month’s Index report. One changed character in any archived answer changes its category hash, which changes the root.
Show earlier editions (6)Hide earlier editions
Non-determinism, sampling and variance
The same AI model, asked the same question twice, does not reliably give the same answer. This is a property of the models, not a flaw in any measurement, and any methodology that does not address it is marking its own homework. We measured it on our own corpus: of 535 engine-prompt pairs asked at least twice over 60 days, only 16.3% produced the identical set of search queries every run (published August 2026, with the full sample). A single point-in-time check of any AI answer is one roll of the dice.
The Index is designed around that fact rather than around pretending it away. Rankings never rest on a single roll: the consensus gate requires at least two independent AI models before a brand ranks at all, so one run’s quirk cannot create a position. Movement is read edition over edition against immutable snapshots, never inside a single run. The exact run count behind every edition is published in its JSON record, the exact model version is recorded on every answer, and model changes are dated in the changelog below with their expected variance noted, so month-over-month movement is always attributable to either the market or the method, never silently to both.
Sampling cadence is deliberate and disclosed: the public Index refreshes monthly in a single verified sweep, and CiteHawk workspaces collect weekly. More runs per refresh would smooth variance further, and the complete verbatim corpus is published precisely so anyone can quantify the remaining variance themselves rather than taking our word for it. That is the trade we choose: fewer, fully published, tamper-evident runs over many unpublishable ones.
Regional rankings
Some categories (banks, insurers, agencies, professional services) get genuinely different answers in different countries, so those categories are also asked as a local buyer would ask, for the United States, the United Kingdom, Australia and Canada. A regional edition is published only when its ranking meaningfully differs from the global one: a different #1, or fewer than three-quarters of the top 10 in common.
Rankings are computed from AI responses only. Claiming a brand cannot change its position, and positions are not for sale.
Brand owners can claim their brand to verify identity details and follow their movement. Identity corrections are reviewed and never affect scores. No payment, partnership, or relationship with CiteHawk influences any ranking.
How claiming works
Any brand on the Index can be claimed, free, by someone who works there. Claiming starts from the brand’s own page: sign up with an email address at the brand’s domain and the claim goes to a human review. Reviews usually finish within two days.
An approved claim unlocks the verified mark on the brand’s page, movement alerts when its position changes, the embeddable certificate badge, views and clicks on its listing, and a short description written by the brand, checked against our listing rules and labelled as the brand’s own. From 1 October 2026 the website link on a brand’s Index page is shown only for claimed brands; unclaimed brands show their domain as plain text. Anyone, claimed or not, can request a correction to a brand’s name, domain or category.
Claiming never changes a ranking. Positions come from AI answers only, the review checks identity and nothing else, and no payment or relationship with CiteHawk moves a brand.
Asking to be removed
A brand that does not want to appear on the Index can ask to be removed. Email support@citehawk.com from an address at the brand’s own domain and name the boards. Removed from the live Index within one business day. Removals apply to the live boards, at the next recompute, and to every future edition; the changelog records each removal without naming the brand. Editions already published stay as they were.
The changelog
Dated version history of the methodology. Two different things are recorded here, and they behave differently: a frozen edition record never changes after publication, so its verbatim answers, capture timestamps, per-board content hashes and the edition root hash stay exactly as they were published. A version bump changes how the current boards are scored, from the next refresh of each board onward. Where an older edition was re-assembled under a newer version, the entry below says so and gives the date it happened.
Show earlier changes (29)Hide earlier changes
What the Index is not
The Index reports what AI models recommend. It is not an endorsement by CiteHawk, and it is not a review site. If an AI model is wrong about a category, the Index will faithfully show you that it is wrong. That is the point.

Want this measurement for your own brand?
CiteHawk tracks the same signals for your brand across up to 8 AI models, every week, receipts included.
Free AI visibility report · No credit card · 50 prompts · 8 engines
