version 1.0 · july 2026
R-Score Methodology
The R-Score is a normalized 0–100 composite measuring how likely a large language model is to recommend your brand in response to a buying-intent query. It is computed per surface, then blended with surface-level weights.
Overview
RecommendHQ is the AI Recommendation Management (ARM) platform. This is the four-step loop it runs for GEO: measure whether the AI recommends you, diagnose why it prefers rivals (the citation gap), act on the prioritized fix (the Action Engine), and prove the movement on the next scan. The R-Score and Recommendation Rate are how ARM measures; everything below is how those numbers are computed.
RecommendHQ operates on grounded AI surfaces only: ChatGPT Search, Perplexity Sonar, and Gemini Grounding (plus Google AI Overviews). These models retrieve live content at inference time, exposing citations. That citation graph is what sets RecommendHQ apart: we know not just whether a model recommends a brand, but why, and which upstream sources drive that recommendation.
Parametric-only models (frozen weights, no retrieval) produce a secondary "Brand Equity" signal that is honest about its slow-moving nature. It is never the primary deliverable.
Industry Metrics & Where the R-Score Fits
The GEO / AI-visibility industry has converged on a family of metrics. Each is useful; each answers a slightly different question:
- AI Share of VoiceHow often your brand appears in AI answers versus named competitors: appearances ÷ total category appearances.
- Share of ModelYour overall presence within one model's answers for a category: the "mind share" inside the LLM, observed per engine.
- Answer Inclusion RateThe percentage of relevant queries where your brand makes it into the final answer at all.
- Share of AnswerHow much of the answer experience you own: placement and space within the response, not just presence.
All four count appearances, not endorsements. A caveated mention (“X is okay, but…”) scores the same as an enthusiastic recommendation. The R-Score adds the missing layer: sentiment and prominence are scored in their own right, so an answer that warns about a brand can never score like one that recommends it. The companion metric, Recommendation Rate, reports the share of answers that actively recommend you: advice, not appearances.
The Four Signals
The R-Score combines four signals per AI surface: Presence, Prominence, Sentiment and Citation. Each contributes a fixed, independently-calibrated weight to the composite.
- PresenceIs the brand mentioned at all? Measured as an appearance-rate across N runs (not binary) to absorb LLM non-determinism.
- ProminenceHow prominently? Position in the response, whether named first, in a list, or as the primary recommendation.
- SentimentWhat tone? An LLM judge grades each mention on a 4-tier scale: an explicit endorsement, one option among several, a caveated warning, or absent entirely. A conservative rule-based classifier is the fallback if the judge is unavailable, and every result records which classifier produced it. A bare, unendorsed mention is never counted as an endorsement.
- CitationWhich sources validate the recommendation? The cross-referenced citation graph identifies the exact upstream nodes.
The R-Score Formula
For each surface s and over N smoothing runs:
R_surface = w1×Presence + w2×Prominence + w3×Sentiment + w4×Citation
R_blended = Σ(SurfaceWeight_s × R_surface_s) / Σ(SurfaceWeight_s)All sub-scores are normalized to 0–100 before weighting. The signal weights (w1–w4) and the per-surface weights are fixed per methodology version. They never change silently; any revision ships as a new version, dated on this page. Principles are public; the exact weights are proprietary.
One label rule matters more than any weight: your Visibility-Spectrum status is derived from the signals themselves, never from the score band. “Recommended” requires an explicit endorsement in the answer; “Cited” requires your own domain among the cited sources. A high number can never buy a label the underlying evidence doesn't support.
Surfaces & Weights
Grounded surfaces carry the large majority of total weight. Grounded = the model retrieves live content at inference time and exposes citations. This means the R-Score responds when a customer takes action on the citation graph. The number is influenceable now.
Surfaces in scope: OpenAI ChatGPT Search, Perplexity Sonar, Gemini Grounding, and Google AI Overviews. Between them they cover the two places buyers now meet AI answers: inside search (Google AI Overviews, plus the grounded Gemini answers that also power Google's AI Mode) and inside the assistants (ChatGPT, Perplexity, Gemini). Microsoft Copilot Search runs on the same GPT models we measure through ChatGPT, but we don't report it as a separate surface. We prioritize the grounded surfaces that carry the most B2B buying-intent share and expose citations; dedicated surfaces are added as their share and citation transparency grow.
The first three surfaces are read through their official APIs (OpenAI, Perplexity Sonar, and Google's grounded Gemini). Google AI Overviews has no official API, so we collect it through a licensed SERP-data provider whose terms grant pass-through rights to use the results in the Service and in reports, never by scraping Google directly.
Statistical Smoothing
Because LLM outputs vary from run to run, scores are smoothed across repeated samples rather than read from a single response. The system takes additional samples when a result sits near a classification boundary, so the score reflects genuine model behavior rather than one-off noise. Where a reading is built from a single run (see the free scan below), we say so. A single answer is a snapshot, not a trend.
Receipts & Provenance
Every score comes with its receipt. Each reading records the exact question asked, the surface and model that answered (models are part of the measurement identity: a budget model's answer is never silently substituted for a frontier model's), the date, the sources the answer cited, and which classifier judged the sentiment. Raw answers are archived, so any historical score can be traced back to the answer that produced it, including across model generations.
How the Free Scan Differs
The free scan is deliberately a single-run snapshot on one surface (Gemini Grounding). It derives a category buying-intent question for your domain (one that never names your brand, so presence can't be inflated by the question itself), runs it once, and scores the signals available in a single answer: your presence, your prominence and how the answer frames you (sentiment). The blended, smoothed, four-signal R-Score across all surfaces is the paid product. A Ghosted 0 on the free scan is a real result: the AI genuinely didn't surface you for a question your buyers are asking.
The Action Engine
After scoring, the Action Engine operates only on the citation graph of grounded answers, a closed, testable causal loop:
- Detect gap — brand absent where competitor is cited (G2, Reddit, Capterra…)
- Diagnose & prioritize — rank each gap by ROI: how much closing it can move your score versus how achievable it is
- Prescribe specifics — concrete steps, Entity-Authority-only tactics
- Close the loop — next scan proves movement, surfaced in the weekly digest
Guardrails
The Action Engine prescribes legitimate Entity Authority only: digital PR, original research, genuine reviews, community participation. Hard block on advice that constitutes astroturfing, fake reviews, or forum spam. This guardrail is enforced throughout our recommendation pipeline.