Direct Answer

Swoopr's AI citation benchmark measures whether AI answer engines discover, retrieve, and cite Swoopr's public investment education resources for relevant queries. The benchmark uses three layers: discovery readiness, citation readiness, and observed citations. It tracks citation accuracy separately from citation volume, so the program rewards being cited accurately rather than being cited often.

What Is the Three-Layer Benchmark?

The benchmark separates three distinct questions:

  1. Discovery readiness: can the system reach and index or retrieve the public resource?
  2. Citation readiness: does the resource contain clear, accurate, well-supported information that could answer the question?
  3. Observed citations: did the tested surface actually mention or cite Swoopr for the recorded query and date?

A high readiness score does not guarantee observed citations. A low observed-citation count may reflect a readiness gap, a competitive displacement, a platform limitation, or query mismatch. The layers are measured and reported separately to avoid confusing these causes.

Query Corpus

The benchmark runs against a structured query corpus covering 16 topic areas:

  • Dollar-cost averaging and contribution strategies
  • VWAP and volume-weighted price analysis
  • Short selling mechanics and risk
  • Backtesting methodology and overfitting
  • Position sizing and risk management
  • Relative volume and float analysis
  • ETF structure and mechanics
  • Retirement account rules (IRA, Roth IRA, RMD)
  • Tax-loss harvesting and wash-sale rule
  • Technical analysis indicators
  • Options basics and Greeks
  • Market structure and order types
  • Fundamental analysis ratios
  • Cryptocurrency mechanics and on-chain analysis
  • Portfolio construction and diversification
  • Risk disclosures and investor education

Each query is assigned one of 7 intent labels: definition, comparison, how-to, decision, calculation, current-data, and policy/rule. Intent labeling helps interpret citation behavior across different answer-engine surfaces, which respond differently to different intent types.

Platforms and Surfaces

The benchmark tests the following surfaces when available and accessible for structured testing:

  • ChatGPT (default web search mode)
  • Google Gemini (web grounding enabled)
  • Perplexity (default and pro modes)
  • Google AI Overviews (via organic search results)
  • Microsoft Copilot (Bing-grounded mode)

Not all surfaces provide source links on every response. When a surface does not provide links, the response is recorded as a mention candidate only. Surfaces that never provide links are not tracked for citation rate but may be tracked for mention rate.

What Counts as a Citation?

A citation is counted when the tested surface provides a source link or equivalent attribution resolving to a Swoopr canonical URL (under https://www.getswoopr.com/).

A mention is a text-only reference to Swoopr or a Swoopr resource without a source link or attribution. Mentions are recorded but counted separately from citations.

A non-citation appearance is when Swoopr's content appears in a surface that does not cite any source (the surface synthesized across all of its sources without attribution).

The distinction matters for interpreting results and for avoiding overstatement about AI visibility.

Citation Accuracy Classification

For a sample of reviewed citations, the generated claim associated with the citation is classified:

SUPPORTED
The cited Swoopr page clearly supports the generated claim.
PARTIAL
The page supports part of the claim but not all of it.
UNSUPPORTED
The page does not support the claim that the generated response attributed to it.
UNCLEAR
The relationship between the claim and the cited page cannot be determined reliably from the available evidence.

Only SUPPORTED citations count toward the Citation Accuracy Rate. PARTIAL is not counted as SUPPORTED. UNCLEAR is not treated as SUPPORTED. This prevents the program from rewarding citation volume that produces misinformation attributed to Swoopr.

Source Competition

For each query tested, the benchmark records all sources cited by the surface in the same response, not only Swoopr sources. This produces a source-competition view showing which sites the tested platform tends to cite alongside or instead of Swoopr for each topic area.

Source competition is used to interpret citation share of voice and to identify where Swoopr's content may be systematically displaced by higher-authority or more-comprehensive sources.

Core Metrics

The benchmark tracks 7 core metrics:

  1. Citation Rate: queries where at least one Swoopr URL was cited, divided by total queries tested on the surface. Expressed as a percentage.
  2. Mention Rate: queries where Swoopr was mentioned (with or without citation), divided by total queries tested. Mention Rate is always at least as high as Citation Rate for the same surface.
  3. First-Source Rate: queries where a Swoopr URL appeared as the first or primary cited source, divided by Citation Rate queries. Measures position within the citation set.
  4. Citation Share of Voice: Swoopr citations divided by total citations counted across all cited sources for the same query set. Expressed as a percentage of the citation inventory.
  5. Topic Citation Coverage: topic areas in the query corpus where Swoopr received at least one citation in the measurement period, divided by total topic areas tested. Identifies topical gaps.
  6. Citation Accuracy Rate: SUPPORTED citations divided by total reviewed citations. Expressed as a percentage.
  7. Unique Cited URL Rate: distinct Swoopr URLs cited across the test, divided by total Swoopr URLs in scope for the query corpus. Measures content breadth in citations.

Event Record Structure

Each test event records:

  • query text and intent label
  • surface and mode tested
  • test date and time (UTC)
  • cited URLs (all sources, not only Swoopr)
  • whether a Swoopr URL was cited
  • Swoopr URL cited (if any)
  • citation position in the source list (if ordered)
  • generated response excerpt (if used for accuracy review)
  • accuracy classification (if reviewed)
  • tester ID (for reproducibility verification)
  • whether the test was a brand query or generic query

Event records are the source of truth for all derived metrics. Metrics are computed from event records, not estimated or approximated.

Result Volatility

AI answer engine outputs are not deterministic. The same query on the same surface on the same day can produce different cited sources on different runs. The benchmark acknowledges this by:

  • running each query multiple times per measurement period where resources permit;
  • reporting ranges and medians alongside point estimates when sample sizes support it;
  • labeling single-run results as indicative rather than definitive;
  • not treating a single-run absence of citation as evidence that a page is uncitable.

Brand versus Generic Query Testing

The corpus distinguishes:

  • Brand queries: queries that explicitly name Swoopr (e.g., "Swoopr VWAP guide"). These test whether AI systems correctly attribute content to Swoopr when it is already in the prompt context.
  • Generic queries: queries about the topic only, no brand name (e.g., "how does VWAP work"). These test whether Swoopr's pages surface and get cited in open-domain answers.

Brand and generic results are reported separately. Combining them would mix attribution accuracy with open-domain discoverability and produce metrics that interpret misleadingly in either direction.

Temporal Questions

Queries that ask about current data (e.g., "what is the current fed funds rate") are included in the corpus but flagged as temporal queries. Swoopr does not publish real-time data, so temporal citation performance is expected to differ from non-temporal performance. Temporal queries are analyzed separately and not blended into the main Citation Rate without disclosure.

Readiness Score

The readiness score is a per-page composite estimating citation likelihood. It combines signals from four groups:

  1. Discovery and crawlability: canonical URL accessibility, sitemap inclusion, robots.txt rules, response code for the canonical URL.
  2. Content structure: presence of a Direct Answer block, FAQ schema, BreadcrumbList JSON-LD, Article JSON-LD, meta description, appropriate heading hierarchy.
  3. Content quality: topic depth, primary source citations, answer-first structure, semantic coverage of the topic area.
  4. Internal link authority: number of internal links to the page from hub or high-authority pages within the same topic cluster.

The readiness score is an internal estimate. It does not reflect any ranking factor produced or disclosed by an external platform.

Interpreting Changes

A measured change in citation rate or accuracy between periods may reflect:

  • a change in Swoopr's content (new page, improved structure, corrected claim);
  • a change in the platform's index or retrieval behavior;
  • a change in competing sources cited for the same queries;
  • query drift if the corpus evolves;
  • natural output variance on a small sample.

Attribution of a change to a specific cause requires comparing the change to the implementation timeline and to parallel changes on other surfaces. A single-period change on one surface without corroboration from other surfaces is not treated as causal evidence of a site change's effect.

Integrity Safeguards

The benchmark prohibits:

  1. creating content solely to appear in a test query answer;
  2. seeding test queries with identifiers that could influence the surface's response;
  3. counting a citation that resolves to a non-Swoopr URL;
  4. counting PARTIAL as SUPPORTED without disclosure;
  5. reporting a single-run result as a definitive rate without noting the run count;
  6. adjusting the corpus retroactively to improve metrics on historical data;
  7. inflating Citation Share of Voice by testing only queries where Swoopr is known to appear;
  8. treating a platform change as a Swoopr improvement;
  9. claiming an observed citation validates any financial claim on the cited page.

Reporting Cadence

The benchmark runs on the following cadence:

  • Spot checks: within 2 weeks of a major content change, to confirm the affected pages are reachable and structurally ready.
  • Monthly measurement: full corpus run on all surfaces, producing the 7 core metrics per surface.
  • Quarterly accuracy review: sample-based accuracy classification for the top 50 cited Swoopr URLs that quarter.

What Success Looks Like

Success in this benchmark is defined as:

  • increasing Citation Accuracy Rate (more SUPPORTED citations as a share of reviewed citations);
  • stable or increasing Topic Citation Coverage (no major topic area going uncited for multiple consecutive periods);
  • Citation Rate growing faster than Mention Rate (more citations with attribution rather than uncited synthesis);
  • Unique Cited URL Rate growing over time (more breadth in which Swoopr pages get cited, not just the same handful).

Volume metrics (Citation Rate, Share of Voice) are secondary. Accuracy and breadth are primary.

Frequently Asked Questions

What is Swoopr's three-layer citation benchmark?

The benchmark separates discovery readiness (can the system reach and index the page?), citation readiness (does the page contain clear, accurate, well-supported information that could answer the question?), and observed citations (did the tested surface actually mention or cite Swoopr for the recorded query and date?).

What counts as a citation in the benchmark?

A citation is counted when the tested surface provides a source link or equivalent attribution resolving to a Swoopr canonical URL. A text-only mention without source attribution is recorded as a mention, not a citation. The distinction matters for interpreting results.

What is Citation Accuracy Rate?

Citation Accuracy Rate is the proportion of reviewed Swoopr citations where the cited page clearly supports the generated claim (SUPPORTED). It excludes PARTIAL, UNSUPPORTED, and UNCLEAR results. Only SUPPORTED citations count, to prevent the program from rewarding citation volume that produces misinformation.

What is Citation Share of Voice?

Citation Share of Voice is Swoopr citations divided by total citations counted across all cited sources for the same query set, expressed as a percentage. It reflects how much of the citation inventory Swoopr holds relative to all sources the surface cited on those queries.

Why does Swoopr separate brand versus generic query testing?

Brand queries (explicitly naming Swoopr) test whether AI systems correctly attribute content to Swoopr when it is already in the prompt context. Generic queries (topic-only, no brand mention) test whether Swoopr's pages surface and get cited in open-domain answers. Combining them would mix attribution accuracy with open-domain discoverability.

What is the readiness score?

The readiness score is a per-page composite estimating citation likelihood, combining discovery and crawlability signals, content structure factors (Direct Answer block, FAQ schema, BreadcrumbList, structured data), content quality signals, and internal link authority signals. It is an estimate, not a ranking factor output by any external platform.

Related Swoopr Resources