Hobby Answer ← Back to hobbyanswer.com
METHODOLOGY · PERMANENT PAGE

How we test what AI recommends

Transparent methodology is part of the product, not an appendix. Every paid Benchmark shows its work — and this page is the standing description of that work, so you can challenge it, re-run it, or hold us to it.

What every measurement shows

Every Hobby Answer Benchmark and public scoreboard includes:

  1. The exact prompts. The real buyer questions we asked, word for word. No paraphrases, no hidden prompt engineering.
  2. Test date and engine. When each question was asked and which engine and model answered, where the platform exposes it.
  3. Number of runs. AI answers vary run to run; where the variation materially affects the result, we run the question repeatedly and report the rate, not a single roll.
  4. Raw answer receipts. The actual answer text (or faithful captured excerpts), dated and kept on file — so "AI said X" is always checkable.
  5. Scoring definitions. How we count a mention, a recommendation, rank position, recommendation strength, and citations (below).
  6. Facts vs. hypotheses. What we observed is labeled as observation; why we think it happened is labeled as reasoned hypothesis. The two never blur.
  7. Limitations. Stated on the result itself, not buried here.

The engines

We test four major AI answer engines: ChatGPT, Claude, Gemini, and Perplexity through reproducible API checks, requesting web-grounded answers where the engine and model support them. Grounding availability can vary by engine, model, or run. A response without usable source evidence can inform mention and recommendation findings, but it is not counted as citation evidence. Each Benchmark identifies the engine, model where exposed, date, and evidence captured for the run.

How scoring works

MeasureDefinition
MentionThe brand or one of its named aliases actually appears in the answer text. Being in the question doesn't count.
RecommendationThe answer presents the brand as something the buyer should consider — not a passing reference.
Rank / positionAmong the distinct brands or products the answer recommends, the position of the tracked brand (1 = first or most prominent).
Share of answerAcross repeated runs of a question, how often the brand is mentioned or recommended — a rate, not a one-off.
Citations / sourcesThe pages and domains the engine searched or cited when building the answer — captured per run, because "where did it get that?" is usually the commercial question.

Honest limitations

Every public scoreboard we publish carries this line, and it applies to everything on this page: a result is a point-in-time visibility test, not a permanent ranking — and never a claim that one product is better than another.

What we will not do

Re-run it yourself

The fastest way to check our work: take any question from a scoreboard or report and paste it into ChatGPT, Claude, Gemini, or Perplexity yourself. Your answer may differ from ours — that's model variation, and it's exactly what the repeated-run rates are for. If you find something that looks wrong, tell us: blake@hobbyanswer.com. We'd rather correct a finding than defend one.

Hobby Answer · hobbyanswer.com · methodology last updated July 10, 2026