METHODOLOGY · PERMANENT PAGE
How we test what AI recommends
Transparent methodology is part of the product, not an appendix. Every paid Benchmark shows its work — and this page is the standing description of that work, so you can challenge it, re-run it, or hold us to it.
What every measurement shows
Every Hobby Answer Benchmark and public scoreboard includes:
- The exact prompts. The real buyer questions we asked, word for word. No paraphrases, no hidden prompt engineering.
- Test date and engine. When each question was asked and which engine and model answered, where the platform exposes it.
- Number of runs. AI answers vary run to run; where the variation materially affects the result, we run the question repeatedly and report the rate, not a single roll.
- Raw answer receipts. The actual answer text (or faithful captured excerpts), dated and kept on file — so "AI said X" is always checkable.
- Scoring definitions. How we count a mention, a recommendation, rank position, recommendation strength, and citations (below).
- Facts vs. hypotheses. What we observed is labeled as observation; why we think it happened is labeled as reasoned hypothesis. The two never blur.
- Limitations. Stated on the result itself, not buried here.
The engines
We test four major AI answer engines: ChatGPT, Claude, Gemini, and Perplexity through reproducible API checks, requesting web-grounded answers where the engine and model support them. Grounding availability can vary by engine, model, or run. A response without usable source evidence can inform mention and recommendation findings, but it is not counted as citation evidence. Each Benchmark identifies the engine, model where exposed, date, and evidence captured for the run.
How scoring works
| Measure | Definition |
| Mention | The brand or one of its named aliases actually appears in the answer text. Being in the question doesn't count. |
| Recommendation | The answer presents the brand as something the buyer should consider — not a passing reference. |
| Rank / position | Among the distinct brands or products the answer recommends, the position of the tracked brand (1 = first or most prominent). |
| Share of answer | Across repeated runs of a question, how often the brand is mentioned or recommended — a rate, not a one-off. |
| Citations / sources | The pages and domains the engine searched or cited when building the answer — captured per run, because "where did it get that?" is usually the commercial question. |
Honest limitations
- Model variation. The same question can produce different answers minutes apart. That's why we report rates over repeated runs and refuse to sell single-run certainty.
- Personalization and geography. Consumer apps can tailor answers by account history and location; our tests are clean-session and primarily U.S.-based unless stated.
- Change over time. Every result is a point-in-time receipt. Models update, sources shift, answers move. That is also exactly why the measurement is worth repeating.
- Source availability. Some channels are only partially visible to any tester (private communities, some social platforms). Where coverage is partial, the report says so.
- API and app differences. Consumer apps and their APIs can behave differently. Our results describe the tested API run; where an app comparison materially changes a finding, we label it separately.
- Attribution. Better AI visibility does not convert to revenue on a schedule anyone can promise. We measure recommendation movement and source visibility first, and connect commercial data where it genuinely exists.
Every public scoreboard we publish carries this line, and it applies to everything on this page: a result is a point-in-time visibility test, not a permanent ranking — and never a claim that one product is better than another.
What we will not do
- No guaranteed placement. Nobody can promise AI will recommend you. We measure, diagnose, and improve the evidence AI can find — then rescan and show the movement, whichever way it went.
- No pay-to-rank. Companies can pay to be measured, represented accurately, or analyzed more deeply. They cannot pay to improve their position on a public scoreboard. Commercial relationships are disclosed wherever a client appears in public work.
- No fabricated evidence. No invented prompts, answers, sources, reviews, or citations — ever. Anything our fact-checking can't verify ships flagged for human review, not published as fact.
- No undisclosed posting. Community reply drafts are disclosed (they open by saying the poster makes the product), and clients post from their own accounts. We never auto-post or astroturf.
Re-run it yourself
The fastest way to check our work: take any question from a scoreboard or report and paste it into ChatGPT, Claude, Gemini, or Perplexity yourself. Your answer may differ from ours — that's model variation, and it's exactly what the repeated-run rates are for. If you find something that looks wrong, tell us: blake@hobbyanswer.com. We'd rather correct a finding than defend one.