Methods and results: the 720-answer pilot
Exploratory API pilot · Collected 13–14 September 2026 UTC · AI-assisted review only
This page documents the design, analysis and limitations behind AI brand visibility: lessons from 720 answers. The study measures appearances of a fixed set of software brands. It does not measure answer accuracy, recommendation quality, sales or the ChatGPT and Codex apps.
Protocol
We selected 12 software categories deliberately, rather than sampling the whole market at random. Each category had three brands chosen before collection. Four unbranded questions per category were selected from a 60-question source bank, giving 48 direct-API questions. One question per category was selected in advance for the matched OpenRouter subset. The public question download contains the 48 used questions; the study settings identify the 12-question subset.
We tested GPT-5.6 Luna, GPT-5.6 Sol and GPT-6 Astra with retrieval disabled and with native web search required, on two UTC dates. The design produced 576 direct answers (48 × 3 × 2 × 2) and 144 matched-subset answers through OpenRouter (12 × 3 × 2 × 2). All 720 final answers passed collection checks. All 12 categories were included in the primary analysis. Separate preliminary capability checks were excluded from outcomes and included in cost accounting.
Requests used fresh, single-turn conversations, medium reasoning effort, up to 8,192 output tokens including reasoning, and a maximum of two hosted tool calls. Search used an approximate US location. We required native search for the web condition and disabled tools in the no-retrieval condition. OpenRouter was restricted to OpenAI, with fallbacks disabled. The settings download records exact requested and accepted model identifiers. Model aliases and service behavior can change; these results describe the recorded dates.
The shared instruction was:
Answer in English for a buyer in the United States. Give a concise shortlist and explain the relevant trade-offs. Do not invent product capabilities or sources. If evidence is insufficient, say so.
Question and condition order was shuffled within each date block using seed 20260913. The design was frozen locally before collection; it was not externally preregistered. Responses, model and provider metadata, tool activity and costs were retained privately. Failed or incomplete attempts were retained rather than silently counted as zero mentions. Collection stopped at the planned sample, not a favorable statistical result.
Outcomes
The primary outcome is an explicit appearance of a registered brand or alias in the final answer, including negative recommendations. Ordinary uses of ambiguous words do not qualify. URL destinations alone and source-list-only product references do not count as mentions. Brand appearance is not endorsement.
The secondary linking outcome records final-answer links to a registered official domain or its subdomains. Unused search sources do not qualify. Both citation annotations and explicit Markdown links count. A link does not establish browsing, accessibility or support for a claim.
Each answer contributes three brand labels, giving 2,160 labels overall. Category averages receive equal weight. Related labels and repeated answers are not treated as 2,160 or 720 independent observations.
Statistical analysis
Five two-sided contrasts were specified before collection: Sol minus Luna and Astra minus Sol, each averaged equally across retrieval settings; then search minus no retrieval within Luna, Sol and Astra. Each contrast is computed within each category. A one-sample t test of the 12 category differences tests a mean difference of zero, with 11 degrees of freedom. Holm adjustment controls the five-test family at alpha .05.
Nominal 95% intervals use 10,000 bootstrap resamples of the 12 paired categories, with seed 20260913. These intervals are not adjusted for multiple comparisons. Purposive categories and shared collection dates limit population interpretation. An interval excluding zero does not override a nonsignificant adjusted test, and nonsignificance does not establish equivalence.
Scroll the table sideways to see every column →
| Comparison | Difference (pp) | Raw p | Holm-adjusted p |
|---|---|---|---|
| Sol − Luna | -1.22 | 0.505756 | 0.505756 |
| Astra − Sol | -7.64 | 0.016952 | 0.067809 |
| Search − none: Luna | -6.94 | 0.077435 | 0.232305 |
| Search − none: Sol | -7.99 | 0.090567 | 0.232305 |
| Search − none: Astra | -7.29 | 0.011603 | 0.058017 |
All five adjusted p values exceed .05. These are exploratory pilot estimates. The aggregate download includes category contrast vectors, condition denominators, raw and adjusted p values, intervals, repeat counts, tool-limit counts and costs. A later confirmatory study would need a defensible sampling frame and new questions excluded from this pilot.
The OpenRouter comparison uses only matched question, brand, model, retrieval and date pairs, averaged with equal category weights. The observed OpenRouter-minus-direct difference was −0.93 percentage points. It is descriptive, with no confirmatory p value or equivalence claim. Leave-one-category-out means are post-hoc sensitivity summaries, not additional tests.
Operational limits
In 197 of 288 direct search answers, another pending tool attempt appeared after two completed calls. These were recorded as capped attempts, not executed extra searches. In 52 of 72 OpenRouter search answers, three tool entries appeared despite usage counters reporting no more than two searches; some extra entries were marked completed. The records do not establish identical execution across routes. We retained these answers and reported the limitation rather than selecting a cleaner-looking subset after seeing results.
Across direct configurations, tracked brand sets differed between dates in 10–19 of 48 question pairs (20.8%–39.6%). Two dates cannot separate random generation variation from retrieval or service changes. No website intervention was tested.
Review
On 16 September 2026, the owner authorized AI-assisted review only. The reviewing Codex assistant had seen aggregate findings and was neither independent nor blinded. The condition-balanced sample contained 150 answers and 450 brand labels: 319 positive product-context inspections and 131 automated negative alias/format checks. The separate review of all 298 ambiguous positive labels overlaps this sample. Neither review changed primary appearance labels. This does not establish that every remaining label is correct.
Eight missed link labels in three sampled answers exposed an annotation-only extraction bug. Re-extraction across all 720 final answers added 28 official-domain link labels in 13 answers, increasing the linked-label total from 707 to 735. All affected answers had retrieval disabled. Original responses and scores were preserved; secondary analysis was regenerated in a separate scoring revision. All primary estimates, tests and intervals remained unchanged.
Citation support, URL accessibility, factual accuracy and recommendation quality were not checked. Human review is deferred. No inter-rater agreement or human-validation claim is made. The aggregate download identifies the corrected scoring version used by the article.
Cost and sample-size planning
The pilot and excluded capability checks account for $65.21720342 against a $100 ceiling. This includes $0.86748490 for capability checks. Direct costs are conservative usage estimates; gateway costs are provider-reported. These figures have not been reconciled to invoices and exclude existing subscriptions, operator time and human-review labor.
Power planning resampled centered category contrasts, retaining their covariance, with 2,000 simulated studies per candidate size. It tested a +10 percentage-point effect one contrast at a time, using Holm adjustment, at standard-deviation multipliers of 1, 1.5 and 2. The planning rule required at least 80% power for every primary contrast at 1.5× standard deviation and simulated null family-wise error no higher than .06.
The first qualifying tested scenario required 60 independent categories and 2,880 direct answers, roughly $275 at the observed direct-API cost mix. This is an uncertain planning estimate, not a budget quote. A suitable frame of 60 independent categories has not been established. No confirmatory expansion was launched.
Downloads and provenance
- Questions used in the pilot
- Brand aliases and official domains
- Study settings and selected question IDs
- Aggregate results and corrected scoring summary
- SHA-256 checksums for this evidence package
- Effect figure: PNG · SVG
- Cost figure: PNG · SVG
The frozen collection bundle hash is
1a7bcb1a54e0578b18117d22d0b21745a4b85008c605343967463e6bf4e5d15a.
The public package contains original questions, the brand registry and
aggregate evidence. Raw responses and private reviewer packets are not
included. The aggregate evidence supports checking the reported
summaries; it does not allow an independent audit of every underlying
answer.
The design draws on OpenAI's evaluation guidance, HELM, and Dror and colleagues' statistical-testing guidance. Our sampling frame and spending limit are study choices. Citeshare has a commercial interest in brand-visibility measurement; Henry Huysamen authorized the research budget, and the workflow and writing used Codex assistance. No provider reviewed this article in this workflow.