Citelyra research · Published · Updated

How to measure AI search visibility: sample sizes and tests that give honest answers

Short answer

To measure how often AI answer engines cite your website, ask 200 or more different customer questions, 1 to 3 times each, and pool the results. To prove that a change worked, compare 40 to 50 changed pages with 40 to 50 similar unchanged pages before and after. Citelyra's simulation study found that before-and-after checks without comparison pages reported false wins up to 30% of the time.

Citelyra is an independent research project studying how AI answer engines such as ChatGPT, Perplexity and Google AI Overviews choose which websites to cite. This page publishes our first study: a Monte Carlo simulation of how AI citations behave, used to work out which measurement methods give reliable results. The results are about methods and sample sizes. They are not measurements of any real brand or AI engine.

Why is AI search visibility a probability, not a ranking?

AI answer engines generate a new answer each time, so the same question can cite your site in one answer and not the next. That makes your visibility a probability: the share of answers that mention or cite you. One check proves nothing. Repeated checks give an estimate with an error range.

Citelyra models each question as a Beta-Binomial process. After a site is cited in k of n answers, the likely citation rate follows a Beta(1 + k, 1 + n − k) distribution. For example, 3 citations in 10 answers gives a best estimate of 33%, and a 90% chance that the true rate is between 14% and 56%.

How many AI answers do you need to measure visibility?

With a fixed budget, asking more different questions gives a more accurate overall score than repeating the same questions. Repeats are only useful for finding the specific questions where you never appear.

Repeats per questionQuestionsError in overall score"Lost question" flags correct
16001.7 points26%
32002.0 points35%
10602.6 points59%
30203.9 points91%
Budget of 600 AI answers, moderate variation between questions, 3,000 simulations per row. Citelyra, October 2026.

Recommendation: run a broad first pass of 200 or more questions with 1 to 3 repeats, then re-run only the questions with zero citations about 10 times each.

Do before-and-after SEO tests work for AI search?

Not on their own. AI engines update their models often, and an update can shift every site's visibility at once. A simple before-and-after check mistakes that shift for the effect of your change. In Citelyra's simulation, when a page change had no real effect, a before-and-after check still reported an improvement:

Pages tested510204080
Without comparison pages7%13%21%25%30%
With matched comparison pages2%2%3%2%2%
Share of tests reporting a false win when the change did nothing. 2,000 simulations per cell. Citelyra, October 2026.

More data made the simple check worse, because it became more confident about the wrong thing. Comparing changed pages with similar unchanged pages (a difference-in-differences design) kept false wins near 2%.

How many page pairs does a test need?

Size of real effectPage pairs for an 80% chance of detecting it
Strong (2.7× more likely to be cited)About 20
Moderate (1.8×)About 40 to 50
Weak (1.35×)More than 80
Each page measured on 5 questions × 3 answers per period. Citelyra, October 2026.

Can machine learning find what makes AI engines cite a website?

It can predict citations, but it cannot reliably find their causes from ordinary website data. Large brands adopt new tactics first and are also cited more, so any new tactic looks effective. In Citelyra's simulation, llms.txt was given zero real effect. A logistic regression on observational data still reported it as an 85% boost, and called it statistically significant in 100% of 200 simulated studies. A gradient boosting model predicted citations well (AUC 0.73) but ranked llms.txt above content freshness.

A randomised test, where pages were chosen for a change at random, estimated the true effect almost exactly: 1.81× against a true 1.81×.

How should you estimate visibility for each question?

Pool the results. With only a few answers per question, raw counts are noisy. An empirical Bayes estimate that borrows information from all questions cut the typical per-question error by 30%, from 17.7 to 12.4 percentage points, in 1,000 simulations of 60 questions × 5 answers.

What does published research say helps AI citations?

Method

Citelyra simulated AI citation outcomes with a multilevel logistic model: logit P(cited) = baseline + question effect + page features + AI update drift + noise. The baseline citation rate was 20%. The spread between questions had a standard deviation of 0.5 to 1.5 on the logit scale. AI updates between test periods had a standard deviation of 0.5. Analysis used Python with numpy, scipy, statsmodels and scikit-learn. Other assumptions change the exact sample sizes, but did not change the direction of any finding.

Frequently asked questions

What is AI search visibility?
The share of AI-generated answers that mention or cite your website for the questions your customers ask.
Does llms.txt improve AI search visibility?
There is no evidence that it does. Google says its search does not use llms.txt, and an Ahrefs study of 137,000 sites found 97% of llms.txt files were never fetched by an AI crawler.
How often should I measure AI visibility?
Monthly is enough. AI answers vary from day to day, so daily tracking mostly measures noise.
Who runs Citelyra?
Citelyra is an independent research project that tests how AI answer engines discover and cite new websites. This site is itself part of that test.

Results on this page come from simulations, not from real AI answers. Citelyra will publish real measurements as they are collected. Data licensed CC BY 4.0: you may reuse it with credit to Citelyra.