Method

How we measure, and what we do not.

Whaily sends your prompts to model APIs and counts what comes back. This page states every number the product reports, the formula behind it, whether it carries a confidence interval, and the three places the method is weaker than it looks. If you are evaluating this category, this is the page to argue with.

Six numbers carry an interval.

A confidence interval says how far a number could move if the same question were asked again. Six numbers in Whaily come back with one. Five of them are a share of a countable denominator, which is what makes a Wilson interval the right tool. The sixth is a median rank, which needs a different method and gets one.

Share of voice
A 95 percent Wilson score interval under every line, for your brand and for each tracked competitor, wherever the share is shown: the window total, the per model row and the daily trend.
The difference between two shares
A 95 percent Newcombe interval (1998, method 10), built from each side’s own Wilson interval. The product says whether that interval clears zero. It never prints a p value.
Burst mention rate
A 95 percent Wilson score interval on the share of a burst’s runs that named the brand, overall and again for each model in the burst.
Burst sentiment shares
A 95 percent Wilson score interval on each sentiment label’s share of the burst answers that were analysed.
Burst cited domain shares
A 95 percent Wilson score interval on the share of a burst’s runs that cited a given domain at least once.
Burst median position
A distribution free interval from the binomial order statistics, which reports the coverage it actually reached rather than claiming 95 percent.

Eight do not, and here they are by name.

These are the numbers you will see most of, including the one on the front of the overview. None of them is drawn with a range, and the product shows nothing rather than a range it did not compute.

Visibility percent
The share of tracked answers in which the brand was named in the prose or its own site was recommended.
Inclusion rate
Of the answers that named any tracked brand, the share that named this one.
Average brand position
The mean place the brand took among the brands an answer named, over the answers that named it. Answers that did not name it are excluded rather than counted as a bad position.
Net sentiment
Of the answers analysed in a window, the share positive minus the share negative, in points, from minus 100 to plus 100.
Top three share
Of the answers analysed in a window, the share where the brand was named among the first three brands. Answers that never named it stay in the denominator.
Cited without being named
Of the answers that linked to a page on your own domain as a source, the share that did not name you anywhere.
NCI score
A 1 to 100 weighted score for one source in your category, built from ninety days of AI citations: how many models cited it, how many providers, the deepest single model’s count, how many distinct prompts sat behind those citations, how many models recommended it, and how recently it was last seen. Version 4.
Competitive rating
A 0.0 to 10.0 score a named model gives a brand on one purchase criterion, with a one sentence reason. Ratings from different models are kept as separate data points and never averaged.

Four of the eight are shares of a denominator, so an interval is computable for them and we have simply not shipped one. Average position, net sentiment, the NCI score and a competitive rating are not shares of anything, and each would need a different method rather than the one above. Until that changes, what you get is the number and the sample size behind it, with no range drawn around it.

Below five answers there is no number at all.

Every aggregate percentage is withheld below five answers in the selected scope. It reads N/A. Not zero, and not a triumphant 100 percent off a single run, which is what a brand new workspace would otherwise show on its first day.

One exception, and it is deliberate. A single point on a time series is exempt, because one run per model per day is the normal cadence and a trend line made entirely of N/A tells you less than a noisy one. So a daily point can sit on a denominator of one while the window total above it still reads N/A.

The four estimators, by name.

Wilson score interval, 95 percent

Five of the six numbers above use it. It is not the normal approximation: at a rate near 0 or near 100 percent the normal interval runs off the end of the scale and is the wrong width, and a burst that comes back with "no model named you in fifty runs" is exactly that case.

Newcombe interval, 95 percent (1998, method 10)

For the difference between two rates. It is built from each side’s Wilson interval, so it inherits Wilson’s behaviour at the extremes. A surface states whether the interval clears zero and stops there. No p value is printed anywhere in the product, because a p value invites "significant" as a verdict where the honest answer is a range.

Binomial order statistic interval

For a median position in a burst. A position is a rank, not a measurement, so a mean and a standard error would answer a question the data cannot. This interval assumes nothing about the shape of the distribution, and because the coverage it can reach is discrete it reports the coverage it actually achieved rather than claiming 95 percent.

A planning table, used before a burst and never on a result

When you size a burst, the panel shows what more runs would buy. That table is the normal approximation at the worst case rate of 50 percent:

  • 100 runs: within about 9.8 points
  • 250 runs: within about 6.2 points
  • 500 runs: within about 4.4 points
  • 1000 runs: within about 3.1 points
  • 2000 runs: within about 2.2 points

A rate near 0 or 100 percent is measured more tightly than this for the same number of runs, so the table is the worst case rather than a forecast.

Three things wrong with this.

None of these is a bug, and none of them is going to come up in a sales call unless we raise it. They are the places where a number here is softer than it looks, so they belong on the same page as the method.

1. The interval on a difference is narrower than it should be.

Newcombe’s method assumes the two shares are independent. On share of voice they are not, twice over: both counts come from the same answers, and they come out of the same denominator, so a naming counted for one brand is a naming not counted for another. The two move against each other, and that negative correlation makes the true error larger than the method assumes. The interval we print is therefore narrower than a correct paired one, which means it can call a difference real where a paired test would not. We state the interval and whether it clears zero. We do not present it as a significance test, and we are telling you this rather than waiting for you to find it.

2. The planning table promises slightly less than it delivers.

The table above is a fixed formula at the worst case rate, not a calculation on your actual rate, and it is the normal approximation rather than the Wilson interval the results come back with. It is therefore always a little wider than the interval you will eventually read. That is the safe direction to be wrong in, and it is still wrong: treat the table as a guide to the shape of the tradeoff, not as the answer.

3. The tightest row in that table is not one you can buy.

The table goes to 2000 runs. No plan allows a burst that large. The cap is on total runs in one burst, across every model in it:

  • Free: 5 runs
  • Starter: 20 runs
  • Pro: 60 runs
  • Agency: 80 runs
  • Enterprise: 200 runs

At 200 runs the worst case half width is about 6.9 points, and at 60 it is about 12.7. Because the cap is on the total, a burst spread across three models gives each model a third of it, and each model’s own interval is set by its own share of the runs. If you need a sample larger than the cap, that is a conversation rather than a button, and the honest thing is to say so here rather than let the table imply otherwise.

What we do not measure.

We call the model APIs, not the chat apps.

That is what makes a result comparable across versions and across days. It also means we do not measure what a logged in person sees in the ChatGPT web interface, Google AI Mode, Copilot or Meta AI, where a router, a system prompt and stored memory change the answer. Nobody measures those from the outside. We would rather name them than let the gap pass unmentioned.

A citation is not a mention.

A citation is a page the answer linked to as a source. A mention is your name appearing in the prose. They come apart constantly, and the case that matters is a brand whose pages are cited while its name is never said. URLs are stripped from the answer before any name matching runs, so a link to your site can never count as your name being spoken. Wherever a surface can show both, it says which one it is showing.

Burst answers stay out of your tracking numbers.

A burst is a deliberate experiment with a run count you chose. Letting those answers into the daily trend would mean a big test bent the line it was run to measure, so burst responses are excluded from the numerator and the denominator of every tracking metric.

Cited without being named is an observation, not an experiment.

The number says a model reached for your page and then did not say your name. It does not say that putting your name on that page would change a later answer. We have run no such experiment, and models change under us, so the metric is a comparison and the surface that shows it is worded the same way.

Nothing is recalculated in secret.

Every computed value stores the version of the method that produced it. When a method changes, the history is backfilled on purpose and the old version is retired by name. A score still carrying a retired version reads as no score rather than as a number to be compared with a current one.

If something here is wrong, we want to know.

This page is maintained against the code that computes the numbers, not against a slide. If you find a claim on it that the product does not honour, write to [email protected] and we will correct it.