Last updated: September 2026
The problem this exists to solve
Answer engines aren’t consistent. Ask ChatGPT the same question twice and you can get two different sets of sources. This one fact breaks most of what gets published as an AEO experiment. One article, one change, one before-and-after, called a win or a loss based on what happened once.
A result like that can’t separate your change from the engine rolling differently that day. So nothing here is tested that way.
The method: Pattern scan
Every hypothesis runs the same way by logging a large sample of citations that already exist and looking for a pattern across all of them.
The inconsistency that breaks single-page tests is what makes this approach work. Run enough queries and the run-to-run variation averages out, leaving whatever is stable underneath.
Here is the process, every time.
- Write the hypothesis as one sentence that can be wrong. The hypothesis is a prediction, written down before any data comes in, and it has to be specific enough that a result could contradict it.
- Write the classification rule before collecting any data. The exact rule for sorting a source into one group or the Other, written down before I see a single result. It doesn’t move once the data comes in.
- Choose the queries before logging. Public, relevant, tied to the hypothesis. The list gets fixed before collection Starts, so I can’t drop queries that produce inconvenient results and keep the ones that don’t.
- Log every citation. For each query, every source the engine cites gets recorded and classified using the rule from step 2. Nothing gets skipped for being awkward to classify. Worked in batches, so mistakes surface early instead of at the end.
- Count and compare. Percentage of citations in each group, calculated once logging is complete.
- Apply the thresholds below. They’re fixed before the test runs. If the result lands somewhere inconvenient, it gets published where it landed.
- Publish the result and the full dataset. Every experiment links to the data underneath it. If you want to check the work, you can.
What count as a citation
A citation is one source, the engine names or links in its answer to one query.
- The same domain cited across ten different queries counts as ten citations, one per query.
- The same domain cited twice within one answer counts once.
- A source named without a link counts. A source linked without being named counts.
Recorded with each one is the query, the date, the engine, and the classification.
Which engines, and when
Each experiment states which engine it ran against and the dates it ran. Models change without notice, so a result is a snapshot of a specific engine in a specific window, not a permanent finding. Where a claim matters enough, I re-run it later and publish both.
Minimum sample size
Every group in a comparison has to clear 30 citations before a grade gets assigned. Below that, one outlier swings the whole percentage and the result isn’t stable enough to call a pattern.
If a group hasn’t reached 30, the experiment is marked Pending. It doesn’t get graded early.
Thirty is a floor, not a target. Most tests run well above it, and the sample size is published with every result so you can weigh it yourself.
The five grades
Every hypothesis gets exactly one of these. No labels invented after the fact.
| Grade | Meaning | When it’s assigned |
| Confirmed | Pattern held up clearly | Gap of 15 points or more in the direction the hypothesis predicted, both groups at or above 30 |
| Busted | Pattern ran the other way | Gap of 15 points or more in the opposite direction, both groups at or above 30 |
| Partly | True in some conditions, not others | Gap between 5 and 15 points, or the effect appears only in a subset of queries. |
| Inconclusive | Not enough signal either way | Gap under 5 points with both groups at or above 30 |
| Pending | Still collecting data | A group hasn’t reached 30 yet |
A hypothesis that predicts a difference and finds none is Inconclusive. One that finds a difference pointing the wrong way is Busted. Those are different results and they get different labels.
What the thresholds can and can’t do
These are raw percentage-point gaps, not statistical significance tests. At samples near the 30 floor, a gap of five or ten points can appear by chance. That’s the reason the Confirmed threshold sits at 15 rather than somewhere lower, and the reason Partly exists as a category instead of rounding a small gap up into a win.
Read Confirmed as “a gap large enough to be worth your attention,” not “proven.”
What this doesn’t prove
Every grade on this site means “associated with.” Never “causes.” If FAQ formatting shows up more often in cited sources, that doesn’t mean FAQ formatting is the reason. Sites with FAQs might just be better-structured overall, more established, or better optimized in ten other ways that aren’t being isolated for. This gets said on every result, not only the ones where hedging is convenient.
Why nulls get published
A test that finds no gap is still a result, and it goes up with the same weight as one that does. A site that only publishes its wins isn’t testing anything. It’s marketing with extra steps.
What “experiment” doesn’t mean here
One article watched over time, cited in two weeks or three months, is an observation. It’s one subject, not a sample. Those live in Field Notes.
When an observation turns into a hypothesis worth testing at scale, it moves here and gets run properly.
When the rules change
Any change to the process, the thresholds, or the grades gets logged below with a date. If the rules move, you’ll see when and why.