DOCUMENTED
Directly supported by the platform provider’s own documentation.
RESEARCH · METHODOLOGY V1.0
We measure AI visibility as a chain — from technical eligibility through discovery, citation and recommendation to a business outcome. We freeze the prompts before we look at the answers, record every source and competitor, separate documented platform rules from our own experiments, and keep the losing results so that progress can be measured instead of asserted.
This page is the method. It is published in full because a measurement you cannot inspect is a claim, not a measurement. The last section shows what it returned when we pointed it at ourselves.
Seven things that are not the same thing
Almost every confusion in this field comes from treating two of these as one. A site can be perfectly crawlable and never be cited. A brand can be cited and never recommended. A recommendation can appear once and be nothing but model variability. We measure the steps separately because they fail separately and they are fixed separately.
Seven levels · every observation gets exactly one
“Visible” is not a measurement. Each run against each prompt is scored at one level, and the levels for cited and recommended are deliberately kept apart: a citation means an engine used you as a source, a recommendation means it told someone to hire you. Collapsing the two is the most common way this gets measured wrong.
The brand does not appear in the response at all.
Surfaced in a list or retrieval result, without meaningful descriptive treatment.
Named and described in the response.
Used as a source or shown as a linked citation.
Actively recommended for the query, not merely listed.
Reaches L4 repeatedly, across several engines or across repeated clean runs.
A visit, enquiry or client can reasonably be attributed to that visibility.
Why there is no “AI visibility score” here
A composite number is easy to sell and it hides the problem. Consider a site with strong crawler access, strong entity understanding, weak unbranded discovery, a citation rate of zero and a recommendation rate of zero. A score compresses that into something like 62 out of 100 — which tells you less than the five facts it replaced, and points at no action at all. We report the dimensions separately and let them disagree with each other.
Four rules · the reason a result means anything
A prompt is written down and frozen before anyone looks at the answer. After that it is an instrument, not a draft. Without these four rules a benchmark measures the person running it.
A single positive answer is a signal, not a victory
These systems are stochastic. Ask the same question twice and you can get two different answers with nothing having changed. So one good result is never reported as a win. It becomes meaningful when at least one of these also holds.
Five classes · applied before advice reaches a client
Every strategic statement carries one of five labels internally. The point is narrow and practical: it stops industry folklore from being repeated to a client as if it were a platform rule.
Directly supported by the platform provider’s own documentation.
Seen in our testing, but not established as a platform rule.
A reasonable reading of the evidence. Not a mechanism.
Insufficient evidence. Says so, and stays unused.
A deliberate test of something unverified, registered before it runs.
Six diagnoses · six different answers
When a competitor is named and you are not, the useful question is which kind of gap it was. These six have different remedies, and picking the wrong one is how budget gets spent on the wrong work.
The site cannot be reliably crawled, indexed or retrieved.
No page or entity directly matches what the query asks for.
The service is claimed, but there is no public proof, method, data or example behind it.
Competitors carry stronger independent mentions, reviews or ecosystem presence.
Competing material is newer or more actively maintained.
No stable evidence explains the difference yet. Often the honest answer.
Two corollaries we apply strictly: an authority gap is not solved with keywords, and an evidence gap is not solved with more schema.
The list that makes the rest of the page worth reading
Most of what is sold as AI search optimisation rests on one of the following. None of them is part of this method, and none of them is something we will tell you we can do.
Baseline 27 August 2026 · sealed 4 September 2026
We ran this method against our own site first. Across 50 controlled observations on 5 engines, Black Oar Studio did not appear once — including on prompts that describe what we actually sell. That is the honest starting position, and it is published here rather than kept in a drawer, because the “before” is what any later movement has to be measured against.
Two things we will not round off. One engine contributed only 2 of 12 runs, because we count a run only when the web search visibly executed — a toggle being switched on is not evidence that a search happened. And a sixth engine was run in full but is excluded from the count entirely, because its source panel was never captured, so we cannot show that it searched at all.
The full dataset — every prompt, every engine, every competitor named, and the source domains behind them — is not published yet. Two of the six engine records are incomplete, and we would rather publish late than publish a reconstruction. It follows as a separate piece.
Read this before trusting any single run — ours included
No manual benchmark reproduces every user's experience of an AI system. The same prompt can return a different answer for reasons that have nothing to do with the site being measured.
Which is why this method reports trends and repeated observations, and never presents one run as a permanent ranking.