Cendar LabAEO. Marketing. AI engineering.Discuss your project
← Field notesAEO measurement

AEO measurement

How to Measure AI Citations and Decide What to Improve Next

Measure AI citations with a defined question sample, recorded test conditions and source-level review. Keep mentions, supporting citations, referral visits and business outcomes separate.

Four separate measurement layers: observed mention, cited source, recorded visit and confirmed business outcome.
Original Cendar Lab measurement diagram. These are different observations, not an automatically connected attribution funnel.

Start with a measurement question, not a dashboard

A B2B team might want to know whether answer systems can find its implementation guidance, whether the systems describe its offer correctly, or whether relevant inquiries arrive after people encounter the brand. Those are different research questions. The first concerns observable discovery and citations, the second concerns accuracy, and the third concerns a customer journey. A single visibility number can obscure all three. Write down which decision the measurement should inform before choosing the interface, sample or reporting frequency.

An actionable question could be: when we ask a defined set of integration-planning questions in a specified search-enabled assistant, does it cite our relevant guide, and does the citation support the associated statement? Another could be: after the revised guide is published, do we observe different supporting sources under comparable conditions? These questions have boundaries. They do not ask whether the brand is universally trusted by an entire model or whether every future user will see the same answer.

This guide proposes a practical observation protocol. It is not a report of live Cendar experiments or a standard mandated by a platform. Adapt the protocol to the decision, available access and acceptable data collection. A small, well-described study can be more useful than an elaborate dashboard with an unexplained denominator. Begin with a record you can audit and a method another reviewer can understand; automate only after you know which observations are meaningful and which are not.

Define mention, citation, support, visit and outcome separately

A mention is an appearance of the organization or product name in the answer. A citation is a reference to a source URL. Support is the relationship between a claim and the material on that cited page. A visit is a request or session your measurement system can observe under its collection rules. A business outcome is a separately confirmed event, such as a relevant inquiry or a signed agreement. One can occur without the others, so each needs its own field and interpretation.

Imagine a hypothetical answer that names a consultancy but cites a third-party directory. That is a brand mention, not a citation to the consultancy's website. Another answer might link directly to a technical guide without naming the company in the prose. That is a site citation even if a name-based search misses it. A third might cite the correct domain beside a claim the page never makes. Counting all three as equally successful would conceal important differences in both discoverability and fidelity.

Business reporting requires another distinction. Clicking a cited link does not make someone a qualified lead. Submitting a form does not mean a project is viable. Accepting a proposal is not proof of payment. Preserve these stages even when a stakeholder asks for a simple summary. A concise report can show several clearly labeled observations without collapsing them into an invented causal chain. The point is to make decisions easier, not to create the largest possible number from loosely related events.

Build a question set around real decisions

Choose questions that represent the audience you intend to serve. For a B2B AEO practice, that can include definitions, methods, limitations, implementation requirements and buying decisions. For an engineering service, it can include architecture tradeoffs or integration prerequisites. Use legitimate customer questions, existing content and available search data as inputs. Remove private details before putting any question into an external system. Do not begin by selecting only prompts that already produce flattering mentions of your company.

Separate branded questions from unbranded ones. Asking what Cendar Lab offers is useful for checking entity accuracy, but it is not the same discovery problem as asking how to scope an AEO engagement without naming a provider. Keep informational and provider-selection questions distinct too. A method guide may appropriately cite documentation without recommending any company. Label the question's purpose so that a technically sound answer is not scored as a failure merely because it did not behave like an advertisement.

Version the set before observing a period. Keep the exact wording, the reason a question belongs and the intended audience. If a prompt becomes obsolete or misleading, change it deliberately and identify the new version. Do not silently remove difficult questions halfway through a report. Add exploratory prompts in a separate group until you decide whether they belong in the recurring panel. This protects the denominator and lets readers distinguish genuine change from a more favorable selection of questions.

Record the conditions under which an answer was produced

An observation needs more than a screenshot of the final sentence. Record the date and time, the product or interface, any model label the interface exposes, the language, the question version and whether search or browsing was available. Describe whether you used a fresh conversation or supplied earlier context. Do not invent a model version when the interface does not show one. Unknown is an acceptable value when it identifies a real limit rather than hiding missing work behind a confident label.

Location, account state and personalization may affect an experience, but the degree depends on the product. Record relevant conditions that you can actually observe without exporting private sessions or credentials. A manual observation in a browser and a programmatic API request can use different retrieval behavior; do not treat them as identical merely because a brand name appears in both. If you need an API experiment, verify that it represents the product behavior your business question is about.

Repeat a question when the design calls for repeated observations, not until you obtain the answer you wanted. A single run can be informative but should not be presented as stable coverage. A budget-limited study can state that each question was observed once and explain the consequence. A repeated panel can report variation between runs. The right amount of repetition depends on variability and the decision at stake; there is no universal number of prompts that certifies a brand's standing in AI answers.

References: Google Search Central — AI features and your websiteGoogle Search Console — Performance report overview

Keep a compact observation record

A simple table is often sufficient to start. Give each question and observation a stable internal identifier, then record the visible answer and cited sources according to an approved retention policy. Separate the raw observation from the reviewer's classification. That lets another person revisit a disputed support judgment without rewriting the historical answer. Do not include account tokens, cookies, customer identifiers or unrelated conversation history in an exported research record, especially when several people need to review it.

For each reference, preserve the actual URL shown, the destination after redirects when checked, and the page title or other relevant context. A provider might reference a third-party copy, a translated version or an outdated path. Grouping at domain level can be useful for a summary, but keep the URL-level detail underneath it. Otherwise you cannot tell whether the answer selected the right guide, an obsolete page or merely the organization's home page beside a very specific technical claim.

Define statuses for incomplete observations. A timeout, a unavailable feature, a refused request and a completed answer without citations are not interchangeable. Mark an inaccessible cited page as unverified, not unsupported, until you can inspect it. Preserve the reason a run was excluded from a calculation. A report that distinguishes zero, unknown and not applicable is easier to trust than one that turns every missing cell into a numerical result for the sake of drawing a chart.

Suggested observation fields, not a platform-mandated schema
Field groupRecord
QuestionID, exact wording, intent, audience and panel version
ConditionsTime, interface, visible model label, language, search availability and conversation context
AnswerPermitted excerpt or reference, brand mention and cited URLs
ReviewClaim supported, partly supported, contradicted, unverified or not applicable
LimitErrors, exclusions, access gaps and reviewer notes

Check what each citation actually supports

Open the cited page and identify the statement the answer attaches to it. Ask whether the page supports that statement under the same conditions. If an article says that a service can be scoped after an API review, an answer claiming that the provider supports every platform without qualification overstates the source. A citation badge beside a sentence does not establish fidelity. The task is to evaluate the relationship between the claim and the source, not simply to confirm that a URL loads.

Use clear review categories. Supported means the relevant passage substantiates the claim within its stated limits. Partly supported means a material condition or portion is missing. Contradicted means the source says something incompatible. Unverified means access or evidence is insufficient to decide. Not applicable can be appropriate when the item is navigation rather than a factual claim. Write a short explanation for difficult cases and, when practical, have another reviewer check a subset without seeing the first reviewer's conclusion.

The classification is an editorial judgment, not an objective property delivered automatically by a screenshot. State the rules and acknowledge disagreement. If you use an AI system to assist review, it should not become the sole authority over whether another AI answer is accurate. Verify material claims against the cited source, especially for legal, financial, health or security subjects. Keep the original observation intact when correcting a classification so that the report remains a record of what happened and how it was assessed.

References: Google Search Central — Creating helpful, reliable, people-first content

Calculate sample statistics with honest denominators

Once the observation rules are stable, you can summarize the sample. A citation occurrence proportion might count completed observations that cite an owned URL and divide by eligible completed observations. A support proportion might count reviewed cited claims judged supported and divide by the cited claims actually reviewed. These denominators differ. State whether you count questions, runs, responses, citations or claims. Also state whether multiple URLs in one response count once or several times; either convention needs to be explicit.

For a deliberately hypothetical example, suppose a panel contains twenty completed observations and five cite a company-owned URL. Five divided by twenty is twenty-five percent of that defined sample. It is not twenty-five percent of all assistant conversations, all users, all searches or the market. If three of those observations concern the same branded question, say so. If several attempts failed, report them separately and explain the exclusion rule rather than hiding failed observations inside a more favorable denominator.

Avoid adding unlike proportions into a universal authority score. A brand mention, an accurate citation and a paid contract do not share a natural unit. You can create an internal prioritization rubric if its weights and purpose are transparent, but label it as your rubric, not a platform ranking. Stakeholders should be able to reconstruct a figure from its underlying observations. If that is impossible, the score is more likely to conceal uncertainty than to improve a marketing decision.

Use Search Console without inventing an AI-only filter

Google's documentation states that appearances in AI features such as AI Overviews and AI Mode are included in overall Search Console traffic, within the Web search type. The Performance report supplies familiar metrics such as clicks, impressions, click-through rate and average position. Those are useful for understanding the site's available search performance. They do not automatically isolate the contribution of AI answers. A rise in Web impressions cannot honestly be relabeled as a measured rise in AI citations.

Review dimensions such as page, query, country and device, and use a period appropriate to the question. The default report covers three months, but you should deliberately choose and document the comparison window. Newest data may be preliminary. Aggregation differs between the chart and some table views, which can explain discrepancies. A position metric is not a universal placement seen by every user. Google's documentation also notes that a query appearing in your report may produce different visible results when you run it yourself.

Keep missing access and missing data separate from a true zero. If a property has not been verified for your team, report that the source was not examined. If you export data, inspect how unavailable values are represented; Google's overview notes that some unavailable values can appear as zeros in exports. Before interpreting a change, check the source definition and aggregation. These details are not reporting trivia: they determine whether a number answers the question stakeholders believe it answers.

References: Google Search Central — AI features and your websiteGoogle Search Console — Performance report overview

Treat referral analytics as one observable path

Some visits from answer products carry a recognizable referring domain. Others may arrive without a useful referrer or may never occur because the person gets what they need from the answer itself. Analytics can describe the requests and events it lawfully observes; it cannot reveal every unseen citation or every offline influence on a buying decision. Keep the metric's name close to that reality. Observed assistant referrals is more defensible than total AI-driven demand when the latter has not been measured.

Choose campaign parameters for links you distribute deliberately, such as an opted-in newsletter, and avoid personal data in those parameters. An email address, contact ID or confidential project name does not belong in a public analytics URL. Do not add tracking tokens to a source URL solely to persuade an assistant to repeat them. A sensible page canonical should remain stable while the analytics system handles allowed acquisition context separately. Review collection, consent and retention before using the resulting information for profiling or sales automation.

A voluntary question in a commercial conversation can provide additional context about how a prospect found the company, but treat it as self-reported information rather than exact attribution. People may remember several sources or omit one. A prospect mentioning an assistant does not prove the assistant caused the purchase. Combining qualitative context with aggregate traffic and real pipeline events can inform decisions, provided you keep the different evidence types visible and do not silently merge them into a single causal claim.

Compare periods without confusing change with causation

Before changing a page, record the version, the intended improvement and the expected observable behavior. After an authorized release, verify the implementation first: the new passage exists, links resolve and the canonical is correct. Only then ask whether later discovery observations changed. A deployment timestamp is not the moment every crawler refreshed its copy. A response that still reflects an older page may be an observation about freshness rather than a reason to repeatedly rewrite the content without understanding the timeline.

Use comparable questions and conditions where possible, and document changes you cannot control. Platforms can update models, search behavior or interface features. Competitors may publish new material, and demand can vary with season or news. A before-and-after comparison can show a change in your observations; by itself it usually cannot isolate your edit as the cause. If you introduce a comparison group, explain why it is comparable and what differences remain instead of presenting it as a perfect experimental control.

Do not cherry-pick the first favorable result after an edit or the worst result before it. Report the whole agreed panel and the variability that matters to the decision. Separate newly added questions from the recurring group. If a page performs differently across products, preserve that difference rather than averaging it away. The objective is to decide what to investigate or improve next, not to manufacture a stable line chart from experiences that are genuinely variable and partly outside your control.

References: Google Search Central — AI features and your websiteGoogle Search Console — Performance report overview

Turn observations into a prioritized improvement backlog

A useful report ends with decisions tied to evidence. If an answer cites the wrong page, review whether your own information architecture points clearly to the intended source. If the answer overstates an offer, examine whether the page hides important conditions or contradicts another published description. If a relevant page is inaccessible, investigate the actual HTTP, crawl and network behavior. These are hypotheses to test, not proof that a single metadata adjustment will change every future answer.

Group findings by responsibility. Editorial work may require a clearer definition or better source placement. Engineering work may involve a redirect, canonical, template or access problem. Commercial work may require a more honest description of scope and availability. Privacy work may require removing unintended data from a public route. Assign a concrete acceptance check to each change and retain the previous version where rollback matters. A backlog item that simply says improve authority is not specific enough to verify.

Prioritize impact, relevance, confidence and effort without pretending to know search volume that has not been researched. An inaccurate description of a paid service may deserve attention before a low-stakes mention count. A blocked core guide may matter more than adding a new optional AI file. Conversely, an answer that accurately cites a competitor can be a useful signal about a content gap rather than a defect in your site. The remedy might be better original work, not a technical workaround.

References: Google Search Central — AI features and your websiteOpenAI — Overview of OpenAI Crawlers

Publish a report that remains useful when the result is mixed

A practical report can begin with the decision, scope and period, then summarize the panel, observations, interpretation and next actions. Include the question-set version, products observed, run counts, exclusions and review criteria. Show important examples with enough context to understand the judgment, while respecting source rights and privacy. Keep the underlying evidence accessible to authorized reviewers. A stakeholder should be able to distinguish the observation from the analyst's recommendation without reading a separate methodology document every time.

It is acceptable to conclude that the sample is too small, a source is unavailable or an effect is not yet clear. Those statements protect the business from false certainty. Avoid promising that a measurement service sees every conversation inside an AI product. Avoid implying that allowing a crawler proves the page was retrieved for a particular answer. Keep a clear distinction between actions your team completed, platform behavior you observed and business outcomes confirmed elsewhere in the organization.

For a first B2B cycle, choose a manageable panel, perform the review carefully and discuss whether the result changed a real decision. If it did not, refine the research question before building more automation. If it did, repeat the protocol at a cadence the team can sustain and expand only where added coverage is useful. The strongest measurement system is not the one with the most impressive score. It is the one that can explain what happened, what remains unknown and what evidence should guide the next action.

Sources and scope

By Cendar Lab. This guide references the documentation below; it does not imply vendor affiliation or a comparative test. Consult the scope note for the distinction between documented facts, recommendations and illustrative examples.

START WITH YOUR QUESTION

Let’s move your project forward.

Tell us what you want to improve. We’ll review your goals and discuss the next step.

A USEFUL FIRST MESSAGE

“We want clearer answers about our services in search and AI discovery. Where should we improve our content and technical setup first?”

What happens after you send it?We review your description and reply with questions about your project. No calendar booking or newsletter signup.

Your project, in a few sentences.

No technical brief needed to start.

Your project inquiry
Required
Required
Required

Share the goal, environment and constraints. AEO, SEO, email marketing, AI, cloud and software questions are welcome.

Add project details (optional)
Optional
Optional
Optional

Please do not send passwords, API keys, confidential documents, or sensitive personal information.