AEO: Measuring Whether Claude, Codex and Grok Mention Your Product
Is your product cited when a coding agent answers a realistic question? That is the question behind aeo, a local command-line tool presented on GitHub by probelabs. It queries Claude Code, Codex and Grok with plausible questions and records whether your brand appears, whether the agent actually searched, and above all what it typed into search. The angle is interesting: instead of guessing, you observe real agent behavior, with a clear split between what the model "already knows" and what it goes looking for.
Two arms: knowledge and search
The tool asks each question twice. First with search disabled: this is the knowledge arm, testing what the model already believes without web access. Then with search allowed: this is the search arm, showing whether the agent searched and which queries it formulated.
This separation matters because the two arms tell different stories. According to the source, a new brand rarely wins the knowledge arm in its first year. The search arm, meanwhile, often looks like a confirmation of an already-named player rather than a discovery. The source presents these as general observations, without numerical data: treat them as a working hypothesis, not a law.
What the tool actually records
For each arm, the runner stores among other things:
- brand_mentioned: is the brand cited?
- searched: did the agent run a search?
- search_queries: the strings typed, verbatim.
- vendors_in_search_queries: which competitor names already appeared in those strings.
- token and spend information, when available.
The query detail is the most actionable part. If the agent never types your brand and your page is not in the backend being used, publishing more articles will not change anything. That is an important distinction: the problem is not always the content, it can be retrieval.
Prerequisites and limits to know
The tool requires Python 3.11 or higher, only the standard library, and the claude, codex and/or grok CLIs installed locally. The source advises against running these CLIs from a VPS or datacenter. Each cell runs in an isolated, empty /tmp directory, to prevent the agent from discovering the brand through its environment.
Several caveats deserve to be spelled out. The number of repetitions (--samples) is configurable, with a default of n=1, which limits the robustness of results. Local CLIs are described as slow, which can complicate large question grids. Finally, aeo is neither Gemini grounding, nor Google AI overviews, nor a connection to claude.ai: it is a local measurement bench, with the limits that implies.
The playbook: measure, publish, verify, re-run
The source proposes a simple sequence:
- Measure the current state with the tool.
- Publish one URL per topic cluster.
- Verify indexing of those URLs.
- Re-run only the relevant seeds.
That last point is essential. A table with no mentions does not mean you should write fifty articles. First verify each URL and distinguish three situations: confirmation (the agent already knows a player and confirms it), discovery (the agent finds new information) and blind spot (the agent never searches, or searches elsewhere). If pages are already live and mentions remain at zero, the source sees a retrieval problem, not a slug problem — without detailing how to prove it formally.
What to publish to get mentioned
The logic that emerges is less "produce volume" than "be retrievable". One URL per cluster, an indexing check, then a targeted re-measurement: it is a short loop, closer to diagnosis than mass production. For creators and e-commerce operators, the value is knowing whether content work actually translates into mentions, or whether the blockage sits upstream, in how agents access information.
Quick comparison of the two arms
| Arm | What it measures | Typical signal |
|---|---|---|
| Knowledge | What the model knows without search | Brand rarely cited in the first year (general observation from the source) |
| Search | Whether the agent searches and what it types | Often a confirmation of an already-named player |
Key takeaways
aeo offers a useful distinction between knowledge and search, and above all verbatim search strings that show why a brand is absent. The recommended method — measure, publish one URL per cluster, verify indexing, re-run the relevant seeds — is more a diagnostic protocol than a content recipe. The limits are real: default sampling at n=1, slow CLIs, and several source claims that remain unquantified. Use it as one measurement instrument among others, not as definitive proof.
This article is published by Roboto, a platform for generating texts, images, videos and voiceovers with AI. Roboto did not test the aeo tool; the observations reported come from the cited source.