Three developers I know shipped the same feature last quarter. One of them found out about it through a support ticket from a stranger. The other two found out because a customer typed a question into ChatGPT, and the model named their tool. That gap in awareness is the whole problem worth solving. Search traffic used to tell you exactly where your product stood. You watched rankings, you read the query report, you knew which page pulled the signups.
Now a chunk of your audience asks an answer engine, gets a paragraph back, and never clicks anything. If your product is not inside that paragraph, the visit never happens and you have no row in your analytics to prove it went missing. Closing that blind spot is what a proper ai visibility platform is built to do, and honestly, most small dev teams still handle it with a spreadsheet and hope. I would not run a launch without some form of tracking in place, even a crude one. Here is how to actually build that tracking, what to measure, and where teams usually waste their time.
Why your old ranking report looks empty now
Classic SEO tools rank pages against keywords. Answer engines do something different. They read a spread of sources, weigh them, and compose a response. Mozilla's own documentation on how search indexing works is a fine reminder that crawlers and ranking systems reward structure and clarity, and answer engines inherited a lot of that logic while adding a summarization layer on top.
The practical result: your page can sit at position four and still never appear in the answer, because the model preferred a forum thread and a documentation page instead. So stop asking whether you rank. Start asking whether you get quoted. Those are separate questions, and only one of them maps to revenue right now.
What should you actually measure every week?
Pick four numbers and stay disciplined about them. More than four and you will stop checking the dashboard by week three.
- Mention rate: how often your product name shows up when someone asks a buying question in your category.
- Citation sources: which URLs the model leaned on when it mentioned you, and which ones it leaned on when it did not.
- Referral traffic from AI surfaces: sessions that arrive with no keyword and an unusual referrer pattern.
- Share of voice against two named competitors: not the whole market, just the two you lose deals to.
The mention rate is the number I care about most. It moves first, and it moves fast when you publish something the models find worth quoting.
Building a prompt set that mirrors real buying questions
This is the part teams get lazy about. They test "best tool for X" and call it done. Real buyers do not talk like that. They type whole situations. Write down fifteen prompts split across three buckets: discovery ("what do people use to solve X"), comparison ("product A vs product B for a small team"), and troubleshooting ("why does X break when I do Y"). Run them across ChatGPT, Claude, Perplexity, and Google's AI features because they do not agree with each other. Perplexity tends to cite sources inline, which makes it the easiest to audit. Claude is more conversational and harder to pin down. Google's AI layer pulls from a narrower slice of the web and skews toward pages it already trusts.
Once a month, rerun the same fifteen prompts and log the results in a table. Same prompts, same wording, same order, so the comparison holds up.
If you change the prompt wording every week, you are not measuring visibility. You are measuring your own mood.
The four tracking methods, ranked by effort
Manual checking is free and takes twenty minutes a week. Copy the answers into a sheet, note whether your brand appears, and note the sources cited. This works fine up to about ten prompts, and then it collapses. Reverse search is the second option. Paste your brand name into each engine and read what comes back. You will spot wrong claims, outdated features, and competitors positioning themselves as the replacement for you. This is where teams find the ugliest surprises.
Third is log analysis. Filter your server logs for bot user agents and unusual referrers, then look for patterns in what gets fetched repeatedly. That tells you which pages the models keep reading, which is a strong hint about what they consider authoritative about your product.
Fourth is a dedicated tracking AEO & GEO Tool. That is the point where automation pays for itself, since a platform can run hundreds of prompts daily, flag changes, and show you the citation trail behind each mention. Manual work does not scale past a certain size, and you will know when you hit that wall.
A concrete scenario worth copying
A two-person team I worked with sold a developer API. They assumed the model knew their documentation because they had published three integration guides. Testing showed the opposite. When asked about their category, the engine kept naming a larger competitor and citing a Stack Overflow thread from years earlier.
They did three things over six weeks. They rewrote their comparison page to state plainly what their API does, who it is for, and how it differs. They fixed their structured data so the product page and pricing page resolved as related entities, using the markup conventions described at Schema.org. They also got a maintainer of a popular open source library to link to their docs from a real tutorial.
By the end of the six weeks, their mention rate on the same prompt set went from occasional to steady, and the referral traffic carried users who already understood the product before landing. I would take that over a ranking bump any day, because those visitors arrive pre-educated and convert without a demo call.
Technical groundwork the models actually notice
Before you chase mentions, make sure the machines can read you without guessing. Clean semantic HTML, one clear H1 per page, real headings, no text baked into images. Fast responses matter too, and the baseline practices in the NIST guidance on software and system quality are a reasonable sanity check for how you serve content.
You can read the general material at NIST if you want a grounding in why reliability and clarity in what you publish keeps paying off. Answer engines favor pages that state a fact clearly in the first few lines, then back it up. So does your reader, which is convenient.
What to do in your first week
- Write fifteen prompts across discovery, comparison, and troubleshooting.
- Run them in four engines and log every mention and source in a shared doc.
- Fix the single worst factual gap the engines repeat about you.
- Add clean structured data to your product and pricing pages.
- Rerun the set in seven days and compare.
That loop costs you an afternoon and gives you something no analytics dashboard currently provides: a straight answer about whether the machines know your product exists. The teams that will win the next two years are not the ones with the biggest backlink profile. They are the ones who checked early, found the wrong answer, and corrected it before a customer did. When was the last time you asked an AI engine what it thinks your product does?
