You should test AI prompts monthly because no analytics platform currently reports whether AI engines recommend your product, and the answer changes over time.
Search Console shows you queries that produced clicks. Ahrefs and SEMrush show you rankings. None of them show you what ChatGPT says when a buyer asks which tool to use in your category which is increasingly where the shortlist gets formed.
The gap is not a temporary tooling problem you can wait out. Even when proper measurement tools arrive, having six months of your own baseline data will be worth more than starting from zero on the day they launch.
The method below takes about fifteen minutes a month and requires nothing but a browser.
The 5 prompts
Run each in a fresh session with no prior conversation history, in both ChatGPT and Perplexity. Use identical phrasing every month changing the wording destroys the comparison.
Prompt 1: The core buying question
What’s the best [category] tool for [your ideal customer’s specific situation]?
Example: “What’s the best project management tool for a 50-person marketing agency?”
Phrase this the way a buyer would, not the way a marketer would. Buyers describe their situation; marketers describe categories. “Best project management software” is a marketer’s phrasing. “Best project management tool for an agency managing 30 client accounts” is a buyer’s.
What to record: Which brands appear, in what order, and whether the reasoning attached to each is accurate.
What it tells you: Your baseline visibility on the single highest-volume question in your category.
Prompt 2: The comparison question
Compare the top [category] platforms for [company size or industry]
Example: “Compare the top CRM platforms for construction companies with under 100 employees.”
This is the question a buyer asks once they have a rough shortlist and want to narrow it. The answer usually returns three to five vendors with a structured breakdown.
What to record: Whether you make the consideration set at all, and which attributes the model treats as the standard comparison axes
What it tells you: The second point is often more useful than the first. If the model consistently compares on criteria you do not address anywhere on your site, you have found a positioning gap
Prompt 3: The alternatives question
[Your largest competitor] alternatives]
Example: “Asana alternatives” or “alternatives to Salesforce for small teams”
This is the highest commercial intent query in B2B software. Someone asking it has already decided to leave a competitor and is actively looking for a replacement.
What to record: Whether you appear, your position, and which competitors appear alongside you.
What it tells you: If you do not appear here, this is usually the fastest available fix. A dedicated alternatives page targeting your largest competitor is one page of work against one of the highest-intent queries in your category.
Prompt 4: The evaluation criteria question
What should I look for when choosing a [category] platform?
Example: “What should I look for when choosing a customer support platform?”
This one does not name vendors, and that is the point. It reveals which evaluation criteria the model treats as standard for your category.
What to record: The criteria listed, in order, and whether your product’s actual strengths appear among them.
What it tells you: If your differentiator is not on the list, buyers are not being told to look for it which means your positioning is fighting the frame rather than fitting it. That is a marketing problem before it is a visibility problem.
Prompt 5: The direct question
Is [your product] good for [your primary use case]?
Example: “Is Notion good for managing client projects at an agency?”
The prompt most companies skip, and frequently the most revealing.
What to record: Everything. Not just whether you appear, but every factual claim made about you.
Three possible outcomes:
Nothing: The model has no meaningful information about your product. This is a visibility problem with a clear path.
Accurate: Right category, right use case, sensible comparisons. Good position.
Wrong: This is the one that surprises people. Discontinued features listed as current. Pricing from two years ago. Placed in the wrong category. Compared against companies you do not consider competitors. In regulated categories, security or compliance claims attributed to you incorrectly which for an enterprise buyer is a serious problem in either direction.
Being described inaccurately is different from being absent, and often more urgent. The model has confidently synthesized something wrong from stale or thin sources, and it will keep saying it until those sources change.
How to record the results
Build a spreadsheet with these columns:
Column: What goes in it
Date: Month of the test
Prompt: 1–5
Engine: ChatGPT / Perplexity / AI Overviews
Appeared? Y/N Position: Where in the list, if named
Competitors named : All of them, in order
Sources cited: From Perplexity, which shows them
Accuracy notes: Any incorrect claims about your product
The sources cited column is the most actionable and the most commonly skipped. Perplexity displays where it drew from. Over a few months, a pattern emerges: the same three or four source types appear repeatedly in your category. Those are where your effort should go.
What the results actually mean
You appear in 4–5 of 5: Prompts. Strong position. This is not your most urgent problem and your energy is better spent elsewhere. Keep testing quarterly to catch drift.
You appear in 2–3 of 5: Typical for a company with reasonable content and limited third-party presence. You are usually appearing on specific queries and missing broad ones, which is the normal shape. The gap is generally comparison content and review presence.
You appear in 0–1 of 5 while competitors appear consistently: The most common result, and the one worth acting on. In almost every case the cause is not content volume — it is content type and third-party corroboration.
You appear but are described inaccurately: Address this before working on visibility. Increasing the reach of a wrong description makes things worse, not better.
Two metrics worth tracking over time
Appearance rate: Out of your five prompts across two engines ten data points how many name you. Simple baseline.
Share of voice: Count every vendor mention across all ten. Yours as a proportion of the total.
Share of voice is the more honest number because it controls for category-level movement: If your appearance rate rises from 3 to 5 but your share of voice falls from 14% to 11%, you gained in absolute terms while competitors gained faster. Only the second number tells you about competitive position.
Common mistakes when running this test
Using a session with history: Prior conversation contaminates the result. Open a new chat every time.
Changing the phrasing between months: You lose the ability to compare. Write the five prompts down once and use them verbatim.
Testing only ChatGPT: The engines diverge, and the divergence is diagnostic. Strong in AI Overviews but absent elsewhere usually means your on-site SEO is fine and third-party presence is thin. Strong in Perplexity but absent in ChatGPT often means recent coverage without established presence.
Using marketer phrasing: Enterprise workflow orchestration solutions” is not what a buyer types. Use their words.
Testing once and concluding: These systems return variable output. Run the important prompts three times in separate sessions and note the pattern rather than the single result.
Only testing broad category queries: Generic queries are dominated by incumbents and are frequently not winnable. Your specific use-case prompts are where movement actually happens and where the intent is higher anyway.
What to do with what you fin
If you are absent while competitors appear, the fixes in rough order of return:
1. Build an alternatives page for whichever competitor came up most often. One page, highest-intent query.
2. Build comparison pages against the two or three competitors that appeared alongside you.
3. Make your basic facts extractable pricing, target customer, key integrations, stated plainly in crawlable text.
4. Check your robots.txt for blocked AI crawlers. This is a five-minute check that occasionally explains everything.
5. Address review presence, particularly recency. Recent reviews carry more weight than volume from years ago.
Then re-run the five prompts in eight weeks. Movement, when it comes, appears first on the specific queries prompts 3 and 5 before the broad ones.
Frequently Asked Questions:
How often should I run these prompts
Monthly is enough to see trends without generating noise. Quarterly is acceptable if you are already well positioned.
Do I need a paid ChatGPT or Perplexity account?
No. Free tiers are sufficient for this test, though paid accounts on Perplexity give more searches per day if you are testing a large prompt set.
Why do I get different answers each time I run the same prompt?
These systems produce variable output by design. Run important prompts several times in separate sessions and record the pattern rather than a single result.
Should I test Google AI Overviews too?
Yes, though they appear inconsistently depending on query type. When they do appear, they reach a large passive audience.
What if my product appears but the description is wrong?
Fix that before anything else. Document it with screenshots, identify the source (Perplexity’s citations help), correct your own pages and third-party listings, then re-check monthly. Updates take weeks, not days.
Can I automate this ?
Partially, via API, but the manual version is more reliable right now and takes fifteen minutes. Automation tools in this space are early and their results vary from what users actually see.