Target outcome
A measurement-quality scorecard and governed keep, rewrite, merge, add, or review queue
Inventory the set
Audit the wording
Pressure-test coverage
Govern the changes
Run in Claude and Run in ChatGPT open a new browser tab with the prompt already typed. Nothing is sent until you submit it.
Your AI visibility report can only describe the prompts you monitor, so a repetitive or brand-leading set turns a precise dashboard into a distorted slice of the market. This audit stays inside that set and ends with a keep, rewrite, merge, add or review queue. Mining Search Console and Bing queries for new candidates is the neighbouring job.
Prompts get managed like keywords: add more, cut the underperformers. But a prompt set is a measurement instrument, not a traffic source. Low visibility often marks the blind spot most worth fixing, and three rewordings of one question overweight a theme. Wording carries the same risk: "Why is our product the best option?" tests agreement with a premise, not discovery. Provider results and observed searches turn those opinions into counts.
1. Inventory, then count. get_prompts returns one row per monitored prompt, 100 rows per request, so paginate. status is a processing state, the monitoring flag is the on or off switch, and no field carries a persona or a market. Then count what needs no interpretation: brand-name share, prompts per topic, prompts with zero answers, prompts left ungenerated.
2. Cluster, then check against providers. Grouping near-duplicates is inference and stays labelled so. get_matrix at prompt granularity returns at most 50 rows and no offset, covering part of a larger set. Clusters scoring alike on every provider are merge candidates.
3. Read the quadrant counts honestly. get_prompt_signals returns counts for attack, defend, monitor and ignore, but prompt lists for attack and defend only. Everything in ignore is a real result: the set is not separating contested questions from settled ones.
4. Compare against observed searches. get_query_fan_out returns the searches seen inside monitored answers with parent prompt IDs and brand presence, plus a priority label: very-high where competitors appear and your brand shows in under a fifth of that query's answers, then high, medium, low.
A real set of 155 monitored prompts spread across 8 topics returned coverage of 0 to 60 percent per topic, with five of the eight at 20 percent or below. Signals on the same set: 153 of 155 in ignore, average openness 99.6 percent.
The coverage column is the audit. A topic at 0 percent across five providers is not a weak position, it is a topic the set never really tests, and three of those beside one topic at 60 percent is the imbalance this workflow exists to find. Openness at 99.6 rules out the comfortable explanation: the questions are not locked up by incumbents, so the thin coverage belongs to the prompt set, not the market.
Three rules fix the order, and none is a judgement call. Rewrite first every prompt whose wording asserts your advantage, since every future report inherits that bias. Then merge, largest cluster first, freeing the most slots. Then add in the priority order get_query_fan_out already assigned, very-high before high before medium. Skip prompts with zero answers: nothing to compare yet.
Then cap the pass: leave at least two thirds of the set untouched in one cycle, or the like-for-like trend loses its base. Log what you held back with prompt ID, old text, owner and date.
The misreading to avoid is treating post-rewrite movement as progress. Changing the set changes what is measured, so a jump the week after a revision usually reflects the new sample. Until the unchanged prompts have comparable answers, read that baseline and nothing else.
I want to audit my monitored prompt set for measurement bias, redundancy and blind spots. List the brands I monitor and ask which to audit if there is more than one. If a call returns nothing, name the check now unavailable and carry on.
1. Paginate get_prompts until the reported total is covered, one row per prompt: ID, text, status, answer count, topic, tags, monitoring flag and last generation date, where status is a processing state and the monitoring flag the on or off switch.
2. Count, do not judge: how many prompts contain my brand name and what share that is, prompts per topic, prompts with zero answers, prompts last generated over 60 days ago. Flag brand concentration above half the set, and any topic above a third. Those thresholds are the audit's own, not Qwairy fields.
3. Cluster the prompts and mark literal duplicates, keeping separate any variant testing a different audience, geography, constraint or journey stage; label this as your reading. Then run get_matrix with prompt granularity and limit 50 over 28 days: it has no offset, so say plainly if my portfolio is larger, and name the clusters scoring alike on every provider.
4. Run get_prompt_signals with limit 20 over the same 28 days and report the attack, defend, monitor and ignore counts first. Prompt lists come back for attack and defend only, so if both counts are zero, say that no quadrant is actionable and move on rather than building a shortlist from another step.
5. Run get_query_fan_out over 28 days, paginate to the total, and map each query to its parent prompt IDs. List the very-high and high priority queries no prompt of mine covers. If nothing comes back, report that no derived searches were observed and drop the add candidates.
6. Produce one row per prompt: Keep, Rewrite, Merge candidate, Add candidate or Human review, citing prompt IDs, separating evidence from inference, never an automatic deletion. Order it as rewrites, then merges by cluster size, then adds in the fan-out's priority order, stopping once the queue would change more than a third of the set. Then ask which markets and buyer types the set must cover, since no field carries that, and list the gaps.
Connect Qwairy to Claude, pull the exact signals in the workflow, and leave with an execution-ready output.
Map monitored AI visibility across TOFU, MOFU, and BOFU. Isolate stage and provider gaps, inspect selected answers, and build an evidence-based content backlog.
Audit B2B SaaS category and comparison visibility, inspect the sources behind monitored recommendations, analyze configured keyword triggers, and build an evidence-based optimization plan.
Use Google Search Console and Bing Webmaster data through Qwairy MCP to discover, review, launch, and measure an evidence-backed AI prompt portfolio.