Target outcome
A claim-level accuracy register with an evidence-backed correction and re-test plan
Define the truth set
Extract the claims
Trace the evidence
Correct and re-test
Run in Claude and Run in ChatGPT open a new browser tab with the prompt already typed. Nothing is sent until you submit it.
An AI answer can name your brand, cite your own site and still present a retired feature as current. Read monitored answers claim by claim against your approved facts and what comes out is a ranked correction backlog.
Visibility reporting cannot catch this failure. A mention is not an accuracy judgment, a sentiment score is not factual support, and a citation does not prove the claim appears on the cited page. Presence gets celebrated while the error that costs a deal sits inside a favourable answer.
Qwairy records what an answer said, when, and which provider, model and URLs came with it. You supply the dated approved fact, and that separation keeps a verdict traceable.
1. Freeze the approved facts and resolve the brand. list_brands returns the monitored brands and their IDs. The truth registry is yours: one approved proposition per row, with scope, effective date and owner.
2. Inventory the prompts and answers. get_prompts returns prompt ID, text, status and answer count; get_answers adds provider, collection date, brandMentioned and sourceCited. Both cap a page at 100 rows, so paginate with offset. A false sourceCited means your domain was not cited, not that the answer lacked sources.
3. Pull the full evidence. get_prompt_answers lists the latest answers for one prompt, default limit 20, text trimmed to a preview. get_answer_details returns the full text, provider, exact modelId, competitors and the sources with isSelf.
4. Verdict each claim against the registry, then audit the sources. A returned source carries no page body, so a URL beside a claim is co-occurrence. get_source_urls and get_source_domains add aggregate context, with no provider or date filter.
5. Rank the failures and plan the re-test. Severity comes from the claim itself; recurrence is counted separately across distinct answers, prompts, providers and model IDs.
Save the baseline before correcting anything: brand ID, truth-pack version, prompt IDs, observed model IDs, baseline answer IDs and dates. Without it a later run has nothing to compare against, and correcting an owned page does not by itself make a future answer change.
Then work the backlog in one fixed order: contradicted claims about price, contract terms, security or a regulated capability first; contradicted or outdated claims about a product still on sale second; unverifiable and ambiguous claims to their named owner as questions, not corrections; non-factual claims closed with no action. Inside a tier, sort by the number of distinct answers carrying the claim, and break ties on the most recent answer date.
I need a claim-level fact-check of what monitored AI answers say about my brand. List the brands I monitor and ask which to use if there is more than one. Use the last 30 days unless I say otherwise, and label every row as observation, approved fact or analysis.
1. Resolve the brandId with list_brands, then ask me for my dated, owner-approved facts on products, plans, availability, security and legal terms. Build the registry from that material only and flag conflicts rather than picking one. If I give you nothing dated, stop: no verdict can be assigned.
2. Inventory with get_prompts and get_answers at limit 100, paginating with offset until the returned total is covered, keeping provider, createdAt, brandMentioned and sourceCited. If the answer total is zero, widen the window once, then report that nothing was collected rather than reasoning from prompt text.
3. On a cohort of purchase-critical and regulated prompts, call get_prompt_answers then get_answer_details, keeping answer text, provider, modelId, createdAt, competitors, source URL, position and isSelf. A prompt that returns no answers is a coverage note, not a clean bill: list it apart and move on.
4. Split each answer into atomic claims, keeping plan, region, time and quantity wording intact, then assign one verdict per claim against the registry for that answer date: supported, contradicted, outdated, unverifiable, ambiguous or non-factual. A claim with no matching approved record is unverifiable, and that is a finding. Never treat the answer or a cited URL as the authority.
5. For failed high-risk claims, treat the source URLs returned with the answer as co-occurrence only; get_source_urls and get_source_domains add aggregate context with no provider or date filter. An answer that returned no sources still gets its verdicts.
6. Count recurrence across answers, prompts, providers and model IDs, then order the backlog: contradicted claims on price, contract terms, security or regulated capability first; contradicted or outdated claims on products still sold second; unverifiable and ambiguous claims routed to their owner as questions third; non-factual claims closed. Sort inside a tier by distinct answers, tie-break on the most recent date. If nothing lands outside supported and non-factual, say the window is clean and skip the backlog.
7. Finish with a re-test sheet freezing prompt IDs, model IDs, baseline answer IDs and truth-pack version, then ask me who owns each contested fact and whether any dated release changes what is true. Say that correcting a page I own does not guarantee a future answer changes.
Connect Qwairy to Claude, pull the exact signals in the workflow, and leave with an execution-ready output.
Use Qwairy's brand perception, sentiment trends, and source analysis with Claude MCP to detect changes in AI framing, inspect cited evidence, and build a remediation plan.
Audit whether AI engines accurately represent your brand. Compare AI-generated descriptions with your intended positioning, track narrative drift over time, and build a correction strategy.
Audit the technical and content signals that support browser-capable AI agents, then build a structured, testable readiness roadmap with Qwairy MCP and Claude.