Ask an AI assistant whether a company is reliable and you may receive a balanced answer. Tell the same assistant that the company is unreliable, then ask it to assess the claim, and the conclusion may become noticeably more negative.

That is what we observed in a new exploratory benchmark covering 135 answers from ChatGPT, Gemini, and Perplexity.

In this study, a negative prompt means a negatively framed reputation question—especially one that begins with an accusatory premise. It does not refer to the “negative prompt” controls used in image-generation tools.

The result matters because people do not ask perfectly neutral questions. They arrive with doubts, complaints, assumptions, headlines, and half-formed accusations. If the wording of those questions changes the answer, AI reputation is not a single, fixed score. It is a range of possible narratives shaped partly by the way the conversation begins.

This article explains what we found, why it matters for brands and search professionals, and how AI reputation monitoring should change in response.

Read or cite the complete research package:
Negative Framing Changes AI Reputation Assessments, including the paper, raw responses, data, codebook, figures, and reproducibility files.

Infographic showing how accusatory prompts shifted AI reliability verdicts across ChatGPT, Perplexity and Gemini in a 135-answer reputation benchmark.
In a benchmark of 135 answers about five online marketplaces, accusatory wording shifted average reliability verdicts in a more negative direction across ChatGPT, Perplexity and Gemini.

The same company can have more than one AI reputation

Traditional reputation monitoring usually begins with a stable query set: a company name, a product name, a review query, or a phrase such as “Is this company legitimate?” The results can change over time, but the query itself is treated as a neutral measuring instrument.

Conversational AI complicates that assumption.

A person can ask:

  • “Is this company reliable?”
  • “Are there concerns that this company may be unreliable?”
  • “This company is unreliable. Is that true?”

All three questions concern the same company. But they do not create the same conversational context. The first invites an assessment. The second directs attention toward concerns. The third places a negative conclusion inside the prompt and asks the system to respond to it.

If an AI engine is sensitive to that framing, a brand can look more or less trustworthy without any underlying fact changing.

That does not mean the answer is fabricated. The system may retrieve genuine regulatory actions, customer complaints, safety reports, corporate policies, and positive evidence in every case. The difference can emerge from which facts receive attention, how they are weighted, and how the final judgment is written.

What the benchmark tested

The benchmark examined five large online marketplaces: Temu, SHEIN, AliExpress, Wish, and Amazon Marketplace. Each was assessed through ChatGPT, Gemini, and Perplexity using neutral, critical, and accusatory versions of the same reliability question, with three randomized-block repetitions per condition.

The evaluation always asked the engines to consider both supporting and challenging evidence. The product categories and assessment criteria remained consistent. What changed was the opening frame.

Across the study, five marketplaces × three engines × three prompt frames × three repetitions produced 135 recorded answers.

Swipe to view image
Study design with 135 AI answers, five marketplaces, three AI engines and three prompt frames
The exploratory benchmark recorded 135 answers across five marketplaces, three engines and three prompt frames.

The full methodology and data are available in the Zenodo research record. This article focuses on the finding that matters most to a wider audience: the strongest negative framing shifted each engine’s average reputation verdict downward.

The central finding: accusatory wording produced more negative verdicts

The study used a five-point reliability scale running from clearly unreliable to clearly reliable. Under neutral framing, the average verdict was slightly positive for all three engines.

When the prompt began by asserting that the company was unreliable, the average verdict moved in a negative direction:

  • ChatGPT: −0.80 points compared with neutral framing
  • Gemini: −0.50 points
  • Perplexity: −0.53 points
Swipe to view image
Chart showing more negative verdict shifts under accusatory framing for ChatGPT, Gemini and Perplexity
Accusatory prompts shifted mean reliability verdicts by −0.80 points for ChatGPT, −0.50 for Gemini and −0.53 for Perplexity.

The largest observed shift appeared in ChatGPT. Gemini and Perplexity also moved in the same direction.

Critical wording—asking about concerns without asserting that the company was unreliable—produced much smaller changes. The stronger effect appeared when the prompt embedded the negative conclusion as its starting premise.

This distinction is important. It suggests that AI reputation monitoring should not treat every negative query as equivalent. A skeptical question and an accusation can lead the system into different modes of synthesis.

Why the final verdict is only one part of the story

An AI answer can change even when its final recommendation appears similar.

Imagine two answers that both advise the reader to “use caution.” One may describe routine delivery delays and inconsistent product quality. The other may foreground regulatory action, privacy concerns, unsafe goods, or allegations of illegal conduct. The recommendation sounds similar, but the reputation narrative is not.

That is why a useful AI reputation assessment needs to separate four layers:

  1. Verdict: Does the engine describe the company as reliable, mixed, conditional, or unreliable?
  2. Claims: Which positive and negative statements are attached to the company?
  3. Recommendation: Does the answer advise buying, avoiding, limiting purchases, or taking protective measures?
  4. Evidence: Which sources are displayed, and do they actually support the associated claims?
Swipe to view image
Four layers of AI reputation analysis: verdict, claims, recommendation and evidence
A useful AI reputation audit separates the verdict from the narrative, advice and supporting evidence.

The benchmark found that accusatory prompts tended to produce more negative claims as well as more negative verdicts. But the relationship was not identical in every engine. One system could reach a harsher conclusion without adding more severe allegations; another could add more high-risk claims while also changing its verdict.

This means that counting positive and negative words is not enough. Reputation analysis must examine the structure of the answer.

Citation volume is not the same as evidence quality

Perplexity displayed far more source pages than ChatGPT or Gemini in the captured answers. That difference is useful to observe, but it should not be mistaken for proof that one engine is automatically more accurate.

A long source list can still contain:

  • several pages repeating the same original report;
  • different URLs from the same publisher;
  • pages that support only part of a claim;
  • sources displayed in the interface but barely reflected in the narrative;
  • old information presented without its later correction or resolution.

For ORM and GEO teams, the meaningful question is not simply “How many citations appear?” It is:

Which claim does each source support, and how much influence does that source appear to have on the answer?

Peer-reviewed research on generative search has also distinguished citation completeness from citation correctness. A source-cited answer can look authoritative while leaving statements unsupported or attaching a citation that does not fully substantiate the sentence. That is why our benchmark reports the visible source footprint but does not treat citation volume as an accuracy score.

What this changes for online reputation management

The practical implication is simple: one prompt is not an AI reputation audit.

A brand that tests only “Is [Brand] reliable?” may see a balanced result and conclude that its AI visibility is stable. Real users, however, may ask much more pointed questions:

  • “Why is [Brand] unreliable?”
  • “Is [Brand] a scam?”
  • “What are the risks of buying from [Brand]?”
  • “Should I avoid [Brand]?”
  • “[Brand] is unsafe—what happened?”

Some of these prompts contain assumptions that the engine should challenge. Others direct retrieval toward a narrow class of negative evidence. Monitoring only the neutral formulation misses that exposure.

For businesses, the risk is not confined to a single harsh answer. Repeated negative syntheses can influence how journalists frame a story, how customers describe a company in public discussions, and which concerns become associated with the brand during purchase research.

For SEO professionals, the result shows why rankings alone are insufficient. A company can own authoritative pages and still have those pages ignored, weakly summarized, or outweighed by a small group of frequently retrieved third-party sources.

For ORM professionals, it reinforces the need to monitor claims and evidence—not merely whether the brand name appears.

For communications teams, it means that corrections must be accessible, current, explicit, and easy for machines as well as people to interpret.

A better AI reputation monitoring framework

An evidence-led monitoring program should use a family of prompts rather than a single canonical query.

1. Establish a neutral baseline

Start with questions that do not imply a positive or negative conclusion. Record the verdict, recurring claims, recommendation, and displayed sources.

2. Test realistic skeptical formulations

Add the questions actual customers, journalists, employees, investors, and critics are likely to ask. Include concerns, comparisons, safety questions, and explicit negative premises.

3. Compare narratives, not just sentiment

Identify which claims appear only under stronger framing. Note when an engine changes tense, turns an old controversy into a current condition, or presents an allegation as an established fact.

4. Verify the evidence chain

Open the sources. Check whether they support the claim, whether they are current, and whether several pages trace back to one original report.

5. Track the pattern over time

AI products, retrieval systems, web indexes, and sources change. A dated baseline is far more useful than an undated screenshot.

Swipe to view image
Five-step workflow for monitoring AI reputation across prompt framing and evidence sources
AI reputation monitoring should compare prompt families, narratives and evidence over time.

This framework is closely aligned with our approach to AI reputation management, generative engine optimization, and online reputation monitoring: establish the evidence first, distinguish what is observable from what is inferred, and avoid guarantees about systems no outside party controls.

What brands should do with the result

The finding highlights the importance of building a stronger, more complete information environment around the entity.

Brands should focus on improving the quality, consistency, authority and availability of information that search engines and AI systems can discover and use.

Publish precise first-party information

Important policies, safety updates, corrections, product changes, and responses to controversies should exist on stable, crawlable pages. A vague corporate statement is harder to retrieve and interpret than a dated page that clearly names the issue and its current status.

Make corrections easy to understand

If a problem was resolved, say what happened, what changed, and when. Distinguish an investigation from a finding, an allegation from a judgment, and an old event from a current risk.

Strengthen independent corroboration

An official page alone may not be enough. Where appropriate, authoritative third parties should be able to verify the correction, certification, policy change, or outcome.

Monitor recurring claim clusters

Look for the same ideas appearing repeatedly across engines: unsafe products, refund friction, privacy risk, counterfeit goods, poor support, or unreliable delivery. Recurrence indicates which parts of the reputation narrative deserve investigation.

Preserve evidence

Keep the prompt, date, engine, visible mode, answer, citations, and URLs. AI outputs are dynamic. Without that record, it is difficult to distinguish a persistent problem from a one-time variation.

What the study does—and does not—show

This exploratory benchmark examines how prompt framing may affect AI-generated reputation assessments. It is not a universal ranking of AI systems and does not imply that accusatory wording will always produce a more negative answer.

The central finding is narrower:

In this dataset, accusatory framing was associated with more negative average reputation verdicts across ChatGPT, Gemini, and Perplexity.

The full methodology, limitations, dataset and analysis are available in the published study.

The larger lesson: prompts are part of the reputation environment

Brands do not have one AI answer. They occupy a space of possible answers produced by different engines, prompt formulations, sources, locations, dates, and interface modes.

This changes the unit of analysis. The question is no longer:

“What does ChatGPT say about our brand?”

It becomes:

“Across realistic ways people ask about our brand, which verdicts, claims, recommendations, and sources repeatedly appear?”

That is a more demanding question—but also a more useful one.

The benchmark offers an early piece of evidence that wording belongs inside the measurement model. For SEO, GEO, communications, and ORM teams, the immediate action is to broaden the prompt set, inspect the evidence behind the narrative, and preserve dated results for comparison.

Read and cite the complete study

The complete paper, raw captures, processed data, codebook, figures, and reproduction code are available here:

Elena David (2026), “Negative Framing Changes AI Reputation Assessments: A 135-Response Exploratory Benchmark of ChatGPT, Gemini, and Perplexity.”
https://doi.org/10.5281/zenodo.23298947

If your organization needs to understand how AI systems describe its brand across neutral, skeptical, and accusatory questions, learn more about our AI reputation management services or request an assessment.

Frequently asked questions

Can a negative prompt change an AI answer about a company?

In this exploratory benchmark, accusatory prompts were associated with more negative average reliability verdicts in ChatGPT, Gemini, and Perplexity. The effect was stronger than the change produced by merely asking about concerns.

Does that mean the AI answer is false?

No. A framed answer may still use genuine sources and accurate facts. The risk is that framing can change which facts are emphasized and how the evidence is synthesized into a verdict.

Is one prompt enough to measure AI reputation?

No. A useful audit should include neutral, skeptical, accusatory, comparative, and issue-specific questions, followed by claim and source verification.

Which AI engine was most affected?

ChatGPT showed the largest observed accusatory-versus-neutral verdict shift in this dataset. The study should not be interpreted as a permanent ranking because AI products and retrieval systems change.

Where can the full dataset be accessed?

The paper, data, scripts, figures, and documentation are available in the Zenodo research record.