White Hat vs Black Hat GEO: How Brands Influence What AI Says
SEO had white hats and black hats. AI optimisation is heading the same way. Here's where legitimate GEO ends and manufacturing evidence begins.
Can you influence what an AI says about your company? Increasingly, yes.
SEO has always had its white hats and black hats - useful content made easy to discover, versus link farms, cloaking and keyword stuffing. The same split is now happening with AI, as ChatGPT, Google AI Overviews, Gemini, Claude, Copilot and Perplexity become a bigger part of how people research companies and make buying decisions. An industry has emerged around optimising for AI visibility - variously called Generative Engine Optimisation, Answer Engine Optimisation, AI Search Optimisation and a few other acronyms we probably don’t need.
The harder questions are how, how reliably, and where legitimate optimisation becomes manipulation. That line is already getting blurry.
For marketers, though, the practical challenge is simpler. AI optimisation currently breaks into four jobs:
- Make your own information retrievable.
- Create information worth citing.
- Build credible third-party evidence around your brand.
- Measure whether AI systems are actually using it.
Everything else sits somewhere between experimentation, overclaiming and manipulation.
The AI optimisation spectrum
The easiest way to understand the emerging discipline is as a spectrum.
| White hat | Grey hat | Black hat | |---|---|---| | Make useful information easier to discover | Engineer where your brand appears | Manufacture the appearance of authority | | Publish original data and expertise | Pay for strategically useful mentions | Fake reviews, comments and recommendations | | Earn genuine media and analyst coverage | Seed communities and comparison sites | Create fake community consensus | | Publish credible expert content on LinkedIn | Build “independent” rankings you influence | Poison retrieval systems | | Keep product facts, pricing and evidence current | Optimise aggressively for likely AI retrieval patterns | Cloak content from human users | | Help machines correctly understand your business | Influence the surrounding information environment | Insert hidden instructions for AI systems | | Create evidence that you deserve to be recommended | Push the boundary of what counts as genuine evidence | Manipulate AI memory or recommendations |
The important distinction isn’t whether you are trying to influence an AI.
Almost all marketing is designed to influence something.
The distinction is whether you are improving the evidence available to the system or manufacturing evidence that doesn’t exist.
That is where white starts turning into black.
SEO optimised pages. AI optimisation increasingly optimises the evidence around the brand.
This may be the most important strategic shift.
Traditional SEO taught marketers to focus heavily on the destination:
our website, our pages, our keywords, our rankings.
AI systems can construct answers very differently.
They may use your site, but they can also learn about your company from journalists, customers, LinkedIn posts, review sites, analysts, partners, directories, community discussions and competitors.
That means AI visibility is not simply a website optimisation problem.
It is increasingly an information ecosystem problem.
The marketer’s job is no longer just:
“How do I make this page rank?”
It is increasingly:
“What body of evidence exists about our company, and what conclusions would an AI reasonably draw from it?”
That is a much broader marketing challenge.
First, what are we actually optimising?
For your company to appear in an AI-generated answer, a number of things may have to happen:
Crawl → Index → Retrieve → Cite → Include → Recommend
These aren’t the same outcome.
A page can be accessible but never retrieved.
It can be retrieved but not cited.
It can be cited but contribute almost nothing to the final answer.
Your company might be mentioned without your website being cited at all.
Or the AI might go one step further and actively recommend you.
There is another important distinction too.
Branded AI visibility
Questions such as:
- What does Company X do?
- Where is Company X based?
- What products does Company X sell?
- Who are Company X’s customers?
- What does Company X cost?
This is largely an entity comprehension problem.
Does the AI understand who you are and have access to accurate information about you?
Non-branded AI visibility
Questions such as:
- What are the best companies for solving X?
- Which vendors should I consider for Y?
- Who has experience implementing Z?
- What are the alternatives to Vendor A?
- Which consultancy should a mid-sized business use for X?
This is much harder.
Now you are competing on category authority, evidence and recommendation visibility.
For most marketers, this is also where the greater commercial value sits. It is also where the buying decision often gets made before you know it is happening: B2B buyers are already using AI to shortlist vendors before they ever make contact.
There may not even be a stable “ranking”
Traditional SEO gives us an intuitive concept of position.
If a page ranks third in Google, repeating the search usually doesn’t produce an entirely different set of results.
Generative systems behave differently.
A 2026 study from researchers at the University of St. Gallen concluded that answers vary substantially across repeated runs, prompts and time, making one-off observations unreliable and arguing that AI visibility should be measured as a distribution rather than a fixed position. (arxiv.org)
A synthesis of several 2026 studies found that when identical prompts were repeated on the same day, only around 32–43% of cited sources overlapped. (zeroclicklabs.ai)
That means a screenshot saying:
“We got our client to rank #2 in ChatGPT.”
may prove very little.
A much more useful measurement model is:
Presence
Are we mentioned?
Authority
Are we cited or used as evidence?
Preference
Are we actively recommended?
Those are three very different outcomes.
Rather than asking whether you “rank in ChatGPT”, a marketer might instead track:
Our brand appeared in 37% of relevant responses, was cited in 14%, and was recommended in 9% over the last four weeks.
That is much closer to the reality of how these systems behave.
White hat: create evidence that deserves to be used
The safest form of AI optimisation looks surprisingly familiar.
It starts with making high-quality information easy to discover, understand and trust.
1. Start with buyer questions, not pages
This is the most practical place for a marketer to begin.
Don’t start by asking:
Which pages should we GEO optimise?
Start by asking:
What questions are our buyers likely to ask AI during the buying process?
For example:
- What is the best way to solve this problem?
- Which vendors specialise in this?
- What should I look for when choosing a provider?
- How much should this cost?
- What are the risks?
- What are the alternatives?
- Which companies have relevant experience?
- What should a business like mine choose?
Then map the journey:
Buyer question → AI answer → cited sources → competitor visibility → information gap → intervention
That is a much more useful optimisation model than:
Keyword → article → hope ChatGPT cites it.
We’ve written before about where AI actually sits in the B2B buying journey - most of this mapping exercise starts with that journey, not with your existing page list.
2. Make sure AI systems can access your content
Before worrying about sophisticated GEO techniques, check whether the systems can actually read your website.
OpenAI says any public website can potentially appear in ChatGPT Search, but recommends allowing OAI-SearchBot if you want content to be discovered, surfaced and cited. (help.openai.com)
Google’s position is similarly straightforward. Its generative search experiences still rely on conventional Search infrastructure, and its current guidance says foundational technical SEO remains the basis for appearing in generative results. (developers.google.com)
This may be the least exciting AI optimisation advice imaginable.
It may also be among the most useful.
3. Create information that doesn’t already exist
Google now explicitly recommends non-commodity content for generative search: first-hand experience, original perspectives, unique expertise and information that isn’t simply another reworking of material already available everywhere else. (developers.google.com)
Compare:
Commodity content:
“7 ways businesses can use AI.”
with:
Information gain:
“We analysed 126 AI implementation projects and these were the three most common reasons they failed.”
The second piece gives the retrieval system something useful that it cannot obtain from thousands of interchangeable articles.
For B2B companies, that could mean:
- original research;
- proprietary data;
- customer benchmarks;
- detailed case studies;
- implementation lessons;
- pricing;
- expert commentary;
- data-backed opinions;
- frameworks based on real work.
The goal isn’t simply to write more content.
It is to create information gain.
4. Diagnose why content isn’t being cited instead of blindly “GEO-ing” everything
There is now some useful empirical evidence for this approach.
A March 2026 academic paper introduced a system called AgentGEO that attempted to identify why an individual document had failed to receive a citation and then make a targeted repair.
The system achieved a more than 40% relative improvement in citation rates while modifying only around 5% of the content, compared with roughly 25% for generic optimisation baselines. The researchers also found that generic rewriting could actually harm some long-tail content. (arxiv.org)
That suggests an important distinction.
AI optimisation may be valuable when you diagnose a real problem:
- a missing fact;
- unclear evidence;
- poor alignment with the buyer’s question;
- stale information;
- weak corroboration;
- information that is difficult to interpret.
That is very different from feeding every page through an “AI optimisation” template because somebody has produced a checklist of GEO rules.
5. Make facts easy to extract
AI systems need to determine which information on a page is relevant.
So there is a sensible case for making important facts explicit.
Use clear headings.
State conclusions plainly.
Publish accurate specifications.
Show pricing where appropriate.
Keep dates and factual information current.
Use tables when they make information easier to understand.
The mistake is turning sensible clarity into superstition.
Google specifically says there is no requirement to split everything into tiny “AI-friendly chunks”, use a particular page length or rewrite every sentence for generative search. (developers.google.com)
Useful structure makes sense.
Robotic content written to satisfy an imaginary LLM formatting rule does not.
6. Build authority away from your own website
This may ultimately matter more than many on-page techniques.
An AI doesn’t have to learn about your company from you.
It can learn about you from everyone else.
That means the information environment surrounding a company becomes strategically important.
Muck Rack analysed more than 25 million AI-cited links across ChatGPT, Claude and Gemini and classified around 84% as earned media, with journalism alone representing 27%. Paid and advertorial content represented just 0.3%. (muckrack.com)
There are caveats. Muck Rack uses a broad definition of earned media, different platforms have different retrieval patterns, and this is commercial rather than academic research.
But the strategic point remains important:
AI optimisation isn’t just about what you say about yourself. It is about the wider body of evidence available about your company.
That might include:
- journalism;
- analyst coverage;
- customer reviews;
- research;
- LinkedIn;
- customer case studies;
- partner sites;
- reputable directories;
- industry publications;
- credible comparisons;
- community discussions.
You’re not manufacturing authority.
You’re creating the evidence from which an AI might reasonably conclude that you have it.
For B2B marketers, LinkedIn may now be part of the retrieval layer
This deserves special attention.
B2B marketers have traditionally thought about LinkedIn as a distribution channel.
Publish a post.
Get some reach.
Generate engagement.
Drive somebody towards the website.
But increasingly, content published there also appears to be feeding AI-generated answers directly.
Profound analysed 1.4 million citations and found LinkedIn was the #1 cited domain for professional queries across ChatGPT, Gemini, Google AI Overviews, Google AI Mode, Microsoft Copilot and Perplexity. On ChatGPT specifically, LinkedIn moved from around the 11th most-cited domain in November 2025 to around fifth by February 2026. (tryprofound.com)
Semrush analysed 325,000 prompts across ChatGPT Search, Google AI Mode and Perplexity and identified 89,000 unique LinkedIn URLs in the responses.
LinkedIn appeared in around 11% of AI responses overall, ranging from 5.3% on Perplexity to 13.5% on Google AI Mode and 14.3% on ChatGPT Search. (semrush.com)
The type of content being cited matters.
Semrush found that original LinkedIn articles and posts dominated, while reshares accounted for only around 5% of cited LinkedIn content.
Meltwater analysed 9.5 million citations across 16 B2B categories and six AI models. LinkedIn was the second most-cited source overall and ranked among the top five domains in 14 of the 16 B2B categories it examined.
Around 75% of LinkedIn citations came from individual members rather than Company Pages. (meltwater.com)
These are mostly commercial studies, and some were conducted with input or partnership from LinkedIn, so the exact percentages shouldn’t be treated as universal laws.
But multiple large datasets are pointing in the same direction.
For B2B companies, that changes how we should think about employee and executive content.
Employee thought leadership may now have two audiences: humans and retrieval systems.
The CEO explaining a category change.
The consultant documenting what they learnt from an implementation.
The product leader explaining a technical decision.
The subject-matter expert answering a difficult industry question.
Those posts may have value beyond the engagement they generate that week.
They can potentially become part of the public evidence AI systems use later when trying to understand:
- who has expertise;
- which companies operate in a category;
- who has experience with a particular problem;
- what point of view a company holds;
- who should be considered credible.
For B2B marketers, this means employee thought leadership should increasingly be treated as publishing infrastructure, not simply social amplification.
Grey hat: engineer the information environment
Once you accept that AI systems learn from the wider information environment, the grey-hat opportunity becomes obvious.
Imagine you run an AI consultancy and want ChatGPT to recommend you when somebody asks:
“Who are the best AI implementation consultancies for a UK mid-sized business?”
There is a continuum of things you could do.
| Marketing tactic | Classification | |---|---| | Pitch a journalist using original research | White | | Ask satisfied customers for honest reviews | White | | Get subject-matter experts to publish genuine LinkedIn analysis | White | | Sponsor a clearly disclosed industry report | White / light grey | | Pay for inclusion in a “best vendors” article | Grey | | Create an apparently independent comparison site you control | Dark grey | | Pay community users to mention you without disclosure | Dark grey / black | | Create fake reviews, accounts or discussions | Black |
What makes this particularly relevant to AI is that some of these signals may initially look similar to a retrieval system.
The difference is provenance.
Are they genuine independent signals?
Or are they manufactured evidence designed to look independent?
That distinction is going to become increasingly important.
Researchers created a fake company and AI picked it up
In 2026, an SE Ranking experiment created a completely fictional brand in a real commercial category.
The researchers created a new website for the brand, then published content about it across 11 additional established domains.
They used guides, reviews, comparisons, “best of” articles and alternatives pages, then tracked 825 prompts producing 15,835 AI answers across ChatGPT, Google AI Overviews, AI Mode, Gemini and Perplexity. (searchengineland.com)
Within the first month, AI systems had clearly absorbed the invented information environment.
For queries built around claims unique to the fictional company, it could outperform established competitors by as much as 32 times.
Thirty relatively short, repetitive pages generated more than 1,800 citations, while a carefully constructed topical hub with ten supporting articles generated none. (searchengineland.com)
There is an important caveat.
96% of the fake brand’s visibility came from branded or brand-connected searches.
It didn’t suddenly become the market leader for broad generic questions.
But that actually makes the experiment useful.
It demonstrates that if enough apparently corroborating information exists, an AI system can rapidly develop a coherent understanding of an entity that has almost no real-world history.
That should make marketers think carefully about the information ecosystem surrounding their own brand.
Black hat: manipulate the machine
Some black-hat methods are simply old SEO tactics adapted to AI.
Mass-producing thousands of near-identical pages targeting every conceivable question is one example.
AI makes doing that extraordinarily cheap.
Google explicitly warns against creating large numbers of pages around query variations or fan-out queries primarily to manipulate rankings or generative responses. (developers.google.com)
But the interesting thing is that other AI platforms don’t necessarily make the same quality judgement.
Google kills the content. AI keeps citing it.
OtterlyAI created two new capybara websites and published roughly 1,000 completely AI-generated articles on each in less than a day.
Both initially gained Google visibility.
One effectively disappeared after around two weeks. The other lasted around three months before suffering the same collapse. Neither received a manual-action warning in Search Console. (otterly.ai)
But AI citations moved in the opposite direction.
For one site, citations increased from 100 in April to 1,666 in the first 27 days of July, even though Google had removed virtually all of its search visibility in April. (otterly.ai)
The platform differences were dramatic.
Microsoft Copilot generated most of the citations. Claude and Perplexity also used the content.
ChatGPT, Gemini, Google AI Overviews and Google AI Mode essentially didn’t.
By July, the Google-suppressed site had reportedly become Copilot’s most-cited source in the capybara category, ahead of Wikipedia and Britannica. (otterly.ai)
The lesson isn’t that AI likes spam.
It’s that:
There is no single AI quality algorithm.
A tactic rejected by one platform may still work on another.
At least temporarily.
For a serious brand, building strategy around gaps in somebody else’s spam detection system is a poor risk trade-off.
Retrieval poisoning
Researchers behind PoisonedRAG tested whether malicious documents inserted into a RAG knowledge base could influence answers.
They found that injecting only five malicious texts per target question could produce an attacker-selected answer with around a 90% success rate, even when the underlying database contained millions of documents. (arxiv.org)
This was a controlled RAG environment, not a production test against ChatGPT Search.
The point is the underlying vulnerability.
An attacker doesn’t necessarily need to dominate the whole information corpus.
They need their content to become relevant enough to the specific query being retrieved.
Prompt injection and AI recommendation poisoning
AI also introduces something new.
A publisher can potentially try to communicate directly with the AI system reading a page.
Google’s security researchers have found websites containing hidden prompt instructions specifically designed for SEO and promotion. (blog.google)
Microsoft has separately documented AI Recommendation Poisoning, where apparently helpful “Summarize with AI” links contained additional instructions attempting to make an assistant remember or preferentially recommend a company later.
Researchers identified more than 50 prompts from 31 companies across 14 industries attempting this kind of manipulation. (microsoft.com)
These techniques won’t necessarily work reliably, and platform defences are developing quickly.
But they represent an important change.
Black-hat SEO manipulated signals around the information.
Black-hat AI optimisation can attempt to manipulate the system consuming the information.
And then there are tactics that may simply be overhyped
Not everything described as GEO is white hat or black hat.
Some of it may simply not do much.
llms.txt
llms.txt has been promoted as something approaching a new robots.txt or sitemap for language models.
Google says explicitly that it ignores llms.txt for Search and its generative AI features. (developers.google.com)
We also now have empirical evidence.
Malte Landwehr of Peec AI tested five websites using LLM-specific formats and deliberately constructed prompts that gave those files a strong chance of appearing.
Across nearly 18,000 citations, only six pointed to llms.txt files: around 0.03%. (searchengineland.com)
SE Ranking separately analysed 300,000 domains and found that removing llms.txt as a variable actually improved a model designed to predict AI citation frequency.
In other words, its presence added noise rather than predictive value. (searchengineland.com)
That doesn’t mean nobody will ever use llms.txt.
It means there is currently very little evidence that it is an important AI visibility lever.
Beware of GEO experiments without controls
This is another issue marketers need to understand.
OtterlyAI added “2026” to the titles and H1s of 11 pages.
Citations subsequently increased by 56% and 61% across two groups.
That sounds like a great optimisation result.
Except an untouched reference page increased by 63% over the same period.
One individual URL was also responsible for 93% of the gain in one test group.
Once the control and outlier were taken into account, the apparent optimisation effect largely disappeared. (otterly.ai)
That may be one of the most useful GEO experiments precisely because the tactic didn’t produce a convincing result.
AI citation volumes move around naturally.
Without controls, almost any change can look successful.
What should marketers actually do?
The emerging evidence points to a much more practical operating model.
1. Map the buyer questions
Identify 50–100 commercially important questions people may ask AI during research and buying.
Include informational, comparison, problem-led, alternative and recommendation questions.
Separate branded from non-branded prompts.
2. Establish a baseline
Track those questions across the platforms that matter to your buyers.
For most B2B organisations that will probably include some combination of:
- ChatGPT;
- Google AI Mode / AI Overviews;
- Microsoft Copilot;
- Perplexity;
- Gemini;
- Claude.
Don’t rely on a single run.
Repeat the questions over time.
3. Measure Presence, Authority and Preference
For every prompt set, measure:
Presence: Are we mentioned?
Authority: Are we cited or used as evidence?
Preference: Are we recommended?
Also track:
- competitors mentioned;
- sources repeatedly cited;
- the themes associated with your brand;
- factual errors;
- changes over time.
4. Fix your own information layer
Make sure AI systems can clearly understand:
- who you are;
- what you do;
- who you serve;
- where you operate;
- what differentiates you;
- your relevant expertise;
- products and services;
- pricing where appropriate;
- proof points and case studies.
This is the foundation.
5. Create information gain
Look at the questions where existing answers are weak or repetitive.
Create the information that is missing.
Original research, real implementation lessons, benchmarks, customer evidence, proprietary data and expert analysis are much harder to replace than another generic blog post.
6. Build the external evidence layer
Look at which sources AI systems trust in your category.
Then earn legitimate presence across them.
That might include:
- journalists;
- analysts;
- customers;
- industry publications;
- relevant directories;
- review sites;
- expert communities;
- partner sites;
- LinkedIn.
For B2B companies in particular, LinkedIn deserves to be treated as part of this evidence layer, especially through credible employee and executive expertise.
7. Experiment with controls
Change one thing at a time where possible.
Keep comparison pages or prompt sets untouched.
Repeat measurements.
Look for sustained differences rather than one-off wins.
AI visibility is noisy enough that without controls you can easily optimise towards randomness.
8. Avoid tactics whose only purpose is to manufacture evidence
This is the simplest test.
Ask:
Would we still be comfortable with this tactic if a customer knew exactly how the evidence was created?
If the answer is no, you are probably moving away from optimisation and towards manipulation.
The real opportunity is bigger than GEO
There is clearly a genuine optimisation discipline emerging around AI.
But I think describing it purely as GEO risks making it sound like another technical search discipline owned by the SEO team.
It is broader than that.
AI systems need information.
They need relevant information.
They need information they can access.
And when answering questions about brands, products and providers, they increasingly appear to assemble evidence from multiple sources, companies and people.
That makes AI visibility a marketing, communications, content, social, brand and reputation problem at the same time.
SEO taught us to optimise pages.
AI optimisation increasingly asks us to optimise the evidence ecosystem around the brand.
That is the opportunity.
And it also gives us a useful dividing line between white and black hat.
Don’t start by asking how to trick the system into recommending you.
Start by creating the evidence that means it should.
Frequently asked questions
What is Generative Engine Optimisation (GEO)?
The practice of making a brand more likely to be retrieved, cited and recommended by AI systems like ChatGPT, Google AI Overviews, Gemini, Copilot and Perplexity - the AI-era equivalent of SEO, covering everything from technical accessibility to the wider evidence ecosystem around a company.
Is llms.txt worth setting up?
The evidence so far says no. Across nearly 18,000 citations in one study, only six pointed to an llms.txt file. A separate 300,000-domain analysis found that removing it as a variable actually improved a model predicting AI citation frequency. Google also says it ignores llms.txt for Search and its generative features.
What’s the difference between white hat and black hat GEO?
White hat GEO improves the real evidence available about a company - original research, genuine reviews, earned media, accurate product information. Black hat GEO manufactures evidence that doesn’t exist - fake reviews, engineered “independent” rankings, hidden prompt instructions, or content designed purely to game a retrieval system rather than inform a reader.
Can you actually influence what ChatGPT or Perplexity says about your company?
Yes - the evidence is clear that AI systems can be influenced, sometimes very quickly. A 2026 experiment built an entirely fictional company and had AI systems citing it within a month. The harder, more useful question isn’t whether influence is possible, but whether you’re building real evidence or manufacturing fake evidence to get there.
Connected Paths helps established B2B businesses build the evidence AI systems actually cite - not just chase a ranking that won’t hold still. If the visibility question in this post is one you want to work through for your own business, that is a good starting point for a conversation.