The Coe Lab
← Back to Blog

How Three Sites With 215,000 Fake Pages Are Poisoning AI Search Recommendations

September 3, 20266 min read
AIPerplexitysearch enginesspamtrust

A new investigation reveals that 60% of Perplexity's product recommendation citations point to obscure domains, with three sites publishing 215,000 machine-generated pages specifically designed to be cited by AI models.

When you ask an AI search engine for the best CRM, the best project management tool, or the best accounting software, where does it get its answers? A new investigation published this week reveals something deeply unsettling: nearly 60% of the sources powering Perplexity's product recommendations point to domains ranked outside the top 100,000 on the internet, and three obscure sites you've never heard of have collectively published over 215,000 machine-generated pages specifically designed to be cited by AI models.

The finding comes from an exhaustive study by researcher Jakob Greenfeld, who ran 380 software category queries through Perplexity's Sonar and Sonar Pro models via OpenRouter. The results paint a picture of an AI recommendation system that is being systematically gamed — and raise serious questions about whether any AI-grounded search engine can be trusted for product guidance.

The Experiment: 380 Categories, 7,534 Citations

On September 2, 2026, Greenfeld put 380 buyer-intent categories — ranging from broad staples like "CRM software" to niche tools like "museum collection management software" — to both Perplexity Sonar and Sonar Pro. Each call asked for a ranked top five with official homepage domains. All 760 calls returned parseable answers, complete with the URLs the models retrieved during grounding.

The scope is impressive: 3,800 recommendation slots naming 1,807 distinct products, backed by 7,534 citations spanning 2,055 distinct domains. Every cited domain was checked against the Tranco top-1M list and the Wayback Machine. The results were damning.

The Numbers: AI Search Is Citing the Wrong Sites

Here's what the data shows about where Perplexity's citations actually land:

  • 23.4% of all citations point to domains that don't appear in the Tranco top 1 million at all
  • 59.8% of citations point to domains ranked worse than #100,000 globally
  • The median Tranco rank of cited domains is 71,611 — these are not household names
  • Wikipedia, for context, was cited exactly three times out of 7,534 citations

Let that sink in. The world's largest encyclopedia, arguably the most comprehensive source of neutral product information on the internet, was cited less often than wifitalents.com — a site that didn't exist before December 2023.

The Third-Largest Source Is a Vendor's Marketing Blog

Perhaps the most striking discovery is guideflow.com, which ranks as the third most-cited domain overall with 194 citations across 96 of the 380 categories. That's more than Gartner. Guideflow sells interactive product demos — it's not a review site, a directory, or a publisher. It doesn't even compete in any of the software categories queried.

But Guideflow runs a massive content-marketing blog with 3,351 blog URLs listed in its sitemap. Its listicles about markets it doesn't operate in became the third-largest evidence base for product recommendations. There's nothing deceptive about what Guideflow is doing — they're just publishing content. The problem is what the retrieval layer does with it.

Three Sites, 215,000 Pages, One Operation

The investigation goes deeper. Three sites in the top ten — wifitalents.com (71 citations), worldmetrics.org (60 citations), and gitnux.org (50 citations) — appear to be a single operation. The evidence is damning:

  • All three were registered through NameCheap between December 2023 and May 2024
  • All three delegate DNS to the same pair of Cloudflare nameservers
  • All three run the same page template with identical navigation structure
  • Each site maintains a blog of exactly six posts, all cross-promoting the others
  • Together they've published 215,128 machine-generated "best [category]" pages

None of these three domains existed before December 2023. In less than three years, they've become cited more frequently by Perplexity than Capterra, LinkedIn, or Zapier — established platforms with editorial processes and human review.

Why This Matters for AI Search

This isn't just about Perplexity. The underlying problem affects every AI search engine that uses web grounding: ChatGPT's search, Google's AI Overviews, Gemini, and any model that retrieves web pages to support its answers. The issue is structural.

Traditional search engines like Google spent decades building ranking signals — backlinks, domain authority, user engagement metrics, editorial quality — that create a barrier to manipulation. A new spam site can't outrank Wikipedia overnight. But AI grounding pipelines don't use these signals the same way. They retrieve pages based on semantic relevance to the query, and if a page is explicitly designed to match the phrasing of a question like "what is the best CRM software," it gets retrieved regardless of whether anyone has ever heard of the site.

The result is a new attack surface. Instead of gaming backlinks, spammers game content volume and keyword alignment. Generate 215,000 pages covering every conceivable software category, and the AI will find you. It doesn't matter that no human has ever visited your site. The AI isn't checking traffic stats — it's checking whether your page semantically matches the query.

The Broader Pattern: AI Grounding Is Gameable

What makes this particularly concerning is that it's not an isolated incident. The same dynamic plays out across AI-powered recommendations of all kinds. Medical advice, financial guidance, legal information — any domain where AI models ground their answers in web retrieval is vulnerable to this kind of manipulation.

The researchers note that the unranked domains being cited are also significantly newer than established sources. The median first Wayback Machine capture for unranked cited domains is 2020, compared to 2011 for ranked ones. And 16.6% of archived unranked domains were first captured in 2025 or later — meaning they were created specifically to be found by AI retrieval systems.

This is a different kind of spam. Traditional SEO spam targets Google's crawler and ranking algorithm. AI-grounding spam targets the retrieval layer of LLM-powered search — a layer that is newer, less mature, and significantly easier to exploit because it lacks the decades of anti-spam development that went into traditional search ranking.

What Needs to Change

The fix isn't simple, but several approaches could help:

  • Domain authority weighting: AI grounding pipelines should incorporate traditional ranking signals like Tranco position, domain age, and backlink profiles when deciding which retrieved pages to cite
  • Source diversity requirements: If 25% of your citations for a single category come from one domain, something is wrong. Grounding systems should enforce citation diversity
  • Transparency: Users should be able to see the Tranco rank and domain age of cited sources alongside the answer, making it obvious when recommendations are grounded in obscure sources
  • Editorial source lists: For product recommendations specifically, AI search engines should prioritize established review platforms over random blog posts, even if the blog posts are semantically more relevant to the query

The Bottom Line

AI search engines are supposed to be better than traditional search. They promise synthesized answers backed by web sources, cutting through the noise to deliver trustworthy recommendations. But this investigation reveals that the grounding layer — the very mechanism meant to ensure accuracy — is being systematically exploited.

When 60% of your citations come from sites ranked outside the top 100,000, and the third most-cited source is a vendor's marketing blog, you don't have a search engine. You have a bullhorn for whoever is willing to publish the most content.

The AI search revolution is still in its early days. But if the grounding problem isn't solved, the promise of trustworthy AI-powered recommendations will remain exactly that — a promise. Until then, the best advice for anyone using AI search for product decisions is simple: check the sources yourself. You might be surprised by what you find.

Related Posts

Claude Fable 5.1 and Mythos 5.1: When AI Started Doing Real Science

Anthropic's new Claude Fable 5.1 and Mythos 5.1 models aren't just better at coding — they're designing proteins, mapping Venus, and optimizing GPU kernels for biologists. The gap between AI as a chatbot and AI as a research collaborator is closing.

Sep 2, 20267 min

When Security Cameras Meet AI: How BirdNET-Go Turns Your Yard Into a Wildlife Lab

A self-hosted AI system that listens through your existing security cameras and identifies birds, bats, and frogs in real time — no cloud, no subscription, no special hardware required.

Sep 1, 20266 min

Understanding ChatGPT Work: OpenAI's Most Powerful and Confusing Product Yet

ChatGPT Work gives you a headless browser, internet-connected code execution, persistent filesystems, and sub-agents. It's also extraordinarily confusing. Here's what we know.

Aug 31, 20266 min