Three coordinated sites manufactured 215,128 machine-generated “best software” pages — and Perplexity has been citing them as authoritative sources. A Trellner Research report published this week found that worldmetrics.org, gitnux.org, and wifitalents.com built a coordinated operation to infiltrate AI-generated tool recommendations. Their pages carry a literal HTML title of “BrandName — Facts & Grounding Page,” written for parsers, not people. If you have asked an AI tool to recommend a database, a monitoring solution, or a CI/CD platform in the past year, there is a measurable chance you were reading fabricated AI content farm rankings.
How AI Recommendation Content Farms Work
The three domains share identical Cloudflare infrastructure, matching page templates, and were all registered within a six-month window between December 2023 and May 2024. Collectively, they generated 215,128 buying guide pages across 380 software categories. Each page lists named “staff writers” — nine distinct people credited across the cluster for a single software category — whose bylines appear to be fabricated. Some pages still contain unrendered template placeholders reading “Within the next 26 days,” visible in the final published output. Quality control was clearly not the point.
The rankings are also internally contradictory: Gitnux’s top pick for project estimation software does not appear in Worldmetrics’ top five for the same category. This inconsistency is invisible to an AI retrieval system. Perplexity’s RAG pipeline retrieves pages that match a query and synthesizes them into a confident answer — it does not compare sources for internal consistency. The content farms optimized for exactly the signals AI systems reward: machine-readable page structure and retrieval-friendly metadata. Making the content accurate or consistent was never the goal.
When a Marketing Blog Outranks Gartner in AI Citations
Trellner queried Perplexity across 380 software categories, collecting 7,534 citations from 2,055 distinct domains. G2 led with 291 citations — a legitimate peer-review platform. Reddit ranked second with 261. Third was Guideflow.com, a vendor’s own marketing blog for demo software, with 194 citations across 96 categories. Gartner, the most authoritative independent B2B analyst in the industry, appeared just 158 times. Wikipedia appeared three times.
That finding resonated on Hacker News, where the report generated 446 points and 215 comments. A vendor’s marketing blog outranking Gartner in AI-generated recommendations is not a quirk — it is proof that AI retrieval systems do not weight source authority. They weight retrieval signal. Content farms understand retrieval signal, and they are optimizing for it aggressively.
Related: Microsoft ThinkingBox Exposes the AI Agent Reliability Gap
The Pre-Panda Era of AI Search
The broader citation data makes the scope clear. Across all 7,534 citations, 59.8% pointed to domains ranked worse than #100,000 on the Tranco global traffic list. Another 23.4% reference domains outside the top million entirely. Of the vendor domains recommended by these cited sources, 1.1% were completely unreachable and 6.1% had redirected to unrelated sites — one to an Indonesian gambling portal, one to a Monaco hotel.
Traditional web search spent roughly fifteen years building defenses against content farms. Google’s Panda update in 2011 targeted low-quality content; Penguin followed in 2012 to address link manipulation. AI-powered search retrieval has been widely deployed for roughly two years. The content farm operators are running the same playbook, and the defenses are not yet built. Developers are effectively in the pre-Panda era of AI search — except the volume of fabricated pages at 215,128 took months to produce, not years. The Trellner dataset is published under CC BY 4.0, enabling independent verification of every claim.
Use AI Recommendations as a Starting Point, Not a Verdict
The practical response is not to abandon AI search tools — it is to stop treating their citations as the end of research. G2 and Reddit appear at the top of legitimate citations precisely because they carry community vetting that content farms cannot fake at scale: verified reviews, upvotes, real user accounts with histories. GitHub repository activity — stars, recent commits, open issues — is another credibility signal that AI tools do not currently surface but that takes thirty seconds to check manually.
For domain credibility, the Wayback Machine is a fast check. Unranked domains cited by AI tools have a median first Wayback capture of 2020; the content farm cluster was registered in 2023-2024. A domain that appeared two years ago and immediately started showing up in AI recommendations warrants independent verification. Before committing to any recommended vendor, confirm the domain resolves — and goes where it claims to go.
Key Takeaways
- Three coordinated sites manufactured 215,128 pages to exploit AI retrieval systems. Perplexity cited them for software recommendations across 380 categories.
- A vendor marketing blog outranked Gartner in AI citations — 194 to 158. AI retrieval systems weight retrieval signal, not source authority.
- 59.8% of Perplexity’s citations point to domains ranked below #100,000 globally. 6.1% of recommended vendor domains have redirected to unrelated sites.
- AI search is in its pre-Panda era. The defenses that took search engines fifteen years to build do not yet exist for AI retrieval.
- Cross-reference AI stack recommendations with G2, Reddit developer communities, and GitHub repository activity before committing to any vendor decision.













