Trellner 于 2026 年 9 月 2 日向 perplexity/sonar 和 sonar-pro 询问 380 个软件类目的最佳产品,共 760 次调用,收集到 7,534 条引用,其中
TR-2026-009 · 2026年9月2日
在380个软件类别中,基于可靠 AI 推荐的数据源中有59.8% 位于前100,000个最受访问网站之外,并且一些被引用最多的网站是为模型读取而非人为阅读而建立的。
我们向两个基于网络的模型询问了380个软件类别中的最佳产品,并保留了它们获取的每个URL。在返回的7,534个引用中,59.8% 指向在Tranco前1M排名中低于第100,000位的域:https://tranco-list.eu/list/K9QPW,23.4% 指向根本未进入前百万名的域。两家用于数据支撑的网站将其主页的HTML标题命名为“Facts & Grounding Page”——其中“grounding”是模型执行的检索步骤——且它们与显然受共同控制的第三个网站一起,共发布了215,128个机器生成的最佳页面;这三个域在2023年12月之前都不存在。
2026年9月2日,我们将380个购买意图类别——从“CRM 软件”到“博物馆藏品管理软件”——通过OpenRouter提交给perplexity/sonar和perplexity/sonar-pro,每个模型每个类别一个提示,总共760次调用。每次调用都要求以JSON形式返回排名前五,并包含每个产品的官方主页域名。760次调用全部返回了可解析的答案,并且两个模型都报告了它们获取的URL,这也是选择它们的原因。类别在看到任何结果之前就已编写,并且从未修改过。
这样产生了3,800个推荐位,涉及1,807个不同产品,以及7,534个引用,跨2,055个不同域。然后我们在2026-09-01的Tranco每日列表和Wayback Machine中查找每个被引用的域,同时获取模型提供的所有1,502个供应商主页,以查看其是否仍然存在。
Google被排除在外。将Gemini模型在OpenRouter上进行grounding意味着通过OpenRouter自身的网络搜索插件进行路由,因此引用将描述该插件而非Google的检索。仅对Perplexity进行了测量,这里不应被解读为对任何其他搜索引擎的声明。
指向已排名域名的5,768条引用的Tranco中位排名为71,611。排名集中在前列并不显著——引用最多的十个域名占据了17.3%的引用——所以情况并不是由一群著名网站提供答案。关键在于剩下的四分之三:在2,055个被引用的域名中,有751个,占36.5%,没有出现在前一百万名之内。
这些域名也更为新近。未排名被引用域名的Wayback首次捕获中位年份是2020,而已排名域名为2011,并且存档的未排名域名中有16.6%首次捕获于2025年或之后,而存档的已排名域名中只有1.6%。
作为对比,维基百科在7,534条引用中只被引用三次。
guideflow.com 销售互动产品演示。它不是评论网站、目录或出版商,也不在我们所询问的任何类别中竞争。然而,它的博客在我们380个类别中的96个类别里被引用了194次——约四分之一——总体排名第三,超过Gartner。每条引用都是不同的URL:96个不同的guideflow.com博客URL,每个类别一个,其中六个是爱沙尼亚语版本的帖子。它的站点地图列出了3,351个博客URL,其中2,176个是不同的帖子。它为‘3D渲染软件’、‘IVR软件’、‘RFID软件’和‘建筑事务所软件’等提供了基础依据。
这里没有任何欺骗行为。Guideflow发布了大型内容营销博客,就像成千上万的公司一样。衡量标准是检索层如何处理它:一个厂商关于自己未经营市场的榜单内容,竟然成为了关于购买哪款产品的问题的第三大证据基础。
前十及其下方的另外三个网站分别是 wifitalents.com(71次引用,27个类别)、worldmetrics.org(60次,22个)和gitnux.org(50次,23个)。三者合计181次引用,占总数的2.4%,出现在380个类别中的41个类别里。
它们似乎是同一操作。三者均通过 NameCheap 注册,时间介于 2023 年 12 月至 2024 年 5 月,三者的 DNS 都委派给同一对 Cloudflare 名称服务器 pam.ns.cloudflare.com 和 sean.ns.cloudflare.com,并且三者都使用相同的页面模板和相同的导航——服务、市场数据、软件建议、编辑流程、公司。每个网站还都保留了恰好六篇博客文章,而且所有十八篇文章都涉及该系列的其他品牌:其他两者各两篇,以及第四个品牌 zipdo.co 两篇,后者位于同一名称服务器对上,并为其主页提供相同的“事实与基础页面”标题。共享名称服务器对是同一 Cloudflare 账户的强烈间接证据,但不是所有权证明,但模板、分类法和博客在各个项目上完全一致。
它们的规模是重点。它们的站点地图列出了 103,578、107,083 和 105,541 个 URL,其中 70,731、71,684 和 72,713 个是 /best/-software/ 页面:三大品牌共生成了 215,128 个购买指南,而每个站点只有六篇博客文章(每个站点地图中的第七个 /blog/ URL 是博客索引)。并不存在 215,128 个软件类别。
它们的自我描述使其不同寻常。2026 年 9 月 2 日获取时,worldmetrics.org 和 gitnux.org 均返回以下形式的 HTML 标题——事实与基础页面,并且除了品牌名称外,元描述完全相同:
关于 Gitnux 的经验证事实:一家独立市场研究公司,发布行业统计、定制研究和软件最佳榜单。公司、法律、方法论和合规性细节均在一份机器可读的记录中。
“基础”不是买家使用的术语。它是检索系统获取文档以条件化答案的步骤名称。关于自身的经验证事实的机器可读记录,也不是面向人类读者的服务。这些页面在标题和描述中,都是针对读取它们的软件。
这种读取也通过普通方式被购买。worldmetrics.org 宣传定制市场研究“从 €5,000 起”、现成报告“从 €499 起”以及供应商选择“从 €2,500 起”,其上方是同一分类法生成的最佳榜单,供模型检索。
我们从三个品牌获取了相同的分类页面:“项目估算软件”。每个页面都以 JSON-LD 显示其排名,因此无需解释即可读取。每个页面都列出了十个工具;显示前五名。
Gitnux 的冠军在 Worldmetrics 的五个中完全没有出现。每个页面都有三名指定工作人员——Worldmetrics 认可 Kathryn Blake、Alexander Schmidt 和 Victoria Marsh;Gitnux 认可 Diana Reeves、Helena Kowalczyk 和 Olivia Thornton;WifiTalents 认可 Ryan Gallagher、Isabella Rossi 和 Natasha Ivanova——一个问题有九个不同的人。每个页面都声明了编辑流程;Gitnux 将其结果标记为“AI 验证 · 专家审核”。三者在署名行都包含未呈现的模板变量,其中两个显示“在接下来的 26 天内”,另一个显示“在接下来的 40 天内”。
模型提供的 1,502 个供应商主页大多正常。我们每个都检查了两次,一次直接访问,一次通过旋转代理,如果任一次访问可达就算作可访问,这样阻止我们某个 IP 的主机不会被记录为无法访问的公司。
1,502 个中有十个完全无法解析地址——其中八个没有分配任何域名服务器——包括 graphiql.com(提供为 GraphiQL 的主页,但实际上没有那个网站)、todo.com(提供为 Microsoft To Do)和 aquasecurity.io(提供为 Trivy)。另外四个可以解析,但从未响应。加上 404 的情况,共有 17 个域名——占 1.1%——已消失或无法访问。另有 92 个,占 6.1%,会重定向到不同的可注册域名;大多数是普通收购和品牌更名,我们发布完整列表,而不是单独猜测每个。
两者都不是,并且在两组结果中存在分歧。要求提供研究数据管理平台时,两者都提到了 Dryad:sonar-pro 提供了真实的仓库 datadryad.org,sonar 提供 dryad.co,该网站重定向到印度尼西亚的在线博彩门户,标题以“BIGSLOT288 | Portal Game Online”开头。要求提供数据质量工具时,两者都提到了 Monte Carlo:sonar 提供 montecarlodata.com,sonar-pro 提供 montecarlo.com,该网站重定向至摩纳哥的 Monte-Carlo Société des Bains de Mer 酒店和赌场集团。
这两个模型不是两个独立的测量。在380个类别中,它们在289个类别中返回了字节相同的引用列表,且它们的URL集合的Jaccard重叠为0.898,因此Perplexity的层级共享一个检索层,应被视为同一个搜索堆栈抽样两次。它们在首选项上的一致性——在380个类别中有290个类别的首个产品相同——是关于该共享检索的事实,而不是独立系统趋同的证据。
该结果仅涵盖Perplexity。我们尚未测量ChatGPT、Gemini、Copilot或Google的AI模式,也没有理由假设它们的检索组合匹配。
这380个类别是我们自己构建的,而不是买家实际询问的样本,偏向利基垂直领域的列表会比常见查询列表呈现更多的长尾来源。
每次页面抓取都通过一个带轮换数据中心代理和具名研究用户代理的请求发出,因此这些网站返回给我们的内容不一定是它们返回给检索爬虫或浏览器的内容。
在十七个无法访问的供应商域名中,有四个是大型网站,包括nasdaq.com和solidworks.com,它们显然仍在运行,只是从未响应自动请求;这些域名被计为无法访问,而非已死。整个运行是一个检索索引变化的单日快照。
每个类别只运行一次提示措辞、一次运行,不重复抽样。早期试点显示,当“最佳”被替换为“最受欢迎”时,产品候选列表会明显变化,而引用组合变化较小,但本次运行未进行测量。
Tranco排名是流行度衡量,而非质量衡量,较低的排名并非指控。这里仅用于区分广泛访问的网页与其他网页;对特定网站的每一个声明都基于该网站自身的页面,这些页面在数据集中已被链接和归档。
我们尚未表明这些会改变答案。我们没有测试移除这些来源是否会产生不同的推荐,而Guideflow和三个最佳列表品牌很可能会列出合理产品。我们测量的是证据库由哪些文档构成。
最终,从共享的基础设施和相同的模板可以推断出这三个品牌由同一个实体控制。我们不知道是谁在运营它们;这三个品牌中没有一个提到所有者。
完整数据集——每个引用、每个推荐、Tranco 和 Wayback 查询、供应商活跃性检查——以及生成上面每个图表的脚本,均位于 /data/manufactured-sources-behind-ai-recommendations/:/data/manufactured-sources-behind-ai-recommendations/,方法说明位于:/data/manufactured-sources-behind-ai-recommendations/METHOD.md,列文档位于:/data/manufactured-sources-behind-ai-recommendations/README.md。以 CC BY 4.0 许可发布。本报告的 PDF 版本可在 trellner.com/data/manufactured-sources-behind-ai-recommendations/manufactured-sources-behind-ai-recommendations.pdf 获取:/data/manufactured-sources-behind-ai-recommendations/manufactured-sources-behind-ai-recommendations.pdf。
TR-2026-009 · 2 September 2026
Across 380 software categories, 59.8% of the sources behind grounded AI recommendations sit outside the 100,000 most-visited websites, and several of the most-cited are sites built to be read by models rather than by people.
We asked two web-grounded models for the best products in 380 software categories and kept every URL they retrieved. Of the 7,534 citations that came back, 59.8% point at domains ranked worse than #100,000 in the Tranco top-1M list:https://tranco-list.eu/list/K9QPW and 23.4% at domains that are not in the top million at all. Two of the sites doing the grounding have given their homepage the HTML title “Facts & Grounding Page” — grounding being the retrieval step these models perform — and they and a third site under apparently common control have published 215,128 machine-generated best pages between them; none of the three domains existed before December 2023.
On 2 September 2026 we put 380 buyer-intent categories — from “CRM software” to “museum collection management software” — to perplexity/sonar and perplexity/sonar-pro through OpenRouter, one prompt per category per model, 760 calls in all. Each call asked for a ranked top five as JSON, with each product’s official homepage domain. All 760 returned a parseable answer, and both models report the URLs they retrieved, which is why they were chosen. The categories were written before any results were seen and never revised.
That produced 3,800 recommendation slots naming 1,807 distinct products, and 7,534 citations spanning 2,055 distinct domains. We then looked up every cited domain in the Tranco daily list for 2026-09-01 and in the Wayback Machine, and fetched every one of the 1,502 vendor homepages the models supplied to see whether it still exists.
Google was left out. Grounding a Gemini model on OpenRouter means routing it through OpenRouter’s own web-search plugin, so the citations would describe that plugin rather than Google’s retrieval. Only Perplexity was measured, and nothing here should be read as a claim about any other engine.
The median Tranco rank of the 5,768 citations that point at a ranked domain is 71,611. Concentration at the top is unremarkable — the ten most-cited domains take 17.3% of citations — so the story is not that a cartel of famous sites supplies the answers. It is what fills the other four-fifths: 751 of the 2,055 cited domains, 36.5% of them, do not appear in the top million.
Those domains are also newer. The median first Wayback capture is 2020 for the unranked cited domains against 2011 for the ranked ones, and 16.6% of the archived unranked domains were first captured in 2025 or later, against 1.6% of the archived ranked ones.
Wikipedia, for comparison, was cited three times in 7,534.
guideflow.com sells interactive product demos. It is not a review site, a directory or a publisher, and it competes in none of the categories we asked about. Its blog was nonetheless cited 194 times across 96 of our 380 categories — a quarter of them — placing it third overall and ahead of Gartner. Each citation is a different URL: 96 distinct guideflow.com blog URLs, one per category, six of them the Estonian-locale copy of a post. Its sitemap lists 3,351 blog URLs, 2,176 of them distinct posts. It supplied the grounding for “3D rendering software”, “IVR software”, “RFID software” and “architecture practice software” alike.
Nothing here is deceptive. Guideflow publishes a large content-marketing blog, as thousands of companies do. The measurement is about what the retrieval layer does with it: a vendor’s own listicles about markets it does not operate in became the third-largest evidence base for a question about which product to buy.
Three other sites in the top ten and just below it are wifitalents.com (71 citations, 27 categories), worldmetrics.org (60, 22) and gitnux.org (50, 23). Together they account for 181 citations, 2.4% of the total, and appear in 41 of the 380 categories.
They appear to be one operation. All three were registered through NameCheap between December 2023 and May 2024, all three delegate DNS to the same pair of Cloudflare nameservers, pam.ns.cloudflare.com and sean.ns.cloudflare.com , and all three run the same page template with the same navigation — Services, Market Data, Software Advice, Editorial Process, Company. Each also keeps a blog of exactly six posts, and all eighteen are about the other brands in the set: two posts each on the other two, and two on a fourth brand, zipdo.co , which sits on the same nameserver pair and gives its own homepage the same “Facts & Grounding Page” title. Sharing a nameserver pair is strong circumstantial evidence of a common Cloudflare account rather than proof of ownership, but the template, the taxonomy and the blogs match item for item.
Their scale is the point. Their sitemaps list 103,578, 107,083 and 105,541 URLs, of which 70,731, 71,684 and 72,713 are /best/ -software/ pages: 215,128 generated buying guides across three brands , against six blog posts each (the seventh /blog/ URL in each sitemap is the blog index). There are not 215,128 software categories.
The self-description is what makes them unusual. Fetched on 2 September 2026, worldmetrics.org and gitnux.org both return an HTML title of the form — Facts & Grounding Page , and an identical meta description apart from the brand name:
Verified facts about Gitnux: an independent market research company publishing industry statistics, custom research, and software Best Lists. Company, legal, methodology, and compliance details in one machine-readable record.
Grounding is not a term buyers use. It is the name of the step in which a retrieval system fetches documents to condition an answer on. A machine-readable record of verified facts about oneself is not a service to a human reader either. These pages are addressed, in their titles and descriptions, to the software that reads them.
That reading is being purchased in the ordinary way as well. worldmetrics.org advertises custom market research “from €5,000”, ready-made reports “from €499” and vendor selection “from €2,500”, above the same taxonomy of generated Best Lists that the models retrieve.
We fetched the same category page from all three brands: “project estimation software”. Each page states its ranking in JSON-LD, so it can be read without interpretation. Each ranks ten tools; the top five are shown.
Gitnux’s winner does not appear in Worldmetrics’ five at all. Each page carries three named staff — Worldmetrics credits Kathryn Blake, Alexander Schmidt and Victoria Marsh; Gitnux credits Diana Reeves, Helena Kowalczyk and Olivia Thornton; WifiTalents credits Ryan Gallagher, Isabella Rossi and Natasha Ivanova — nine distinct people for one question. Each page announces an editorial process; Gitnux labels its result “AI-verified · Expert reviewed”. All three carry an unrendered template variable in the byline line, reading “Within the next 26 days” on two of them and “Within the next 40 days” on the third.
The 1,502 vendor homepages the models supplied are mostly fine. We checked each twice, once directly and once through a rotating proxy, counting a site as reachable if either attempt reached it, so that a host blocking one of our IPs is not recorded as a dead company.
Ten of the 1,502 resolve to no address at all — eight of them are not delegated to any nameserver — including graphiql.com (offered as the home of GraphiQL, which has no such site), todo.com (offered for Microsoft To Do) and aquasecurity.io (offered for Trivy). Four more resolve but never answer. With the 404s, 17 domains — 1.1% — are gone or unreachable. Another 92, 6.1%, redirect to a different registrable domain; most of those are ordinary acquisitions and rebrands, and we publish the full list rather than guess at each.
Two are not, and in both the two tiers disagreed. Asked for research data management platforms, both named Dryad: sonar-pro gave the real repository at datadryad.org , sonar gave dryad.co , which redirects to an Indonesian online-gambling portal whose title begins “BIGSLOT288 | Portal Game Online”. Asked for data quality tools, both named Monte Carlo: sonar gave montecarlodata.com , sonar-pro gave montecarlo.com , which redirects to Monte-Carlo Société des Bains de Mer, the Monaco hotel and casino group.
The two models are not two independent measurements. They returned a byte-identical citation list in 289 of the 380 categories and their URL sets overlap at a Jaccard of 0.898, so the Perplexity tiers share a retrieval layer and should be read as one search stack sampled twice. Their agreement on the top pick — the same product first in 290 of 380 categories — is a fact about that shared retrieval, not evidence that independent systems converge.
The result covers Perplexity only. We have not measured ChatGPT, Gemini, Copilot or Google’s AI Mode, and there is no reason to assume their retrieval mixes match.
The 380 categories are our own construction, not a sample of what buyers actually ask, and a list weighted towards niche verticals will surface more long-tail sources than a list of common queries would.
Every page fetch went out through a rotating datacentre proxy under a named research user-agent, so what these sites returned to us is not necessarily what they return to a retrieval crawler or to a browser.
Four of the seventeen unreachable vendor domains are large sites, nasdaq.com and solidworks.com among them, that are plainly alive and simply never answered an automated request; they are counted as unreachable, not as dead. The whole run is one day’s snapshot of a retrieval index that changes.
One prompt wording, one run per category, no repeat sampling. An earlier pilot suggested the product shortlist moves noticeably when “best” is swapped for “most popular” while the citation mix moves much less, but this run does not measure it.
Tranco rank is a popularity measure, not a quality measure, and a low rank is not an accusation. It is used here only to separate the widely-visited web from everything else; every claim about a specific site rests on that site’s own pages, which are linked and archived in the dataset.
We have not shown that any of this changes the answers. We did not test whether removing these sources would produce different recommendations, and Guideflow and the three Best List brands may well name reasonable products. What we measured is which documents the evidence base is made of.
Finally, common control of the three brands is inferred from shared infrastructure and an identical template. We do not know who operates them; none of the three names an owner.
The full dataset — every citation, every recommendation, the Tranco and Wayback lookups, the vendor liveness checks — and the scripts that produced every figure above are at /data/manufactured-sources-behind-ai-recommendations/:/data/manufactured-sources-behind-ai-recommendations/, with the method:/data/manufactured-sources-behind-ai-recommendations/METHOD.md and column documentation:/data/manufactured-sources-behind-ai-recommendations/README.md alongside. Released under CC BY 4.0. A PDF version of this report is available at trellner.com/data/manufactured-sources-behind-ai-recommendations/manufactured-sources-behind-ai-recommendations.pdf:/data/manufactured-sources-behind-ai-recommendations/manufactured-sources-behind-ai-recommendations.pdf.