Haus Research 对 Perplexity 的 sonar 与 sonar-pro 两个搜索模型提出 310 个关于 210 家科技公司的事实问题,抓取全部 cited URL 核验。
HR-2026-09 · 2026年9月2日
在1826个引用中,Perplexity的搜索模型附带了句子中的数字,其中34.7%指向无法打开或没有该句子任何图形的页面;按按主张而非按引用计分,872个主张中有14.4%失败。
我们向 Perplexity 的两个搜索模型询问了关于 210 家科技公司的 310 个事实问题,收集了它们引用的所有来源,检索所有来源,并核对页面是否说出了被引用的内容。在附带 1,826 个引用的句子中——那些无需二次意见即可核对的引用——34.7% 指向一个普通读者无法打开或打开后没有包含其句子数字的页面。模型总共放置了 2,511 个引用标记。
上面的单位是引用,而非主张。872个带有数字的主张中有三分之二带有多个标记,我们分别对每个标记进行评分。而是按主张计分,当任何一页显示其图形时,主张算为通过,14.4%未通过。我们先引用,因为标记是个独立的来源主张:这句话来自该网址。
失败的主要原因并非死链。只有1.3%的被引用URL是死链。两大类是读者无法进入的页面,以及读者能进入但未写明的页面。
十个问题模板,每个都是有人会查到的事实:成立、最新融资轮次、人员数、入职价格、总部、收入、披露的泄露情况、现任CEO、收购、付费运行时间SLA。每家公司都有一个;其中100个在不同模板上有第二题。310个问题,温度为0,分别对应perplexity/sonar、perplexity/sonar-pro,以及作为对照的GPT-4.1带网页插件。
两种困惑度模型都将内嵌声明标记为[n],n则索引它们返回的引用数组。这正是审计成为可能的关键:答案底部不是参考文献,而是明确断言该句子来自该URL。我们将每个回答拆分为句子,并为每个标记生成一对主张-引用对。两个模型都未发出指向自身引用列表末尾以外的标记。
然后我们抓取了每一个独特的被引用的 URL——仅 sonar 就有 2,915 个——并将每个 URL 分类为死链、受限、空白、无法访问或有效。任何失败的 URL 会再获得两次机会:一次是延长超时时间,再一次是通过轮换代理重试,这样就不会因为某个数据中心地址被拒而错误记录为被封锁。第三次尝试救回了 192 个 URL。分类只能朝页面有利的一方移动。
标题检查根本不需要任何模型。我们从每条声明中提取其具体细节——金额、百分比、规模、年份,以及任何连续三位或以上的数字——并询问被引用页面的可见文本是否包含其中至少一项,同时进行规范化处理,使得 $185 million、$185M 和 185000000 都匹配。只要有一个数字就算通过。单独的年份也算通过。因此 34.7% 是下限:每一个失败的配对都是页面中没有包含引用该句的任何数字。
在 perplexity/sonar 的 2,915 个独特引用 URL 中:
每六个引用中就有一个受限。这不是来源的问题——PitchBook、ZoomInfo、Crunchbase 和 Reuters 有权收费——但这是引用的问题。读者无法打开的脚注是一种无法检验的来源声明,而引用正是为了防止这种情况而存在。
汇总到答案层面,sonar 的 310 个答案中,有 84.2% 引用了至少一个普通读者无法打开的 URL,有 10.6% 引用了至少一个彻底失效的 URL。
失效的 URL 值得列出,因为它们中大约一半是同一类型的页面——sonar 的 38 个中有 20 个——而它们的 URL 已经暴露了这一点。komo.ai/directory/-offices、temperstack.com/plans/、devhelm.io/sla/、apollo.io/where-is/、portersfiveforce.com/blogs/brief-history/,以及 matrixbcg.com 和 canvasbusinessmodel.com 上的相同路径。这些页面是针对每家公司按问题类型生成的,按规模发布以捕捉我们提出的精确查询,并且下架的成本与上线时一样低。
在撰写当天,我们重新抓取了三个网址。问 Elastic 的总部在哪里,sonar 引用了 komo.ai/directory/elastic-offices:https://komo.ai/directory/elastic-offices: 404。问 Reddit 的总部在哪里,两种模型都引用了 apollo.io/where-is/reddit:https://www.apollo.io/where-is/reddit: 410 页面已移除。问 Discord 最便宜的付费计划,sonar-pro 引用了 temperstack.com/plans/discord:https://www.temperstack.com/plans/discord/: 404。
在那些能够打开并可读的页面对中,16.1% 没有包含声明本身的任何数值。
最干净的例子是价格。问 Vercel 最便宜的付费计划的入门价,sonar 回答“免费的 Hobby 计划是 $0/月,因此第一个付费等级从 $20/月 起”,并引用 vercel.com/docs/plans:https://vercel.com/docs/plans。我们在撰写时抓取了该页面。它返回 HTTP 200,页面列出了计划,但页面中并未出现 $20、$20/月 或 20/月。这个数字可能是对的,但引用并不能作为证据。
第二种模式更有启示性,因为它在不同公司间重复出现。问总部地址,两种模型都生成了一个街道地址并将其归因于公司的维基百科文章:
四篇文章都是在撰写时抓取的,没有一篇包含所归因的街道号、街道名称或邮政编码。几条回答自己也说明了这一点,用了类似“多个来源列出”或“多个商业目录列出”的措辞,但仍然附上了维基百科标记。声明和引用是通过同一过程生成的,而这个过程不是检索。
同样情况的温和版本:sonar 说 GitLab 的 CEO 是 Bill Staples,并且他在 2024 年 12 月 5 日上任,引用 GitLab 自己的执行团队页面:https://about.gitlab.com/company/team/e-group/。该页面确实提到了 Bill Staples,但没有提到日期。句子的一半有来源。
汇总两个 Perplexity 模型,按问题类型,引用页面包含声明数值的对的比例:
排序不是随机的。它追踪的是事实在一个规范位置上写得有多好。员工人数和创立年份放在专门记录这些信息的页面上的结构化字段中。CEO的入职日期和办公室门牌号是大家都会重复但没人公开的东西,所以模型会重现共识,然后指向一个从未包含共识的页面。
我们的飞行员建议,Sonar-Pro的主张基础性不如声纳。在全尺寸下,这一差距消失了。声纳在65.9%的数值对(95%区间62.8–68.9)上通过,Sonar-Pro有64.7%(61.5–67.7)。这两个间隔重叠得很舒服,两个模型的引用率几乎相同:每个答案有9.8和9.7个来源。在本次测量中,它们是相同的乘积。试点结果是一个小样本,告诉我们我们想听的,且没有重复。
带有网页插件的GPT-4.1在一个值得注意的方面表现不同:它每个回答引用2.0个来源,而非9.8个,且其中36.4%是公司自有域名,Perplexity为23.4%。它不发出内联标记,因此无法进行声明级检查,这本身就是结论——一个引用在末尾的列表答案无法逐句审计。
在Sonar的3,031次引用中,分布在989个不同的主机上,其中23.4%位于公司自有域名,23.1%位于B2B目录、收入估算或潜在客户列表——Tracxn、PitchBook、Clay、GetLatka、ZoomInfo、CB Insights、Growjo、Crunchbase及其众多模仿者。最大的单一主机 linkedin.com 为5.6%,其次是 en.wikipedia.org 为4.3%,然后是 tracxn.com 为3.0%。
这些目录页面也是该系列中最不持久的材料:66.0%的页面被打开,而整体引用量为78.7%。它们来自数据库,按排名大规模发布,未经通知被封锁或退休,且目前有四分之一的真实公司问题来源都集中在这些数据库。
我们还进行了传统的可靠性判断作为辅助指标:给每个模型展示每条声明及其引用的页面,然后询问该页面是否支持该声明。在每个模型随机抽取的400对可读配对中,它认为50.8%的Sonar声明得到支持,24.5%部分支持,24.8%不支持。结合所有能够打开页面的配对比例,这个端到端的支持率是40.0%。
这两个指标测量的内容不完全相同,差距正是一点。在可读页面上,确定性检查通过了84.6%的Sonar配对,因为只要一个匹配的图表就足够;而被询问页面是否支持整个声明的裁判则通过50.8%。端到端来看,这是65.9%对比40.0%,相差二十六个百分点。
无论如何,我们还是以确定性数字为主。一个模型评估另一个模型的工作是这个设计中最薄弱的环节,而差距就是真实存在的不确定性的大小。较宽松的数字是读者用grep可以复现的,而且已经糟糕到足以成为研究结果。
我们在Wayback Machine的CDX索引中查找了Sonar引用的随机1,500个URL。在解析成功的1,432个中,有25.1%(95%区间23.0–27.5)从未被捕获过——在页面生命周期的任何时候都没有一次。
对于目录页和主名单页,这一比例为39.3%。其中的两成存在只是因为托管公司持续维护,而我们已经知道结局:本研究中失效的URL几乎全部来自这一类。
因此,AI答案下的引用层主要是由为了被发现而不是为了长期保存而建立的页面组成,其中四分之一没有任何备份。当这些URL失效时,依赖它的声明并不会随之消失。句子仍然存在,标记仍然存在,唯一消失的是验证能力。
这是一个快照,拍摄于2026年9月2日。这不是衰减率,我们也没有在后来重新运行相同的URL。
它涵盖了关于科技公司的英文问题。这里没有任何内容能说明其他语言或行业会发生什么情况。
我们抓取 HTML 时不运行 JavaScript,这就是为什么空类存在并被单独报告,而不是折叠到无效计数中。
答案来自通过 OpenRouter 的 perplexity/sonar 和 perplexity/sonar-pro,而不是来自 Perplexity 消费者产品,后者会在其自己的设置下检索和引用资料。
我们以识别出的机器人身份从数据中心地址抓取,因此受限制的份额比使用浏览器和家庭连接的人能访问的份额要大。
这 210 家公司是手工挑选的知名科技公司,而不是从定义的总体中抽取的,因此主题的组合是我们的,而不是总体的。
而且一个声明即使引用不可靠也可以是真实的,其中大部分正是如此。我们并不是在衡量 Perplexity 是否正确。我们是在衡量它提供的证据是否能作为证据。在我们检查的三分之一引用中,它不能。
完整的数据集在 CC BY 4.0 许可下发布:问题:/data/perplexity-citation-audit/questions.csv,每个引用的 URL 及其解决方式:/data/perplexity-citation-audit/citations.csv,每个声明–引用对及其两个裁决:/data/perplexity-citation-audit/claim_citations.csv,以及 Wayback 查询:/data/perplexity-citation-audit/wayback.csv。收集和分析脚本:/data/perplexity-citation-audit/README.md,完整方法:/data/perplexity-citation-audit/METHOD.md,计算出的数据:/data/perplexity-citation-audit/numbers.json,以及每个命名示例的写时证据:/data/perplexity-citation-audit/evidence.json 都在同一目录下。
该报告的 PDF 版本可在 hausresearch.com/data/perplexity-citation-audit/perplexity-citation-audit.pdf 获取:/data/perplexity-citation-audit/perplexity-citation-audit.pdf。
HR-2026-09 · 2 September 2026
Of 1,826 citations Perplexity's search models attached to a sentence stating a figure, 34.7% pointed at a page that either would not open or did not contain a single figure from that sentence; scored per claim rather than per citation, 14.4% of 872 claims fail.
We asked Perplexity’s two search models 310 factual questions about 210 technology companies, collected every source they cited, fetched all of them, and checked whether the page said the thing it was cited for. Of the 1,826 citations attached to a sentence stating a figure — the ones checkable without a second opinion — 34.7% pointed at a page that would not open to an ordinary reader, or opened and contained none of the numbers in the sentence they were attached to. The models placed 2,511 citation markers in all.
The unit above is the citation, not the claim. Two thirds of the 872 claims carrying a figure have more than one marker on them, and we score each marker separately. Score instead per claim, counting a claim as passing when any one of the pages it points at carries one of its figures, and 14.4% fail. We lead with the citation because a marker is an individual claim of provenance: this sentence came from that URL.
The failure is not mainly dead links. Only 1.3% of cited URLs were dead. The two large categories are pages a reader cannot get into, and pages a reader can get into that do not say it.
Ten question templates, each a fact somebody would actually look up: founding, latest funding round, headcount, entry price, headquarters, revenue, disclosed breaches, current CEO, acquisitions, paid-tier uptime SLA. Every company got one; 100 of them got a second on a different template. 310 questions, put at temperature 0 to perplexity/sonar , perplexity/sonar-pro and, as a control, GPT-4.1 with a web plugin.
Both Perplexity models mark their claims inline as [n] , and n indexes the citation array they return. That is the part that makes an audit possible: it is not a bibliography at the bottom of the answer, it is a specific assertion that this sentence came from that URL. We split each answer into sentences and produced one claim–citation pair per marker. Neither model ever emitted a marker pointing past the end of its own citation list.
Then we fetched every unique cited URL — 2,915 of them for sonar alone — and classified each as dead, gated, empty, unreachable or live. Anything that failed got two more chances: a longer timeout, then a retry through a rotating proxy so that no page was recorded as blocked merely because one datacentre address was unwelcome. That third pass rescued 192 URLs. The classification can only ever move in a page’s favour.
The headline check needs no model at all. From each claim we pulled its specifics — money amounts, percentages, magnitudes, years, any run of three or more digits — and asked whether the cited page’s visible text contains at least one of them, normalising so that $185 million , $185M and 185000000 all match. One figure is enough to pass. A bare year is enough to pass. The 34.7% is therefore a floor: every failing pair is one where the page contains not a single number from the sentence that cited it.
Across perplexity/sonar ’s 2,915 unique cited URLs:
One citation in six is gated. That is not a fault of the source — PitchBook, ZoomInfo, Crunchbase and Reuters are entitled to charge — but it is a fault of the citation. A footnote a reader cannot open is a claim of provenance with no way to test it, which is the condition a citation exists to prevent.
Aggregated to the answer, 84.2% of sonar ’s 310 answers cited at least one URL an ordinary reader could not open, and 10.6% cited at least one that was outright dead.
The dead ones are worth naming, because about half of them are the same kind of page — 20 of sonar ’s 38 — and their URLs give them away. komo.ai/directory/ -offices . temperstack.com/plans/ . devhelm.io/sla/ . apollo.io/where-is/ . portersfiveforce.com/blogs/brief-history/ , and the identical path on matrixbcg.com and canvasbusinessmodel.com . These are pages minted per company per question type, published at scale to catch exactly the query we asked, and taken down as cheaply as they went up.
Three we re-fetched on the day of writing. Asked where Elastic is headquartered, sonar cited komo.ai/directory/elastic-offices:https://komo.ai/directory/elastic-offices: 404. Asked for Reddit’s head office, both models cited apollo.io/where-is/reddit:https://www.apollo.io/where-is/reddit: 410 Gone. Asked for Discord’s cheapest paid plan, sonar-pro cited temperstack.com/plans/discord:https://www.temperstack.com/plans/discord/: 404.
Of the pairs whose page did open and was readable, 16.1% contained none of the claim’s own figures.
The cleanest example is a price. Asked for the entry price of Vercel’s cheapest paid plan, sonar answered that “the free Hobby plan is $0/month, so the first paid tier starts at $20/month”, and cited vercel.com/docs/plans:https://vercel.com/docs/plans. We fetched that page at write time. It returns HTTP 200, it names the plans, and the strings $20 , $20/month and 20/month do not appear anywhere in it. The number is probably right. The citation is not evidence for it.
The second pattern is more revealing, because it repeats across companies. Asked for headquarters, both models produce a street address and attribute it to the company’s Wikipedia article:
All four articles were fetched at write time and none contains the street number, the street name or the postal code attributed to it. Several of the answers say so themselves, in phrasing like “multiple sources list” or “several business directories list”, and then attach a marker to Wikipedia anyway. The claim and the citation were produced by the same process, and that process is not retrieval.
A softer version of the same thing: sonar said GitLab’s CEO is Bill Staples and that he took the role on 5 December 2024, citing GitLab’s own executive team page:https://about.gitlab.com/company/team/e-group/. That page names Bill Staples. It does not carry the date. Half the sentence is sourced.
Pooling both Perplexity models, by question type, share of pairs whose cited page contained one of the claim’s figures:
The ordering is not random. It tracks how well a fact is written down in one canonical place. Headcount and founding year sit in structured fields on pages built to hold them. A CEO’s start date and an office’s street number are the kind of thing everyone repeats and nobody publishes, so the model reproduces the consensus and then points at a page that never carried it.
Our pilot suggested that sonar-pro grounded its claims less well than sonar . At full scale that gap disappears. sonar passes on 65.9% of numeric pairs (95% interval 62.8–68.9), sonar-pro on 64.7% (61.5–67.7). The intervals overlap comfortably, and the two models cite at nearly identical rates: 9.8 and 9.7 sources per answer. On this measurement they are the same product. The pilot result was a small sample telling us what we wanted to hear, and it did not replicate.
GPT-4.1 with a web plugin behaves differently in one respect worth noting: it cites 2.0 sources per answer rather than 9.8, and 36.4% of them are the company’s own domain against Perplexity’s 23.4%. It emits no inline markers, so no claim-level check is possible on it, which is itself the finding — an answer whose citations are a list at the end cannot be audited sentence by sentence.
Across sonar ’s 3,031 citations, spread over 989 distinct hosts, 23.4% point at the company’s own domain and 23.1% at a B2B directory, revenue estimator or lead list — Tracxn, PitchBook, Clay, GetLatka, ZoomInfo, CB Insights, Growjo, Crunchbase and their many imitators. The largest single host is linkedin.com at 5.6%, then en.wikipedia.org at 4.3%, then tracxn.com at 3.0%.
Those directory pages are also the least durable material in the set: 66.0% of them opened, against 78.7% of citations overall. They are generated from databases, published at scale to rank, gated or retired without notice, and they are where a quarter of the sourcing for questions about real companies now goes.
We also ran a conventional groundedness judgment as a secondary metric: a separate model shown each claim and its cited page, asked whether the page supports it. On 400 randomly sampled readable pairs per model it called 50.8% of sonar ’s claims supported, 24.5% partial and 24.8% unsupported. Chained with the share of pairs whose page opens at all, that is an end-to-end rate of 40.0%.
The two metrics are not measuring quite the same thing, and the distance between them is the point. On a readable page, the deterministic check clears 84.6% of sonar ’s pairs, because one matching figure is enough to pass it; the judge, which is asked whether the page supports the whole assertion, clears 50.8%. End to end that is 65.9% against 40.0%, a spread of twenty-six points.
We lead with the deterministic figure anyway. A model grading another model’s work is the weakest joint in this design, and the spread is the honest size of the uncertainty it introduces. The generous number is the one a reader can reproduce with grep , and it is already bad enough to be the finding.
We looked up a random 1,500 of sonar ’s cited URLs in the Wayback Machine’s CDX index. Of the 1,432 that resolved, 25.1% (95% interval 23.0–27.5) have never been captured at all — not once, at any point in the life of the page.
For the directory and lead-list pages the figure is 39.3%. Two in five of them exist only as long as the company hosting them keeps them up, and we already know how that ends: the dead URLs in this study are almost entirely from that same population.
So the citation layer under AI answers is being assembled largely out of pages built to be found rather than kept, and a quarter of it has no copy anywhere. When one of these URLs goes, the claim that leaned on it does not go with it. The sentence stays, the marker stays, and the only thing that has disappeared is the ability to check.
It is one snapshot, taken on 2 September 2026. It is not a decay rate, and we have not re-run the same URLs later.
It covers English-language questions about technology companies. Nothing here establishes what happens in other languages or sectors.
We fetch HTML without running JavaScript, which is why the empty class exists and is reported separately rather than folded into the dead count.
The answers came from perplexity/sonar and perplexity/sonar-pro through OpenRouter, not from the Perplexity consumer product, which retrieves and cites under its own settings.
We fetch as an identified bot from datacentre addresses, so the gated share is larger than the share a person with a browser and a home connection would meet.
The 210 companies were picked by hand as well-known technology firms rather than drawn from a defined universe, so the mix of subjects is ours and not a population.
And a claim can be true with a bad citation, which most of these are. We are not measuring whether Perplexity is right. We are measuring whether the thing it offers as proof functions as proof. On a third of the citations we checked, it does not.
The full dataset is published under CC BY 4.0: questions:/data/perplexity-citation-audit/questions.csv, every cited URL and how it resolved:/data/perplexity-citation-audit/citations.csv, every claim–citation pair with both verdicts:/data/perplexity-citation-audit/claim_citations.csv, and the Wayback lookups:/data/perplexity-citation-audit/wayback.csv. The collection and analysis scripts:/data/perplexity-citation-audit/README.md, the full method:/data/perplexity-citation-audit/METHOD.md, the computed figures:/data/perplexity-citation-audit/numbers.json and the write-time evidence:/data/perplexity-citation-audit/evidence.json for every named example are in the same directory.
A PDF version of this report is available at hausresearch.com/data/perplexity-citation-audit/perplexity-citation-audit.pdf:/data/perplexity-citation-audit/perplexity-citation-audit.pdf.