Vercel Labs 发布实验性语法高亮工具 gpu-lexer,用 27.4KB 的 WebGPU 模型对任意语言源码做 token 标注,无需按语言选择语法。
在 5.56M 字符输入上高亮耗时 402ms,快于 Prism.js 和 Shiki; 在留出文件上与 Shiki 的标注一致率为 88.02%,Top-25 加权一致率 90.35%,作者强调这是实验而非语法等价的高亮器。
Shu Ding:https://x.com/shuding 在 Vercel Labs:https://github.com/vercel-labs
gpu-lexer 将源代码拆分为简单的部分——单词、空白、换行符和符号。然后,一个微小的 WebGPU 模型结合局部和整个文件的上下文为每个部分打标签。它适用于任何语言:不选择语法,而是根据周围源代码猜测每个部分的类型,即使在训练期间从未见过该语言或语法。相邻的标签将成为返回给你的代码的语法片段。
这是一个实验,而不是等同于语法的高亮工具。在未进行训练的文件上,目前模型的 11.98% 的标记标签与 Shiki 不同。这衡量的是与 Shiki 的一致性——而非客观正确性——未见过的语言或真实世界的代码可能差异更大。
“语言无关”意味着使用一个共享的分词器和分类器,而不是每种语言都准确率相等。88.02% 是与 Shiki 相匹配的保留标记标签的比例。也支持混合语言的代码,包括 HTML、Vue 和 Svelte 中的嵌入和区域。
每种语言出现在一个区段;保留标签较少的结果稳定性较低。训练标记包括仅上下文标记和回放。
在 2026 年 9 月 9 日经过一次预热后的单次浏览器运行。输入为 10 个 concatenated 版本的 three.min.js:https://unpkg.com/three@0.97.0/build/three.min.js(5.56M 字符)。MacBook Pro,Apple M4 Pro,20 核 GPU,24GB,macOS 26.6.2,Chrome 152。每个引擎在独立的 worker 中运行;DOM 渲染被排除。gpu-lexer 和 Shiki 返回标记数据,Starry Night 返回 HAST 树,而 Sugar High、Prism.js 和 Highlight.js 返回高亮 HTML。Sugar High 2.3.1,Prism.js 1.30.0,Highlight.js 11.12.0,Starry Night 3.11.0,Shiki 4.4.3。
2026 年 9 月 9 日测量的最小化和 Brotli 压缩浏览器包。主要网页语言包括 javascript、typescript、css、html、json 和 markdown。gpu-lexer 对每种语言使用相同的包。Starry Night 总量包括其 Oniguruma WASM 负载。
Shiki 是 100% 标准化的参考。每个库的标记名称都映射到相同的九类:普通、注释、字符串、数字、关键字、类型、函数、常量和运算符。分数比较 GitHub Innovation Graph 上 1,103 个保留文件中的非空白源代码部分:https://innovationgraph.github.com/global-metrics/programming-languages 前 25 名的 2026-Q1,按每种语言的提交者数量加权。不支持的语言得分为零;语料库大小不影响权重。
实验性软件。高亮显示是概率性的,可能与 Shiki 不同,并且不是解析器,也不能替代编译器、代码检查器或安全分析工具。
Shu Ding:https://x.com/shuding 在 Vercel Labs:https://github.com/vercel-labs。
Shu Ding:https://x.com/shuding at Vercel Labs:https://github.com/vercel-labs
gpu-lexer splits source code into simple parts—words, whitespace, newlines, and symbols. Then a tiny WebGPU model combines local and whole-file context to label each part. It is designed for any language : instead of choosing a grammar, it guesses each part's type from the surrounding source, even when it never saw that language or syntax during training. Adjacent labels become the syntax spans returned to your code.
This is an experiment , not a grammar-equivalent highlighter. On files kept out of training, 11.98% of the current model's token labels differ from Shiki . This measures agreement with Shiki—not objective correctness—and unseen languages or real-world code may differ more often.
1 "Language-agnostic" means one shared tokenizer and classifier, not equal accuracy for every language. 88.02% is the share of held-out token labels that matched Shiki. Mixed-language code is supported too, including embedded and regions in HTML, Vue, and Svelte.
Each language appears in one band; results with fewer held-out labels are less stable. Training tokens include context-only tokens and replay.
One browser run after one warm-up on September 9, 2026. The input was 10 concatenated copies of three.min.js:https://unpkg.com/three@0.97.0/build/three.min.js (5.56M characters). MacBook Pro, Apple M4 Pro, 20-core GPU, 24GB, macOS 26.6.2, Chrome 152. Each engine ran in a dedicated worker; DOM rendering was excluded. gpu-lexer and Shiki returned token data, Starry Night returned a HAST tree, while Sugar High, Prism.js, and Highlight.js returned highlighted HTML. Sugar High 2.3.1, Prism.js 1.30.0, Highlight.js 11.12.0, Starry Night 3.11.0, and Shiki 4.4.3.
Minified and Brotli-compressed browser bundles measured on September 9, 2026. Major web includes javascript, typescript, css, html, json, and markdown. gpu-lexer uses the same bundle for every language. Starry Night totals include its Oniguruma WASM payload.
Shiki is the 100% normalization reference. Each library's token names are mapped to the same nine classes: plain, comment, string, number, keyword, type, function, constant, and operator. Scores compare non-whitespace source parts across 1,103 held-out files in the GitHub Innovation Graph:https://innovationgraph.github.com/global-metrics/programming-languages top 25 for 2026-Q1 , weighted by each language's pusher count. Unsupported languages score zero; corpus size does not affect the weights.
Experimental software. Highlighting is probabilistic, may differ from Shiki, and is not a parser or a substitute for compiler, linter, or security analysis.
Shu Ding:https://x.com/shuding at Vercel Labs:https://github.com/vercel-labs.