{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-07-28T06:20:51.496Z","headline":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers","description":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers。该来源目前仅提供简短信息，Aioga 已保留发布时间、来源和原文入口，并将继续跟踪后续更新。来源：MarkTechPost（RSS）。","url":"https://www.aioga.com/news/cmrx8yoqi008nrot3k4ncdpk2/","mainEntityOfPage":"https://www.aioga.com/news/cmrx8yoqi008nrot3k4ncdpk2/","datePublished":"2026-07-23T08:01:22.000Z","dateModified":"2026-07-23T08:01:22.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://www.marktechpost.com/2026/07/23/meet-gigatoken-a-rust-bpe-tokenizer-that-encodes-text-at-24-53-gb-s-up-to-989x-faster-than-huggingface-tokenizers","https://aihot.virxact.com/items/cmrx8yoqi008nrot3k4ncdpk2"],"canonicalUrl":"https://www.aioga.com/news/cmrx8yoqi008nrot3k4ncdpk2/","directAnswer":{"@type":"Answer","text":"Aioga 编辑摘要：Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than Aioga 将其归入「AI资讯」方向，重点关注它对真实使用和行业竞争的影响。","url":"https://www.aioga.com/news/cmrx8yoqi008nrot3k4ncdpk2/","dateCreated":"2026-07-23T08:01:22.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"marktechpost.com source article","url":"https://www.marktechpost.com/2026/07/23/meet-gigatoken-a-rust-bpe-tokenizer-that-encodes-text-at-24-53-gb-s-up-to-989x-faster-than-huggingface-tokenizers","datePublished":"2026-07-23T08:01:22.000Z","provider":{"@type":"Organization","name":"marktechpost.com","url":"https://www.marktechpost.com/2026/07/23/meet-gigatoken-a-rust-bpe-tokenizer-that-encodes-text-at-24-53-gb-s-up-to-989x-faster-than-huggingface-tokenizers"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmrx8yoqi008nrot3k4ncdpk2","datePublished":"2026-07-23T08:01:22.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmrx8yoqi008nrot3k4ncdpk2"}}],"aggregationSource":"MarkTechPost（RSS）","originalPublisher":{"name":"marktechpost.com","url":"https://www.marktechpost.com/2026/07/23/meet-gigatoken-a-rust-bpe-tokenizer-that-encodes-text-at-24-53-gb-s-up-to-989x-faster-than-huggingface-tokenizers"},"article":{"id":"cmrx8yoqi008nrot3k4ncdpk2","slug":"cmrx8yoqi008nrot3k4ncdpk2","url":"https://www.aioga.com/news/cmrx8yoqi008nrot3k4ncdpk2/","title":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers","title_en":"","summary":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers。该来源目前仅提供简短信息，Aioga 已保留发布时间、来源和原文入口，并将继续跟踪后续更新。来源：MarkTechPost（RSS）。","source":"MarkTechPost（RSS）","sourceUrl":"https://www.marktechpost.com/2026/07/23/meet-gigatoken-a-rust-bpe-tokenizer-that-encodes-text-at-24-53-gb-s-up-to-989x-faster-than-huggingface-tokenizers","aiHotUrl":"https://aihot.virxact.com/items/cmrx8yoqi008nrot3k4ncdpk2","publishedAt":"2026-07-23T08:01:22.000Z","category":"AI资讯","score":0,"selected":false,"articleBody":["Tokenization is the one part of the language modeling stack that almost nobody profiles. Gigatoken：https://github.com/marcelroed/gigatoken, released by Marcel Rød：https://x.com/marcelroed (a PhD student from Stanford) under an MIT license, argues that this was a mistake. The library encodes text at gigabytes per second on a single machine, against baselines that are already multithreaded Rust.","The GPT-2 tokenizer benchmarking yields remarkable results: evaluated on the 11.9 GB owt_train.txt corpus using a 144-core AMD EPYC 9565 dual-socket setup, Gigatoken processes data at a staggering 24.53 GB/s . In comparison, OpenAI’s tiktoken：https://github.com/openai/tiktoken achieves 36.0 MB/s, while HuggingFace tokenizers：https://github.com/huggingface/tokenizers registers at 24.8 MB/s on the identical hardware configuration. These marks demonstrate performance advantages of 681x and 989x, respectively.","On an Apple M4 Max with 16 cores, the same GPT-2 workload runs at 8.79 GB/s, or 1,268x HuggingFace tokenizers and 140x tiktoken. On a consumer AMD Ryzen 7 9800X3D, it runs at 6.27 GB/s, or 106x and 68x. The speedup is not an artifact of one CPU or one vocabulary.","Gigatoken is a byte-pair encoding (BPE) tokenizer written in Rust with Python bindings. It ships on PyPI as gigatoken (version 0.9.0, released 21 July 2026) and installs with pip install gigatoken. The repository is 66.2% Rust and 33.3% Python. It supports 23 distinct tokenizer families in the published benchmarks, covering GPT-2, GPT-OSS, Llama 3 through 4, Qwen 2 through 3.6, DeepSeek V3/R1/V4, GLM 4 and 5, Kimi K2, Nemotron 3, Phi-4, OLMo 2/3, ModernBERT, Gemma and Mistral.","There are two ways to use it. Compatibility mode wraps an existing HuggingFace or tiktoken tokenizer and preserves exact output parity, at a real cost to throughput. The author (Marcel：https://x.com/marcelroed) states on Hacker News：https://news.ycombinator.com/item?id=49010167 that compatibility mode delivers roughly 200–300x depending on usage, because it still pays Python overhead for list creation and string-to-bytes conversion. The native Gigatoken API lets Rust read files directly and is where the published numbers come from.","Every number below is drawn from the repository’s benchmarks section：https://github.com/marcelroed/gigatoken/#benchmarks and the pretokenizer optimization log：https://github.com/marcelroed/gigatoken/blob/main/pretokenizer_optimization_log.md. Switch CPUs, walk the optimization history, or estimate how long your own corpus would take.","The gains do not come from a better BPE merge loop. They come from two places that most tokenizers treat as solved.","(1) Pretokenization : Most implementations delegate this to a regex engine. Gigatoken hand-writes it. The optimization log tracks single-threaded GPT-2 pretokenizer throughput on 100 MB of OpenWebText, and the progression is instructive:","Net effect on the pretokenizer alone: 2.27x over the winnow + NEON baseline, and 22.3x over the regex implementation.","(2) Pretoken caching : If a word has been seen before, its encoded tokens are looked up rather than recomputed. The author (Marcel：https://x.com/marcelroed) notes this is hard in practice, because the cache grows quickly and pretoken distributions are long-tailed. On top of that, interactions with Python are minimized and threads are designed to interact minimally with each other.","The optimization log is unusually honest about what failed. A hot/cold split using #[cold] and #[inline(never)] regressed to 580 MiB/s and was reverted, because the inline barrier stopped LLVM from optimizing the combined ASCII and unicode loop. A two-pass classification buffer with SWAR transition counting was algorithmically correct but ran at 354 MiB/s, since the extra memory traffic beat the branch savings. Profile-guided optimization had no measurable effect, because the inner loop is already branchless and the word-boundary branch is data-dependent.","The comparison is not strictly apples-to-apples. Gigatoken encodes entire un-split files, finding its own document boundaries and parallelizing automatically. In contrast, HuggingFace (encode_batch_fast) is evaluated on the first 100 MB and tiktoken (encode_ordinary_batch) on the first 1 GB, both pre-split on . Because the baselines omit caching, their throughput remains uniform. Measurements report the best of three interleaved rounds using fresh processes with parallelism enabled.","Vocabulary types also introduce constraints; SentencePiece tokenizers are only partially optimized. On EPYC, Gemma 1 processes at 2.51 GB/s (7.3x speedup), Gemma 3 at 3.43 GB/s (9.6x), and CodeLlama at 3.47 GB/s (10.0x). Although substantial, these gains are an order of magnitude lower than the headline BPE performance.","An independent reproduction on KrabArena：https://krabarena.com/claims/gigatoken-ran-26-2x-faster-than-tiktoken-on-a-174-mb-gpt-2-owt-tokenizer-slice verified the results. On a 4-vCPU Intel Xeon VM (2.20 GHz) with a 174 MB OpenWebText slice, Gigatoken 0.9.0 achieved a median of 277.8 MB/s, outperforming tiktoken 0.13.0 (10.62 MB/s) by 26.2x and tokenizers 0.23.1 (3.33 MB/s) by 83.4x. All trials successfully validated 35,356 documents, confirming the performance trend scales with core count.","Check out the GitHub repo ：https://github.com/marcelroed/gigatoken and the release thread ：https://x.com/marcelroed/status/2079642154960564352 . All credit for this research goes to the author of this project.","Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences."],"articleImages":[{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2019/06/Screen-Shot-2021-09-14-at-9.02.24-AM-300x300.png","alt":"","afterParagraph":14,"url":"/media/articles/cmrx8yoqi008nrot3k4ncdpk2/787a6d54564e8e19.webp"},{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2026/07/blog6171-1-100x70.png","alt":"Designing High-Performance GPU Kernels with TileLang: Tensor-Core GEMM, Fused Softmax, FlashAttention, and Autotuning","afterParagraph":15,"url":"/media/articles/cmrx8yoqi008nrot3k4ncdpk2/80925b48ffa65ad2.webp"}],"mediaStatus":"ok","articleBodyZh":["分词是语言建模堆栈中几乎没有人进行性能分析的部分。Gigatoken：https://github.com/marcelroed/gigatoken，由Marcel Rød（https://x.com/marcelroed，斯坦福大学的博士生）在MIT许可证下发布，认为这是一个错误。该库能够在单机上以每秒数GB的速度对文本进行编码，而其基准已经是多线程的Rust实现。","GPT-2分词器的基准测试结果非常惊人：在使用144核AMD EPYC 9565双插槽配置评估11.9 GB的owt_train.txt语料库时，Gigatoken的数据处理速度高达24.53 GB/s。相比之下，OpenAI的tiktoken：https://github.com/openai/tiktoken达到36.0 MB/s，而HuggingFace分词器：https://github.com/huggingface/tokenizers在相同硬件配置下为24.8 MB/s。这些数据分别显示了681倍和989倍的性能优势。","在配备16核心的Apple M4 Max上，相同的GPT-2工作负载运行速度为8.79 GB/s，比HuggingFace分词器快1,268倍，比tiktoken快140倍。在消费级AMD Ryzen 7 9800X3D上，速度为6.27 GB/s，分别为106倍和68倍。这种加速并不是某个CPU或某个词汇表的偶然现象。","Gigatoken是一个用Rust编写、带有Python绑定的字节对编码（BPE）分词器。它通过PyPI以gigatoken（版本0.9.0，发布于2026年7月21日）形式发布，可用pip install gigatoken安装。该仓库由66.2%的Rust和33.3%的Python构成。已发布基准测试中支持23个不同的分词器家族，涵盖GPT-2、GPT-OSS、Llama 3至4、Qwen 2至3.6、DeepSeek V3/R1/V4、GLM 4和5、Kimi K2、Nemotron 3、Phi-4、OLMo 2/3、ModernBERT、Gemma和Mistral。","使用它有两种方式。兼容模式包装现有的HuggingFace或tiktoken分词器，并保持完全相同的输出，但会牺牲吞吐量。作者Marcel（https://x.com/marcelroed）在Hacker News（https://news.ycombinator.com/item?id=49010167）上表示，兼容模式根据使用情况大约提供200–300倍的速度提升，因为它仍然需要承担Python在列表创建及字符串到字节转换上的开销。原生Gigatoken API允许Rust直接读取文件，而发布的性能数据就是基于此模式得出的。","下面的每个数字都来自于存储库的基准部分：https://github.com/marcelroed/gigatoken/#benchmarks 以及预分词器优化日志：https://github.com/marcelroed/gigatoken/blob/main/pretokenizer_optimization_log.md。可以切换 CPU，查看优化历史，或估算您自己的语料库需要多长时间。","性能提升并不是来自更好的 BPE 合并循环，而是来自大多数分词器认为已经解决的两个方面。","(1) 预分词：大多数实现将其委托给正则表达式引擎。Gigatoken 是手写实现的。优化日志跟踪了在 100 MB OpenWebText 上单线程 GPT-2 预分词器的吞吐量，其进展具有启发性：","仅在预分词器上的净效果：相比 winnow + NEON 基线提升 2.27 倍，相比正则实现提升 22.3 倍。","(2) 预分词缓存：如果某个单词之前见过，其编码的 token 将被查找而不是重新计算。作者（Marcel：https://x.com/marcelroed）指出，这在实践中很困难，因为缓存会快速增长且预分词分布是长尾的。此外，与 Python 的交互被最小化，并且线程设计为彼此最少交互。","优化日志对失败的尝试异常诚实。使用 #[cold] 和 #[inline(never)] 热/冷分离导致性能下降到 580 MiB/s，因此被回退，因为 inline 屏障阻止了 LLVM 对 ASCII 与 Unicode 组合循环的优化。带 SWAR 过渡计数的两遍分类缓冲算法上是正确的，但运行速度为 354 MiB/s，因为额外的内存访问抵消了分支节省。基于 Profile 的优化没有可测量效果，因为内层循环已经没有分支，而单词边界分支是数据依赖的。","比较并非严格的同类比拼。Gigatoken 会对整个未拆分的文件进行编码，自己寻找文档边界并自动并行化。相比之下，HuggingFace（encode_batch_fast）是在前 100 MB 评估的，tiktoken（encode_ordinary_batch）是在前 1 GB 评估的，都是预先以 进行拆分。由于基线忽略了缓存，其吞吐量保持均匀。测量报告使用启用并行的新进程进行三轮交错的最佳结果。","词汇类型也引入了限制；SentencePiece 分词器仅部分优化。在 EPYC 上，Gemma 1 的处理速度为 2.51 GB/s（加速 7.3 倍），Gemma 3 为 3.43 GB/s（加速 9.6 倍），CodeLlama 为 3.47 GB/s（加速 10.0 倍）。尽管提升显著，但这些增益比起头条 BPE 性能低一个数量级。","在 KrabArena 上的独立复现：https://krabarena.com/claims/gigatoken-ran-26-2x-faster-than-tiktoken-on-a-174-mb-gpt-2-owt-tokenizer-slice 验证了这些结果。在一台配置 4 核 vCPU（2.20 GHz）的 Intel Xeon 虚拟机上，使用 174 MB 的 OpenWebText 切片，Gigatoken 0.9.0 达到了 277.8 MB/s 的中位速度，超过 tiktoken 0.13.0（10.62 MB/s）26.2 倍，超过 tokenizers 0.23.1（3.33 MB/s）83.4 倍。所有试验都成功验证了 35,356 份文档，确认性能趋势随着核心数量增加而提升。","查看 GitHub 仓库：https://github.com/marcelroed/gigatoken 以及发布帖：https://x.com/marcelroed/status/2079642154960564352。所有这项研究的功劳归属于该项目的作者。","Asif Razzaq 是 Marktechpost Media Inc. 的首席执行官。作为一位富有远见的企业家和工程师，Asif 致力于利用人工智能的潜力推动社会公益。他最近的努力是推出人工智能媒体平台 Marktechpost，该平台以对机器学习和深度学习新闻的深入报道而著称，既技术严谨又易于广大受众理解。该平台每月浏览量超过 200 万次，显示了其在受众中的受欢迎程度。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Aioga 编辑摘要：Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than Aioga 将其归入「AI资讯」方向，重点关注它对真实使用和行业竞争的影响。","background":"背景分析：AI 行业动态需要结合来源、时间、实际可用性和后续反馈判断，标题或单次发布本身不能替代完整证据。","viewpoint":"Aioga 判断：这条动态更适合作为行业观察信号，当前信息足以建立线索，但不足以推导长期结论。","implications":"影响分析：对相关团队而言，短期应先核对来源、可用范围和实际成本，再判断是否值得接入或跟进。","nextStep":"后续观察：继续观察原文更新、官方说明、用户反馈和同类产品的后续动作。","evidenceRefs":["title","summary","articleBody"],"confidence":"medium","status":"published","aiGenerated":false,"autoApproved":true,"generatedBy":"rule-safe-fallback","generatedAt":"2026-07-28T06:29:10.602Z","sourceHash":"8268292731ce3c73","validation":{"passed":true,"mode":"rule-safe-fallback","checks":["schema","length","source-attribution","no-html"]}},"tags":["AI资讯","MarkTechPost（RSS）"],"translations":{"zh-CN":{"title":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers","summary":"Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","category":"AI资讯","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers - Aioga AI资讯","description":"Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","url":"https://www.aioga.com/news/cmrx8yoqi008nrot3k4ncdpk2/"},"en":{"title":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers","summary":"Aioga tracks this update from MarkTechPost（RSS） under AI News. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","category":"AI News","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers - Aioga AI News","description":"Aioga tracks this update from MarkTechPost（RSS） under AI News. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","url":"https://www.aioga.com/en/news/cmrx8yoqi008nrot3k4ncdpk2/"},"ja":{"title":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers","summary":"Aiogaは「AIニュース」の動きとして、MarkTechPost（RSS） からの更新を追跡しています。Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","category":"AIニュース","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers - Aioga AIニュース","description":"Aiogaは「AIニュース」の動きとして、MarkTechPost（RSS） からの更新を追跡しています。Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","url":"https://www.aioga.com/ja/news/cmrx8yoqi008nrot3k4ncdpk2/"},"ko":{"title":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers","summary":"Aioga는 MarkTechPost（RSS）의 업데이트를 AI 뉴스 흐름으로 추적합니다. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","category":"AI 뉴스","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers - Aioga AI 뉴스","description":"Aioga는 MarkTechPost（RSS）의 업데이트를 AI 뉴스 흐름으로 추적합니다. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","url":"https://www.aioga.com/ko/news/cmrx8yoqi008nrot3k4ncdpk2/"},"es":{"title":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers","summary":"Aioga sigue esta actualización de MarkTechPost（RSS） dentro de Noticias IA. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","category":"Noticias IA","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers - Aioga Noticias de IA","description":"Aioga sigue esta actualización de MarkTechPost（RSS） dentro de Noticias IA. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokeniz...","url":"https://www.aioga.com/es/news/cmrx8yoqi008nrot3k4ncdpk2/"},"fr":{"title":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers","summary":"Aioga suit cette mise à jour de MarkTechPost（RSS） dans la catégorie Actu IA. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","category":"Actu IA","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers - Aioga Actualités IA","description":"Aioga suit cette mise à jour de MarkTechPost（RSS） dans la catégorie Actu IA. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace token...","url":"https://www.aioga.com/fr/news/cmrx8yoqi008nrot3k4ncdpk2/"},"de":{"title":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers","summary":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","category":"AI资讯","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers - Aioga KI-News","description":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","url":"https://www.aioga.com/de/news/cmrx8yoqi008nrot3k4ncdpk2/"},"pt-BR":{"title":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers","summary":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","category":"AI资讯","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers - Aioga Notícias de IA","description":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","url":"https://www.aioga.com/pt-BR/news/cmrx8yoqi008nrot3k4ncdpk2/"},"ru":{"title":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers","summary":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","category":"AI资讯","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers - Aioga Новости ИИ","description":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","url":"https://www.aioga.com/ru/news/cmrx8yoqi008nrot3k4ncdpk2/"},"ar":{"title":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers","summary":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","category":"AI资讯","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers - Aioga أخبار الذكاء الاصطناعي","description":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","url":"https://www.aioga.com/ar/news/cmrx8yoqi008nrot3k4ncdpk2/"},"hi":{"title":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers","summary":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","category":"AI资讯","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers - Aioga AI समाचार","description":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","url":"https://www.aioga.com/hi/news/cmrx8yoqi008nrot3k4ncdpk2/"},"it":{"title":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers","summary":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","category":"AI资讯","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers - Aioga Notizie IA","description":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","url":"https://www.aioga.com/it/news/cmrx8yoqi008nrot3k4ncdpk2/"},"nl":{"title":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers","summary":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","category":"AI资讯","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers - Aioga AI-nieuws","description":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","url":"https://www.aioga.com/nl/news/cmrx8yoqi008nrot3k4ncdpk2/"},"tr":{"title":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers","summary":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","category":"AI资讯","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers - Aioga AI Haberleri","description":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","url":"https://www.aioga.com/tr/news/cmrx8yoqi008nrot3k4ncdpk2/"},"vi":{"title":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers","summary":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","category":"AI资讯","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers - Tin tức AI Aioga","description":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","url":"https://www.aioga.com/vi/news/cmrx8yoqi008nrot3k4ncdpk2/"},"id":{"title":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers","summary":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","category":"AI资讯","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers - Berita AI Aioga","description":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","url":"https://www.aioga.com/id/news/cmrx8yoqi008nrot3k4ncdpk2/"},"th":{"title":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers","summary":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","category":"AI资讯","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers - ข่าว AI Aioga","description":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","url":"https://www.aioga.com/th/news/cmrx8yoqi008nrot3k4ncdpk2/"},"pl":{"title":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers","summary":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","category":"AI资讯","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meet Gigatoken： A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s， up to 989x Faster than HuggingFace Tokenizers - Aioga Wiadomości AI","description":"Aioga tracks this update from MarkTechPost（RSS） under AI资讯. Gigatoken is a Rust BPE tokenizer encoding text at 24.53 GB/s, up to 989x faster than HuggingFace tokenizers","url":"https://www.aioga.com/pl/news/cmrx8yoqi008nrot3k4ncdpk2/"}}}}