Tencent Hunyuan released HyOCR-1.5, the first expert model in the end-to-end OCR large model field
to fully open source training, inference, and model weighting. With just 1B parameters, it covers more than 8 text-centric tasks. Introduced the DFlash speculative decoding framework, achieving 6.37× acceleration under Transformers and 2.14× acceleration under vLLM, with end-to-end inference reaching 1.408 seconds per page. Supports 4K resolution and 128K context windows, and expands low-resource OCR (331 languages), ancient script recognition, and multi-image Q&A capabilities through Agentic Data Flow. On OmniDocBench v1.6, it ranked first end-to-end with a score of 94.74.
腾讯混元发布 HyOCR-1.5,这是端到端 OCR 大模型领域首个将训练、推理、模型权重完整开源的专家模型。
仅 1B 参数,覆盖 8 种以上 text-centric 任务。
引入 DFlash 投机解码框架,在 Transformers 下实现 6.37× 加速,vLLM 下 2.14× 加速,端到端推理达每页 1.408s。
支持 4K 分辨率与 128K 上下文窗口,通过 Agentic Data Flow 扩展低资源 OCR(331 种语言)、古文字识别与多图问答能力。
在 OmniDocBench v1.6 上以 94.74 分居端到端第一。
The readable text on this page was extracted from the public source and organized with attribution, publication time and the original link. Copyright remains with the original author and publisher.