欧洲 AI 实验室 Desert Ant Labs 正式成立并上线首批 18 个端侧专用模型(12 个稳定版、6 个 beta),覆盖音频、视觉和文本任务,通过 Swift、Kotlin 和
今天我们推出了 Desert Ant Labs,一家欧洲前沿的人工智能实验室,致力于打造有明确方向的设备端智能。我们认为实现高效智能的最佳路径从设备端开始。
我们正在为音频、视觉和文本构建小型、专用模型——每个模型在毫秒级别内作出响应,运行成本几乎为零,因此您可以在每次产品交互中嵌入智能,而不受代币成本或推理速度的限制。模型小到可以在五年前的手机上运行,快到可以在每一帧或每次按键时使用,并且比您已经支付的 API 调用更优。
首批18个模型今天上线(12个稳定版,6个测试版),可通过一个支持 Swift、Kotlin 和 JavaScript 的 SDK 访问。每个任务使用一个模型,每个模型都被打造成在设备上完成该任务的最快方式:
这里只是举几个例子。其余十四个模型的完整规格和基准测试可在 desertant.com/models:https://desertant.com/models/ 以及 Hugging Face:https://huggingface.co/desert-ant-labs 查阅。每个模型每月最多 10 万活跃设备免费使用。无需代币,无需登录。
我们在欧洲进行开发,在这里“设备端”是数据主权的默认选择。数据永远不会离开用户手中,功能永远不依赖于他人的云端,而从未上传的数据永远不会被强制获取。
五年来,我们一直在以设备优先的方法开发视频应用 Detail:https://detail.co。但当我们引入如 Auto Edit 生成短视频剪辑、或播客音频增强等功能时,不得不依赖云端 API。当 Detail:https://www.apple.com/newsroom/2025/12/apple-unveils-the-winners-of-the-2025-app-store-awards/ 越来越受欢迎时,我们的基础设施开销也随之增加。
每隔几个月,我都会寻找有用的设备端模型。我会在 Hugging Face:https://huggingface.co 寻找能识别无意义词或清理录音的模型。每年六月,我们都会得到新的出色工具,但行业的进展速度还不够快。基础已经具备:芯片、Core ML:https://developer.apple.com/documentation/coreml/、研究成果都在。但缺失的是,从这些基础到在应用中实际实现功能之间所需的一切:一个可以直接嵌入、只需几行代码就能部署的模型。
所以,我们自己训练了模型。事实证明,训练模型是一项产品设计挑战,而产品正是我们所擅长的。我们设计了模型和本地推理,在速度、质量和成本上超过云服务,并且在任务本身上也超过其他本地和云模型,其体积仅为它们的一小部分。
我们用 Clear:https://desertant.com/models/clear/ 替换了 Dolby,实现了更好、更快的音频增强,并用 Voz:https://desertant.com/models/voz/ 让我们的设备端转录速度提升了 5 倍。我们还用 Clips:https://desertant.com/models/clips/ 替换了 Claude Sonnet,这是我们的 284MB 模型,可在 5 秒钟内将 10 分钟的视频分成十几段剪辑——比 Sonnet 快 10 倍,能耗减少 470 倍,并保持相同的质量:https://desertant.com/about/。
将在 iOS 27 发布的 Detail 6,用我们自己的模型取代了所有云 API,全部运行在设备上。
过去几年,我们都像使用普通 API 一样使用大语言模型(LLM)进行开发。而在泛用前沿智能的大量宣传中,我们几乎忘记了它们并不是唯一选择。
我接触的每个开发者都有一个心愿清单,如果成本不是问题,他们希望在设备上构建的模型,或者他们愿意用本地模型替换掉因功能而消耗大量令牌的任务。每天重复十万次的操作:清理录音、标记照片、从句子中提取日期、在文本到达服务器前捕捉名字。上述任务都不需要前沿模型。
NVIDIA 自家的研究人员:https://arxiv.org/abs/2506.02153 拆解了三个智能代理系统,并估算有 40% 到 70% 调用大型模型的请求可以交给小型、专用模型处理。
今年,业内将在数据中心上花费约 4500 亿美元:https://introl.com/blog/hyperscaler-capex-600b-2026-ai-infrastructure-debt-january-2026。同时,全球出货超过十亿部的:https://my.idc.com/getdoc.jsp?containerId=prUS53965725 手机、平板和笔记本电脑上配备了越来越强大的芯片,非常适合完成这类任务。人们手中的计算能力,比地球上所有 AI 数据中心的计算能力加起来还要多。
我们在免费推理上有不公平的优势。没有每次调用成本,因此某个功能会在每条消息上运行,而不仅仅是你负担得起检查的那几条。没有往返,你的客户数据永远不会离开设备。当推理免费时,我们构建产品的方式会完全改变。
要使用本地模型进行构建,开发者体验必须大幅提升。你需要能够商业使用的模型,在速度和质量上超过替代方案,可以用几行代码就放入应用中,并且易于发现。
将最初的一百个模型想象成小脑:https://www.ncbi.nlm.nih.gov/books/NBK538167/,小脑负责持续运作的工作——平衡、时机、那些你从未思考的技能,让大脑的其他部分可以自由思考。这就是我们首先构建的:快速、专用的模型,用于全天运行的工作,直接在设备上免费运行。
然后是大脑皮层,这一层决定哪个模型来回答。首先使用一个小型本地模型,需要时使用更大的模型,仅当工作必须离开设备时才使用云。随着开放研究进展和设备硅能力增强,本地模型会逐步增长,我们也会训练更大的模型。前沿智能,从小端开始构建。
云端实验室提供中性模型,因为按令牌计费需要中性模型。每个 Desert Ant 模型自带我们选择的默认值,以及你需要更改默认值的调节手段。我们将模型和运行时一起优化:在 iPhone 上,Clear 和 Voz 在神经引擎上运行,在浏览器中,Clear 的相同权重通过 WebAssembly 运行。
准备好开始了吗?你可以通过我们的原生 Swift:https://desertant.com/swift/、Kotlin:https://desertant.com/kotlin/ 和 JavaScript:https://desertant.com/js/ SDK 在你的应用中实现 Desert Ant 模型,SDK 可在 GitHub 获取:https://github.com/Desert-Ant-Labs。
我们的文档:https://desertant.com/docs/ 是为开发者和代理编写的,你可以在 Mac 上使用 CLI:https://github.com/Desert-Ant-Labs/desert-ant-cli 尝试模型,或者在 Hugging Face 浏览器中体验:https://huggingface.co/desert-ant-labs。
正在使用我们的模型构建酷炫的东西,或者想和我们一起构建吗?请联系我们:https://desertant.com/contact/。
Today we're launching Desert Ant Labs, a European frontier AI lab building opinionated on-device intelligence. We believe the best path to efficient intelligence starts on-device.
We're building small, specialized models for audio, vision, and text – each model answers in milliseconds, and costs nothing to run, so you can put intelligence in every product interaction, without being limited by token cost or inference speed. Small enough to run on a five-year-old phone, fast enough to use on every frame or keystroke, and better than the API call you're already paying for.
The first 18 models are live today (12 stable and six in beta), accessible via one SDK for Swift, Kotlin, and JavaScript. One model per task, each built to be the fastest way to complete that task on a device:
And that's just to name a few. You can find full specs and benchmarks for the other fourteen, on desertant.com/models:https://desertant.com/models/ and Hugging Face:https://huggingface.co/desert-ant-labs. Every model is free up to 100k monthly active devices. No tokens, no logins.
We're building this in Europe, where "on-device" is the sovereign default. The data never leaves your customer's hands, the feature never depends on someone else's cloud, and what's never been uploaded can never be compelled.
For five years we've been building our video app, Detail:https://detail.co, with an on-device first approach. But when we introduced features like Auto Edit to create short clips, or audio enhancement for podcasts, we had to fall back to cloud APIs. And as the popularity of Detail:https://www.apple.com/newsroom/2025/12/apple-unveils-the-winners-of-the-2025-app-store-awards/ grew, so did our infrastructure bills.
Every few months I'd hunt for useful on-device models. I'd surf Hugging Face:https://huggingface.co for a model that could find filler words or clean up a recording. And, every June, we'd get great new tools to build with but the industry wasn't moving fast enough. The foundation was there: the chips, Core ML:https://developer.apple.com/documentation/coreml/, the research. What was missing was everything between that foundation and actually implementing a feature in your app: a model you could drop in and ship with a few lines of code.
So, we trained the models ourselves. It turns out training a model is a product design challenge, and product is what we know. We designed models and local inference that beat cloud services on speed, quality, and cost, and outperform other local and cloud models on the task itself, at a fraction of their size.
We replaced Dolby for better, faster audio enhancement with Clear:https://desertant.com/models/clear/, and made our on-device transcriptions 5x faster with Voz:https://desertant.com/models/voz/. We also replaced Claude Sonnet with Clips:https://desertant.com/models/clips/, our 284MB model that turns a 10-minute video into a dozen clips in 5 seconds – 10x faster and using 470x less energy:https://desertant.com/about/ than Sonnet, with the same quality.
Detail 6, which will launch with iOS 27, replaces all of our cloud APIs with our own models, running entirely on the device.
We've all spent the past few years building with LLMs as if they were just another API. And, amid the hype around generalist frontier brains, we almost forgot they're not the only option.
Every developer I talk to has a wishlist of on-device models they'd build if cost wasn't a factor, or a feature they're bleeding tokens on that they'd happily swap for a local model. A call that runs the same way a hundred thousand times a day: cleaning a recording, tagging a photo, pulling a date out of a sentence, catching a name before the text hits your servers. None of these needs a frontier model.
NVIDIA's own researchers:https://arxiv.org/abs/2506.02153 pulled apart three agent systems and estimated that 40 to 70% of their calls to a large model could go to a small, specialized one instead.
The industry will spend about $450 billion:https://introl.com/blog/hyperscaler-capex-600b-2026-ai-infrastructure-debt-january-2026 on data centers this year. Meanwhile, the world ships more than a billion:https://my.idc.com/getdoc.jsp?containerId=prUS53965725 phones, tablets, and laptops with increasingly capable chips, perfectly suited to these kinds of tasks. There's more compute available in people's hands than in every AI data center on earth.
We have an unfair advantage with free inference. No per-call cost, so a feature runs on every message instead of the ones you can afford to check. No round-trip, and your customer's data never leaves the device. When inference costs nothing, the way we build products changes entirely.
To build with local models, the developer experience has to get a lot better. You need models you can use commercially, that beat the alternatives on your task in speed and quality, that you can drop into your app with a few lines of code, and are easy to discover.
Think of the first hundred models as the cerebellum:https://www.ncbi.nlm.nih.gov/books/NBK538167/, the little brain. The little brain handles the always-on work – balance, timing, the skills you never think about, so the rest of the brain is free to think. That's what we're building first: fast, specialized models for the work that runs all day, on the device, for free.
Then comes the cortex, the layer that decides which model answers. A small local model first, a bigger one when the job requires it, and the cloud only when the work has to leave the device. As open research advances and device silicon becomes more capable, the local models grow, and we'll train larger ones ourselves. Frontier intelligence, built from the small end up.
Cloud labs ship neutral models because per-token pricing needs a neutral model. Every Desert Ant model ships with a default we choose, and the levers you need to change that default. We optimize the model and the runtime together: on an iPhone, Clear and Voz run on the Neural Engine, and in the browser, Clear's same weights run through WebAssembly.
Ready to get started? You can implement Desert Ant models in your app with our native Swift:https://desertant.com/swift/, Kotlin:https://desertant.com/kotlin/, and JavaScript:https://desertant.com/js/ SDK, available on GitHub:https://github.com/Desert-Ant-Labs.
Our docs:https://desertant.com/docs/ are written for developers and agents and you can try the models on your Mac with the CLI:https://github.com/Desert-Ant-Labs/desert-ant-cli, or in your browser on Hugging Face:https://huggingface.co/desert-ant-labs.
Building something cool with our models, or want to build them with us? Get in touch:https://desertant.com/contact/.