Xinzhan Speed showcases the AI90 inference acceleration solution at WAIC 2026, reducing the first token latency by 50 times and increasing throughput by 5....
产品更新IT之家(RSS)
Today AI Intelligence Brief
Xinzhan Speed showcases the AI90 inference acceleration solution at WAIC 2026, reducing the first
token latency by 50 times and increasing throughput by 5.1 times by offloading the KV Cache to SSD to form a three-level storage. It also simultaneously releases the PT200Z AI SSD, based on pSLC flash memory, with a maximum DWPD of 100.
The readable text on this page was extracted from the public source and organized with attribution, publication time and the original link. Copyright remains with the original author and publisher.
Xinzhan Speed showcases the AI90 inference acceleration solution at WAIC 2026, reducing the first token latency by 50 times and increasing throughput by 5....
IT之家(RSS)2026-07-20T03:57:36.000Z
Scan to open this article
Aioga aggregates global AI updates and preserves source information for verification and citation.