TuringData Launches ContextCube at Tech Week Singapore 2026, Bringing Shared KV Cache to GPU Inference Clusters

4 hours ago 14
LIKE WEBLYF.COM ON FACEBOOK

SINGAPORE, Sept. 29, 2026 /PRNewswire/ -- TuringData today launched ContextCube, a purpose-built KV cache appliance that gives AI inference clusters a shared, persistent pool of context.

Unveiled at Tech Week Singapore 2026, ContextCube enables participating inference nodes to retrieve matching cached KV over RDMA instead of repeatedly computing or storing the same context. This reduces time to first token, frees GPU cycles for token generation and supports greater throughput and concurrency.

As prompts grow longer and requests move between servers, valuable KV cache is often evicted or becomes inaccessible because it is held locally. ContextCube addresses this challenge by making previously computed context reusable across the cluster.

Designed on the NVIDIA CMX inference-storage reference architecture, ContextCube adds a shared Layer 3.5 alongside GPU HBM, DRAM and local NVMe storage. Active computation remains on GPUs while reusable context is retained in a large, cluster-wide cache.

 TuringData)
TuringData launched ContextCube, a KV cache appliance that gives AI inference clusters a shared, persistent pool of context, at Tech Week Singapore 2026. (Photo: TuringData)

Key capabilities include:

Hundreds of terabytes of shared KV cache, with retention ranging from hours to weeks depending on configuration, workload and cache policy 120 GB/s KV bandwidth per appliance* DPU-based data movement using four NVIDIA BlueField-3 DPUs 26 U.2 NVMe drives and two 200 Gb/s ports per DPU Compatibility with inference engines including vLLM and SGLang, and KV cache software including TuringData Cache Fabric, LMCache and Mooncake Starting deployments from one appliance

ContextCube is designed for context-intensive workloads, including AI agents, coding assistants, multi-turn conversations, long-document processing, retrieval-augmented generation and large-scale "token factories."

"As context windows grow and AI agents move into production, the same context is being computed again and again on every server," said Jenvik Li, Founder and CEO of TuringData. "ContextCube gives the whole cluster one place to keep that work, so GPUs can spend their time generating new tokens instead of rebuilding old ones."

ContextCube is available to order now. Visit www.turingdata.io/contact-us.

About TuringData

Headquartered in Singapore, TuringData builds the data layer for AI infrastructure, helping enterprises and AI cloud providers improve GPU utilisation and token economics across training and inference. Learn more at www.turingdata.io.

*Actual performance depends on configuration and workload. NVIDIA and BlueField are trademarks and/or registered trademarks of NVIDIA Corporation in the United States and other countries.

Source