Hacker News

Show HN: External KV Cache Offloading Cuts Long Horizon Inference Costs by 50%

Hacker News - Sun, 07/26/2026 - 9:02am

Hey HN, we’re the developers of OpenLake, an open source storage engine for offloading LLM KV caches from GPU memory into a shared tier of RAM and NVMe. We built OpenLake because KV caches are outgrowing GPU memory.

A single 256K token conversation on Gemma 4 31B produces approximately 43GB of KV state, more than half the memory of an 80GB H100. The problem becomes even harder across a cluster: a prefix cached on one GPU host is unavailable when the next request lands on a different GPU, forcing the new GPU to repeat work the fleet has already completed.

Once the KV cache is offloaded, network bandwidth becomes a major constraint on read latency. To move less data across the wire, we built deferred materialization: a custom CUDA kernel that losslessly compresses KV blocks before they leave GPU memory and decompresses them on the GPU after retrieval. In our tests, this achieved:

- 1.72× lossless KV compression. - Approximately 600GB/s decompression throughput on an H100 - 80GB/s of effective KV throughput over a physical 50GB/s link

At 128K context, retrieving cached KV reduces TTFT from 44 seconds to 0.6 seconds, a 66× improvement. Across the complete workload, GPU time reduces from 1,169 seconds to 606 seconds, saving 48.2% of GPU cost.

OpenLake is written in Rust and uses io_uring with one pinned runtime per physical core. We provide connectors for vLLM and SGLang so the cache can be enabled without modifying the inference engine itself.

I would love to hear how others are handling KV reuse across GPU hosts, especially for long contexts, and get to know your thoughts.

Thanks!

GitHub: https://github.com/openlake-project/openlake

Here is our blog: https://cloud.theopenlake.com/blog/taming-the-beast-managing...

Comments URL: https://news.ycombinator.com/item?id=49057767

Points: 10

# Comments: 0

Categories: Hacker News

Show HN: Temporal Context Map

Hacker News - Sun, 07/26/2026 - 8:59am

Timeline with depth.

Comments URL: https://news.ycombinator.com/item?id=49057749

Points: 2

# Comments: 0

Categories: Hacker News

Ask HN: Are you using Rust on embedded devices yet? If not, why?

Hacker News - Sun, 07/26/2026 - 8:11am

I started dabbling with ESP32-based MCUs from Waveshare with Rust, and I'm quite impressed with the state of things (esp-rs, embassy, probe-rs etc). Note that I haven't really done embedded in any other languages, so I don't really have anything to compare it with.

One thing I did notice was that creating beautiful, interactive user interfaces on MCUs with displays is a little harder, libs like embedded-graphics don't look that great.

This made me wonder, are people using Rust in their professional embedded work?

Comments URL: https://news.ycombinator.com/item?id=49057337

Points: 2

# Comments: 0

Categories: Hacker News

Pages