This RAM/Memory reduction is most useful for Valkey 9.1 environments that are key string heavy. They are used extensively in caching, sessions/authentication, real-time apps, ecomm, Ad tech/martech, and gaming.
These scenarios see a reduction of 37.5%, which has been attained in our testing environments versus version 7.2.. The biggest jump in the reduction came between 7.2 and 8.1 where the difference was 29.8%. Although the reduction doesn’t apply to every Valkey environment, those that are string heavy should experience 37.5% RAM reductions.
RAM usage scenario:
Starting point: 1 TB on Valkey 7.2
This post is based on a rigorous benchmark and a structural teardown of the source, aimed at answering three questions:
Five versions were tested — 7.2.14, 8.0.10, 8.1.9, 9.0.5, and 9.1.1 — all built from source, all using jemalloc 5.3.0 as the allocator, so version is the only variable.
Test Environment settings:
| Version | Unique Keys | used_memory | Bytes / Key | vs Prev | vs 7.2.14 |
|---|---|---|---|---|---|
| 7.2.14 | 632,181 | 61.78 MB | 102.48 B | — | baseline |
| 8.0.10 | 632,565 | 57.00 MB | 94.49 B | −7.8% | −7.8% |
| 8.1.9 | 632,573 | 43.43 MB | 71.98 B | −23.8% | −29.8% |
| 9.0.5 | 631,672 | 43.36 MB | 71.99 B | +0.0% | −29.8% |
| 9.1.1 | 631,985 | 38.58 MB | 64.02 B | −11.1% | −37.5% |
We can see that across roughly 632K unique keys, the size of per key averages 102.48 bytes in version 7.2.14 down to 64.02 bytes in 9.1.1 — a cumulative drop of 37.5%. In other words, the same hardware stores roughly 60% more of these small keys on 9.1 than it did on 7.2. We can see 8.1 alone accounts for the single biggest jump, at -23.8%, because that’s where the hash table got rewritten. However 9.0 shows basically zero change, since that release simply wasn’t targeting memory.

In 7.2, dictEntry is a classic three-pointer struct — *key, *val, *next — 24 bytes total. The catch is that the key itself is a pointer to a separately allocated SDS blob: every lookup pays for an extra malloc and an extra memory hop just to read the key.

8.0 introduces embeddedDictEntry: the key’s bytes are stored inline right after the struct. *val, *next, plus the inlined key data come to 16 bytes plus the key length — one allocation instead of two.
Result: −8 bytes per key, one fewer memory hop on every lookup, automatic on upgrade, no config changes needed.
This is the single largest jump in the whole series. In 8.0, the dictionary is a classic chained hash table: reaching a value means “bucket → entry → key → value” — four memory hops (plus two more per hash collision).

8.1 rewrites this as a new hashtable structure where one bucket is exactly 64 bytes — the size of a CPU cache line, — with the key and value embedded directly inside it.The bucket’s 8-byte metadata region packs a “has-child-bucket” bit, 7 filled-bits, and 7 bytes of secondary hash, letting most non-matching candidates be skipped without an extra memory access. Key and value are embedded in the same serverObject, shortening the path to roughly two hops — “bucket → serverObject” — and a hash collision usually stays within the same cache line.
dictEntry as a struct disappears entirely. The result is roughly −20 bytes per key-value pair (−30 bytes for keys with a TTL) — the single biggest jump in the entire 7.2→9.1 series.



In earlier versions, embedded strings (embstr) still wasted an 8-byte pointer field simply referencing the contiguous payload. 9.1’s fix: for short strings, the string bytes are stored directly inline in those same 8 bytes normally used for the pointer — no pointer needed, and no separate allocation for the value.
Result: up to a 20% memory reduction for strings under 128 bytes — and the benchmark’s “key + 16-byte value” pairs sit right in the middle of that range.
The three structural changes account for most of the improvement:
8.0 inlined the key directly into dictEntry, eliminating a separate allocation and one memory hop (−8 bytes/key); 8.1 rewrote the dictionary into a 64-byte, cache-line-aligned hashtable that embeds both key and value, collapsing the lookup path from four hops to about two and removing dictEntry altogether (roughly −20 to −30 bytes per entry); and 9.1 reused the embedded-string pointer field to store short string data inline instead of pointing to it, cutting up to 20% off strings under 128 bytes.
Together, these changes show that Valkey’s memory efficiency gains came from targeted, low-level restructuring of core data structures rather than a single sweeping rewrite.
Two approaches to the same data:


Same data, same record count. Approach A uses 220 MB; Approach B uses 100 MB — a 54% reduction. That gap comes from the fact that every key pays a fixed per-key overhead: four keys pay it four times, one key pays it once.
Beyond raw memory — where Hash wins because it shares per-key overhead, there’s per-field TTL, where String is native with EXPIRE per key, while Hash didn’t natively support field-level TTL before 7.4.
For atomic multi-field operations, Hash wins with a single HSET or HMGET, while String needs multiple round trips.
For cluster slot locality, String needs a manual {hashtag} to co-locate related keys, while a Hash is automatically one key, one slot.
So there’s no universally correct choice — it’s memory versus flexibility, and it interacts with your cluster routing strategy too.

Testing five Valkey versions (7.2.14 through 9.1.1) under identical conditions, per-key memory for a simple SET workload dropped from 102.48 bytes to 64.02 bytes — a cumulative 37.5% reduction, with 8.1 alone responsible for the largest single jump (−23.8%) and 9.0 showing essentially no change. Three targeted structural rewrites drove this: 8.0 inlined the key into dictEntry, removing a separate allocation (−8 bytes/key); 8.1 replaced the chained hash table with a 64-byte, cache-line-aligned hashtable that embeds key and value together, cutting the lookup path from four hops to two and eliminating dictEntry entirely (−20 to −30 bytes/entry); and 9.1 reused the embedded-string pointer field to store short string data inline, saving up to 20% for strings under 128 bytes. Separately, how the same data is modeled also matters as much as the version: storing a user’s data as 4 individual String keys costs 220 MB, versus 100 MB for the same fields grouped into a single Hash — a 54% saving driven by paying the fixed per-key overhead once instead of four times. That memory advantage comes with trade-offs, though: Hash lags String on native per-field TTL support (pre-7.4) but wins on atomic multi-field operations and automatic cluster slot locality, so the right choice depends on whether memory efficiency or per-field flexibility matters more for the use case.