Valkey Memory Optimization, Version by Version and Encoding: How Many Bytes Did Each Release Actually Save?

September 29, 2026
Author
Percona Team
Share this Post:

Introduction

This RAM/Memory reduction is most useful for Valkey 9.1 environments that are key string heavy. They are used extensively in caching, sessions/authentication, real-time apps, ecomm, Ad tech/martech, and gaming.
These scenarios see a reduction of 37.5%, which has been attained in our testing environments versus version 7.2.. The biggest jump in the reduction came between 7.2 and 8.1 where the difference was 29.8%. Although the reduction doesn’t apply to every Valkey environment, those that are string heavy should experience 37.5% RAM reductions.

RAM usage scenario:

Starting point: 1 TB on Valkey 7.2

  • Valkey 7.2: 1,000 GB
  • Valkey 8.0: ~800 GB
  • Valkey 8.1: ~640 GB
  • Valkey 9.1: ~630 GB

This post is based on a rigorous benchmark and a structural teardown of the source, aimed at answering three questions:

    1. For the same data, how much memory does each version actually save?
    2. Where exactly does that saving come from at the data-structure level?
    3. How do different data structures and encoding result in significantly different memory footprints?

 

1. Test Setup: 

Five versions were tested — 7.2.14, 8.0.10, 8.1.9, 9.0.5, and 9.1.1 — all built from source, all using jemalloc 5.3.0 as the allocator, so version is the only variable.

Test Environment settings:

  • Standalone mode, not cluster, to rule out cluster-level overhead
  • –maxmemory 0, –save “” –appendonly no, to keep persistence and eviction out of the picture
  • valkey-benchmark -t set -n 1000000 -r 1000000 -d 16: one million SETs, 16-byte values
  • Both -n and -r are set to 1,000,000 on purpose — this causes random-key collisions, so only about 63.2% (~632K) land as unique keys. This matches the official Valkey benchmark methodology exactly
  • The baseline (empty instance) memory is subtracted for each version, giving the true marginal bytes per key

 

2. Results: Memory per Key, Version by Version

Version Unique Keys used_memory Bytes / Key vs Prev vs 7.2.14
7.2.14 632,181 61.78 MB 102.48 B — baseline
8.0.10 632,565 57.00 MB 94.49 B −7.8% −7.8%
8.1.9 632,573 43.43 MB 71.98 B −23.8% −29.8%
9.0.5 631,672 43.36 MB 71.99 B +0.0% −29.8%
9.1.1 631,985 38.58 MB 64.02 B −11.1% −37.5%

We can see that across roughly 632K unique keys, the size of per key averages 102.48 bytes in version 7.2.14 down to 64.02 bytes in 9.1.1 — a cumulative drop of 37.5%. In other words, the same hardware stores roughly 60% more of these small keys on 9.1 than it did on 7.2. We can see 8.1 alone accounts for the single biggest jump, at -23.8%, because that’s where the hash table got rewritten. However 9.0 shows basically zero change, since that release simply wasn’t targeting memory.

 

3. Let us see how the change happens version by version.

7.2 → 8.0: Embedding the Key into dictEntry

Valkey 7.2 dictEntry diagram: a 24-byte struct of three pointers (*key, *val, *next), with *key pointing to separately allocated SDS key data and *val pointing to the robj value

 

In 7.2, dictEntry is a classic three-pointer struct — *key, *val, *next — 24 bytes total. The catch is that the key itself is a pointer to a separately allocated SDS blob: every lookup pays for an extra malloc and an extra memory hop just to read the key.

 

Valkey 8.0 embeddedDictEntry diagram: *val and *next pointers with the key bytes stored inline in one allocation (16 bytes plus key bytes), pointing to the robj value

 

8.0 introduces embeddedDictEntry: the key’s bytes are stored inline right after the struct. *val, *next, plus the inlined key data come to 16 bytes plus the key length — one allocation instead of two.

Result: −8 bytes per key, one fewer memory hop on every lookup, automatic on upgrade, no config changes needed.

8.0 → 8.1: A Cache-Line-Sized Hash Table

This is the single largest jump in the whole series. In 8.0, the dictionary is a classic chained hash table: reaching a value means “bucket → entry → key → value” — four memory hops (plus two more per hash collision).

 

Valkey 8.0 chained dict diagram: hash table slots pointing to dictEntry structs, with collisions chained to the next entry

 

8.1 rewrites this as a new hashtable structure where one bucket is exactly 64 bytes — the size of a CPU cache line, — with the key and value embedded directly inside it.The bucket’s 8-byte metadata region packs a “has-child-bucket” bit, 7 filled-bits, and 7 bytes of secondary hash, letting most non-matching candidates be skipped without an extra memory access. Key and value are embedded in the same serverObject, shortening the path to roughly two hops — “bucket → serverObject” — and a hash collision usually stays within the same cache line.

dictEntry as a struct disappears entirely. The result is roughly −20 bytes per key-value pair (−30 bytes for keys with a TTL) — the single biggest jump in the entire 7.2→9.1 series.

 

Valkey 8.1 hashtable diagram: one 64-byte cache-line bucket with an 8-byte metadata region and seven entry slots, pointing to a serverObject with the key and value embedded, about two memory hops per lookup

9.0 → 9.1: Reusing the Pointer Field Itself

 

Valkey 9.0 and earlier embstr diagram: robj header (type/encoding, refcount, LRU) with an 8-byte *ptr field pointing to the SDS header and string bytes placed right after the robj

Valkey 9.1 robj diagram: the same 8 bytes previously used for *ptr now hold short string bytes inline, removing the pointer

In earlier versions, embedded strings (embstr) still wasted an 8-byte pointer field simply referencing the contiguous payload. 9.1’s fix: for short strings, the string bytes are stored directly inline in those same 8 bytes normally used for the pointer — no pointer needed, and no separate allocation for the value.

Result: up to a 20% memory reduction for strings under 128 bytes — and the benchmark’s “key + 16-byte value” pairs sit right in the middle of that range.

The three structural changes account for most of the improvement: 

8.0 inlined the key directly into dictEntry, eliminating a separate allocation and one memory hop (−8 bytes/key);    8.1 rewrote the dictionary into a 64-byte, cache-line-aligned hashtable that embeds both key and value, collapsing the lookup path from four hops to about two and removing dictEntry altogether (roughly −20 to −30 bytes per entry);   and 9.1 reused the embedded-string pointer field to store short string data inline instead of pointing to it, cutting up to 20% off strings under 128 bytes. 

Together, these changes show that Valkey’s memory efficiency gains came from targeted, low-level restructuring of core data structures rather than a single sweeping rewrite.

 

4. Same 1 Million Records, Two Encodings, 54% Less Memory

Two approaches to the same data:

  • Approach A: 4 separate String keys per user (SET user:1:name Alice, etc.)
  • Approach B: 1 Hash key per user, 4 fields (HSET user:1 name Alice age 30 city Toronto score 100)

Approach A: four separate Valkey String keys per user using SET user:1:name Alice, SET user:1:age 30, SET user:1:city Toronto, SET user:1:score 100

Approach B: one Valkey Hash key per user using HSET user:1 name Alice age 30 city Toronto score 100

 

Same data, same record count. Approach A uses 220 MB; Approach B uses 100 MB — a 54% reduction. That gap comes from the fact that every key pays a fixed per-key overhead: four keys pay it four times, one key pays it once.

Beyond raw memory — where Hash wins because it shares per-key overhead, there’s per-field TTL, where String is native with EXPIRE per key, while Hash didn’t natively support field-level TTL before 7.4. 

For atomic multi-field operations, Hash wins with a single HSET or HMGET, while String needs multiple round trips. 

For cluster slot locality, String needs a manual {hashtag} to co-locate related keys, while a Hash is automatically one key, one slot. 

So there’s no universally correct choice — it’s memory versus flexibility, and it interacts with your cluster routing strategy too.

 

Comparison table of Valkey String (separate keys) versus Hash (grouped fields) across memory at scale, per-field TTL, atomic multi-field operations, cluster slot locality and best fit

Summary

Testing five Valkey versions (7.2.14 through 9.1.1) under identical conditions, per-key memory for a simple SET workload dropped from 102.48 bytes to 64.02 bytes — a cumulative 37.5% reduction, with 8.1 alone responsible for the largest single jump (−23.8%) and 9.0 showing essentially no change. Three targeted structural rewrites drove this: 8.0 inlined the key into dictEntry, removing a separate allocation (−8 bytes/key); 8.1 replaced the chained hash table with a 64-byte, cache-line-aligned hashtable that embeds key and value together, cutting the lookup path from four hops to two and eliminating dictEntry entirely (−20 to −30 bytes/entry); and 9.1 reused the embedded-string pointer field to store short string data inline, saving up to 20% for strings under 128 bytes. Separately, how the same data is modeled also matters as much as the version: storing a user’s data as 4 individual String keys costs 220 MB, versus 100 MB for the same fields grouped into a single Hash — a 54% saving driven by paying the fixed per-key overhead once instead of four times. That memory advantage comes with trade-offs, though: Hash lags String on native per-field TTL support (pre-7.4) but wins on atomic multi-field operations and automatic cluster slot locality, so the right choice depends on whether memory efficiency or per-field flexibility matters more for the use case.

 

0 0 votes
Article Rating
Subscribe
Notify of
guest

0 Comments
Oldest
Newest Most Voted

Far
Enough.

Said no pioneer ever.
MySQL, PostgreSQL, InnoDB, MariaDB, MongoDB and Kubernetes are trademarks for their respective owners.
© 2026 Percona All Rights Reserved