The Memory Layout of Rust Enums: Discriminants, Niche Optimization, and Alignment

The Memory Layout of Rust Enums: Discriminants, Niche Optimization, and Alignment

Rust enum sizing equals the largest variant payload plus discriminant and trailing alignment padding. Ignoring variant balance bloats hot collections, blowing past L1 cache boundaries. Exploiting invalid bit-pattern niches like non-null pointers eliminates discriminant overhead, preserving cache residency in high-throughput systems.

In production ring buffers, packet decoders, and event-driven actor engines, naive Rust enum declarations routinely trigger catastrophic memory inflation and silent CPU cache thrashing. Software engineers casually add a single 512-byte diagnostics payload to an otherwise lean state machine enum, unaware that rustc sizes tagged unions to accommodate the maximum variant footprint across all instances. When instantiating an array or Vec of ten million elements, this unexamined variant balance forces gigabytes of dead zero-padding into physical memory, evicting hot instructions and data from L1 and L2 caches, choking hardware prefetchers, and degrading system throughput by orders of magnitude.

use std::mem::{align_of, size_of};
use std::num::NonZeroU64;

#[repr(C)]
pub struct DiagnosticPayload {
    pub error_code: u64,
    pub raw_buffer: [u8; 504],
}

// Case 1: The bloated variant hazard
pub enum IngestionEvent {
    Heartbeat,                       // 0 bytes payload
    Ack(u64),                        // 8 bytes payload
    Error(DiagnosticPayload),        // 512 bytes payload
}

// Case 2: Standard alignment padding tax
pub enum ScalarMetric {
    Inactive,                        // 0 bytes payload
    Sample(u64),                     // 8 bytes payload
}

// Case 3: Zero-cost pointer niche optimization
pub type DirectReference<'a> = Option<&'a DiagnosticPayload>;

// Case 4: NonZero integer niche optimization
pub type CompactIdentifier = Option<NonZeroU64>;

fn main() {
    println!("=== RUST ENUM MEMORY PROFILE ===");
    println!("Variant 1 (IngestionEvent):    size = {}B, align = {}B", size_of::<IngestionEvent>(), align_of::<IngestionEvent>());
    println!("Variant 2 (ScalarMetric):     size = {}B, align = {}B", size_of::<ScalarMetric>(), align_of::<ScalarMetric>());
    println!("Variant 3 (DirectReference):  size = {}B, align = {}B", size_of::<DirectReference>(), align_of::<DirectReference>());
    println!("Variant 4 (CompactId):        size = {}B, align = {}B", size_of::<CompactIdentifier>(), align_of::<CompactIdentifier>());
}

Execution of this profile on x86_64 yields the following raw physical allocations:

  • IngestionEvent: 520 bytes (align: 8 bytes)
  • ScalarMetric: 16 bytes (align: 8 bytes)
  • DirectReference (Option<&T>): 8 bytes (align: 8 bytes)
  • CompactIdentifier (Option<NonZeroU64>): 8 bytes (align: 8 bytes)

Why Does Enum Variant Size Disparity Cause Memory Bloat?

Enum memory bloat occurs when Rust sizes tagged unions according to the largest variant payload plus required alignment padding. Unbalanced variants force dead padding into contiguous memory buffers, multiplying allocation size across arrays, thrashing L1 instruction and data caches, and degrading hardware memory bus throughput.

Standard x86_64 and ARM64 microarchitectures service memory hierarchies via fixed 64-byte cache lines. When the processor pulls data from DRAM into the L1 data cache (L1d), it does not fetch individual fields; it fetches a 64-byte burst over the memory interconnect. In the case of IngestionEvent, every single slot in a contiguous buffer requires 520 bytes, which spans nine discrete cache lines ($520 / 64 = 8.125$). Even if 99.99% of network ingress traffic consists of zero-payload Heartbeat signals, the hardware cache is forced to retain 512 bytes of completely uninitialized, useless padding per element.

A collection of 1,000,000 Heartbeat instances under this schema consumes 520 megabytes of physical RAM instead of the single megabyte theoretically required to hold the tag data. The effective L1d cache capacity plummets by over 99%. Hardware prefetchers fail to detect coherent spatial locality across nine cache lines per logical event, forcing the CPU pipeline to stall on memory load operations while waiting hundreds of clock cycles for DRAM line fills.

How Does Rust Utilize Niche Optimization To Eliminate Discriminants?

Niche filling optimization occurs when the rustc compiler repurposes illegal bit patterns within payload fields to represent distinct enum variants without allocating separate discriminant memory. Non-null pointers, NonZero integers, and bounded types embed variant identities directly into otherwise unused bit states, preventing cache-busting alignment bloat.

In traditional C-style tagged unions, representing an optional pointer requires a tag field (enum Tag { None, Some }) followed by the raw address. Because an 8-byte pointer mandates 8-byte alignment, the 1-byte tag incurs 7 bytes of interior alignment padding, forcing the composite structure to 16 bytes. Rust avoids this overhead entirely by understanding type-level invariants through the compiler’s layout engine.

A reference (&T), a non-null pointer (core::ptr::NonNull<T>), and a Box<T> share a strict invariant: their internal 64-bit addresses can never be 0x0000000000000000 (null). A null address is an illegal state a “niche.” The compiler’s layout algorithm flags this zero value as an available bit pattern. When evaluating Option<&T>, None is mapped directly to 0x0000000000000000, while Some(&T) contains the actual, non-zero pointer address. The runtime discriminant is stored directly inside the pointer’s memory location without adding a single byte of memory overhead.

This optimization extends to integer types bounded by non-zero invariants, such as std::num::NonZeroU64. Because zero is prohibited by the type constructor, Option<NonZeroU64> requires exactly 8 bytes. Conversely, an ordinary u64 occupies all $2^{64}$ bit combinations ($0$ through $2^{64} – 1$), leaving zero niches. Therefore, Option<u64> must allocate an independent 1-byte tag, which trailing padding forces to 16 bytes total.

Why Does Struct Alignment Force Trailing Padding In Enums?

Discriminant alignment padding occurs when the compiler places integer tags alongside large, strictly aligned variant fields, mandating trailing padding bytes to maintain proper memory alignment for array allocations. Hardware memory architectures require word-boundary alignment, causing single byte tags to consume four or eight full bytes.

Hardware platforms mandate alignment for multi-byte primitives to avoid unaligned memory access penalties. On modern x86_64 processors, reading an 8-byte scalar crossing a 64-byte cache line boundary requires two discrete memory transactions, split-lock bus contention, and microcode assists. On architectures like ARM64, misaligned access can trigger unaligned fault traps handled by the kernel at massive cycle costs.

In ScalarMetric, the discriminant tag requires 1 byte. The payload for ScalarMetric::Sample(u64) requires 8 bytes, which dictates an alignment requirement of 8 bytes for the entire enum. The compiler places the 1-byte tag at offset 0, followed by 7 bytes of interior padding to place the u64 payload at offset 8. The total size reaches 16 bytes. If an enum’s variants have large alignment requirements, even tiny variants waste substantial space to preserve alignment across consecutive array elements.

The standard mitigation for high-throughput architectures involves boxing the heavyweight payload:

// Heavyweight variant moved out of the hot cache line
pub enum OptimizedIngestionEvent {
    Heartbeat,
    Ack(u64),
    Error(Box<DiagnosticPayload>), // 8-byte pointer replaces 512-byte array
}

Replacing DiagnosticPayload with Box<DiagnosticPayload> reduces the size of the enum from 520 bytes to 16 bytes. An array of 1,000,000 events drops from 520 MB to 16 MB. The hot contiguous memory path now fits entirely within modern L3 CPU caches, moving the 512-byte payload to cold, heap-allocated memory that is only traversed when exceptions occur.

Technical Troubleshooting FAQ

Why does rustc output “warning: large size difference between variants”?

This diagnostic occurs when the compiler detects that one variant dominates the enum layout, forcing smaller variants to allocate excessive unused padding. Fix this issue by wrapping the oversized variant’s payload in an indirect heap pointer, such as Box<T>, to collapse the enum size to pointer-width dimensions.

Why does size_of::<Option>() equal 1 while size_of::<Option>() equals 2?

The bool primitive only uses bit patterns 0x00 (false) and 0x01 (true), leaving 254 unused niche states (0x02 to 0xFF) within its single allocated byte to store None. The u8 primitive consumes all 256 possible bit patterns, forcing the compiler to append a separate discriminant byte plus padding.

The architectural decision between contiguous bloated cache lines versus pointer indirection via heap allocation is a brutal zero-sum compromise. Contiguous bloated layouts destroy L1d cache residency and bottleneck hardware memory buses on dead padding, while pointer indirection preserves cache density at the expense of heap allocator churn, memory fragmentation, and TLB miss latency on the cold execution path.

References

You may also like

See All Posts →