Sequential Consistency in C++ Memory Model

Understanding std::memory_order_seq_cst through a practical lock-free Producer-Consumer pattern and formal synchronization analysis.


Table of Contents

  1. Sequential Consistency in C++ Memory Model
    1. Table of Contents
    2. Overview
      1. Trade-offs: Intuition vs. Hardware Cost
    3. Practical Example: Lock-Free Producer-Consumer
    4. Formal Proof of Correctness
      1. 1. Intra-Thread Sequencing (Sequenced-Before)
      2. 2. Inter-Thread Synchronization (Synchronizes-With)
      3. 3. Transitive Execution Chain
    5. Transitioning to Acquire-Release Semantics

Overview

Modern C++ Concurrency Lock-Free

In C++ multithreading, sequential consistency (std::memory_order_seq_cst) is the default memory ordering for atomic operations. It provides an intuitive mental model: operations appear to execute in a strict single-thread-like sequential order, shared consistently across all threads.

Trade-offs: Intuition vs. Hardware Cost

  • Advantage: Sequential consistency matches our intuitive understanding of program execution. Every thread observes operations in source-code order, as if all threads execute on a single global clock.
  • Disadvantage: Achieving this global order requires the CPU and compiler to introduce heavyweight synchronization memory barriers, which inhibits certain reordering optimizations and incurs performance overhead.

Practical Example: Lock-Free Producer-Consumer

Below is a classic Producer-Consumer implementation using atomic flags for signaling instead of std::mutex or std::condition_variable.

#include <atomic>
#include <iostream>
#include <string>
#include <thread>

std::string work;
std::atomic<bool> ready{false};

void consumer() {
    // Consumer polls ready flag
    while (!ready.load()) {} 
    
    // Guaranteed to see "done" because of sequential consistency
    std::cout << work << std::endl; 
}

void producer() {
    work = "done";       // Non-atomic assignment
    ready.store(true);   // Atomic store (seq_cst by default)
}

int main() {
    std::thread t1(consumer);
    std::thread t2(producer);
    
    t1.join();
    t2.join();
}
Execution Output & Explanation

Output:

done

Even though work is a non-atomic variable, the program is completely free of data races. The atomic operations on ready establish a strict boundary ensuring work is fully written before it is read.


Formal Proof of Correctness

To prove why this pattern is deterministic and free of data races, we analyze the execution order using formal relations within the C++ Memory Model.

1. Intra-Thread Sequencing (Sequenced-Before)

Within a single thread, evaluation steps follow program order:

  • In producer(): work = "done" is sequenced-before ready.store(true).
  • In consumer(): while (!ready.load()) is sequenced-before std::cout << work.
\[\text{work = "done"} \xrightarrow{\text{happens-before}} \text{ready.store(true)}\] \[\text{while (!ready.load())} \xrightarrow{\text{happens-before}} \text{std::cout << work}\]

2. Inter-Thread Synchronization (Synchronizes-With)

Under sequential consistency, an atomic store synchronizes with an atomic load that reads the written value across threads:

\[\text{ready.store(true)} \xrightarrow{\text{synchronizes-with}} \text{while (!ready.load())}\]

This inter-thread synchronization establishes a global happens-before relation between the threads.

3. Transitive Execution Chain

Combining intra-thread sequencing and inter-thread synchronization yields the total ordering chain across threads:

\[\text{work = "done"} \xrightarrow{\text{happens-before}} \text{ready.store(true)} \xrightarrow{\text{happens-before}} \text{while (!ready.load())} \xrightarrow{\text{happens-before}} \text{std::cout << work}\]

Because work = “done” strictly happens-before std::cout « work, the consumer thread is mathematically guaranteed to observe "done".


Transitioning to Acquire-Release Semantics

While std::memory_order_seq_cst provides strong guarantees, it forces all atomic operations globally into a single total order.

In performance-critical code, this same Producer-Consumer guarantee can be achieved with lower overhead using Acquire-Release Semantics:

  • ready.store(true, std::memory_order_release); in the producer.
  • ready.load(std::memory_order_acquire); in the consumer.

Acquire-Release synchronizes only dependent operations between specific threads without enforcing a global clock across all threads, reducing CPU pipeline stalls.


This site uses Just the Docs, a documentation theme for Jekyll.