I'm implementing a chunk-oriented shared ring-buffer for single producer, multiple consumers.
The buffer keeps track of the current iteration (i.e. wrap count) and the last written chunk index. When writing, the producer first marks the chunk's iteration field as dirty, fills in the other fields (offset/size/etc), then marks the iteration field as non-dirty, with the the same iteration number of the whole buffer. After that, it increments the last written chunk index appropriately. When reading, the consumers would read the chunk descriptor, verify the iteration is the expected one and it's not being written to, read the data, and verify again.
struct chunk_descriptor_t {
atomic_uint_fast32_t iteration;
atomic_uint_fast32_t chunk_offset;
atomic_uint_fast32_t chunk_size;
atomic_uint_fast32_t user_data;
};
struct shared_buf_descriptor_t {
atomic_uint_fast32_t total_data_size;
atomic_uint_fast32_t num_chunks;
atomic_uint_fast32_t iteration;
atomic_uint_fast32_t last_written_chunk_idx;
struct chunk_descriptor_t chunks_descriptors[];
};
typedef enum e_iter_flags {
ITER_VALUE_MASK = (1 << 30) - 1,
ITER_PRODUCER_DIRTY = 1 << 30,
ITER_PRODUCER_RUNNING = 1 << 31
} iter_flags_t;
The producer would write the next chunk like this:
static void add_chunk(shared_buf_producer_t producer, const char *data,
uint32_t chunk_size, uint32_t user_data) {
uint_fast32_t chunk_idx = (producer->last_written_chunk_idx + 1) %
producer->shbuf->header.num_chunks;
struct chunk_descriptor_t *chunk =
&producer->shbuf->descriptor->chunks_descriptors[chunk_idx];
// 1. mark the chunk dirty. This must happen first.
chunk->iteration =
ITER_PRODUCER_DIRTY | (producer->iteration & ITER_VALUE_MASK);
// 2. update the chunk fields and write data. This must happen after 1 and before 3.
// The order inside doesn't matter.
chunk->chunk_size = chunk_size;
chunk->user_data = user_data;
chunk->chunk_offset = producer->write_offset;
memcpy(producer->shbuf->header.data + producer->write_offset, data,
chunk_size);
// 3. mark the chunk clean. This must happen after 2 and before 4.
chunk->iteration = producer->iteration & ITER_VALUE_MASK;
// 4. update the last written index. This must happen last.
producer->shbuf->descriptor->last_written_chunk_idx = chunk_idx;
// producer bookkeeping
producer->write_offset += chunk_size;
producer->last_written_chunk_idx = chunk_idx;
}
So I need to ensure the memory ordering. Currently, memory_order_seq_cst is used for all stores. However, this would also order the stores in (2), where the order inside isn't important, as long as (2) as a whole comes after (1) and before (3).
I wonder if I can relax the ordering in (2). And if so, should those be memory_order_relaxed stores, or other fields should not be atomic at all maybe?