How to code a C# mutex block that is conditional on an entity id?

Viewed 520

I'm looking for a C# pattern for coding a synchonized operation including writes to two different databases for a particular entity such that I can avoid race conditions for simultaneous operations on the same entity.

E.g. Thread 1 and thread 2 are processing an operation on entity X at the same time. The operation writes information for X to database A (in my case, an upsert to MongoDB) and database B (an insert to SqlServer). Thread 3 is processing the same operation on entity Y. The desired behavior is:

  • Thread 1 blocks thread 2 while processing writes to A and B for entity X.
  • Thread 2 waits until thread 1 completes writes to A and B and then makes writes to A and B for entity X.
  • Thread 3 is not blocked and processes writes to A and B for entity Y while thread 1 is processing

The behavior I'm trying to avoid is:

  • Thread 1 writes to A for entity X.
  • Thread 2 writes to A for entity X.
  • Thread 2 writes to B for entity X.
  • Thread 1 writes to B for entity X.

I could uses a mutex across all threads, but I don't really want to block the operation for a different entity.

2 Answers

Using the lock statement is insufficient for multiple processes1. Even named/system semaphores are limited to a single-machine and thus insufficient from multiple servers.

If duplicate processing is OK and a "winner" can be selected, it may be sufficient just to write/update-over or use a flavor of optimistic concurrency. If stronger process-once concurrently guarantees need to be maintained, a global locking mechanism needs to be employed - SQL Server supports such a mechanism via sp_getapplock.

Likewise, the model can be updated so that each agent 'requests' the next unit of work such that dispatch can be centrally controlled and that an entity, based on ID etc., is only given to a single agent at a time for processing. Another option might be to use a Messaging system like RabbitMQ (or Kafka etc., fsvo); for RabbitMQ, one might even use Consistent Hashing to ensure (for the most part) that different consumers receive non-overlapping messages. The details differ based on implementation used.

Due to the different nature of a SQL RDBMS and MongoDB (especially if used as "a cache"), it may be sufficient to loosen the restriction and/or design the problem using MongoDB as a read through (which is a good way to use caches). This can mitigate the paired-write issue, although it does not prevent global concurrent processing of the same items.

1Even though a lock statement is globally insufficient, it can be still be employed locally between threads in a single process to reduce local contention and/or minimize global locking.


The answer below was for the original question, assuming a single process.

The "standard" method of avoiding working on the same object concurrently via multiple threads would be with a lock statement on the specific object. The lock is acquired on the object itself, such that lock(X) and lock(Y) are independent when !ReferenceEquals(X,Y).

The lock statement acquires the mutual-exclusion lock for a given object, executes a statement block, and then releases the lock. While a lock is held, the thread that holds the lock can again acquire and release the lock. Any other thread is blocked from acquiring the lock and waits until the lock is released.

lock (objectBeingSaved) {
  // This code execution is mutually-exclusive over a specific object..
  // ..and independent (non-blocking) over different objects.
  Process(objectBeingSaved);
}

A local process lock does not necessarily translate into sufficient guarantees for databases access or when then the access spills across processes. The scope of the lock should also be considered: eg. should it cover all processing, only saving, or some other work unit?

To control what objects are being locked and reduce the chance of undesired/accidental lock interactions, it's sometimes recommend to add a field of the most specific visibility to the objects explicitly (and only for) the purpose of establishing a lock. This can also be used to group objects which should lock on each other, if such is a consideration.

It's also possible to use a locking pool, although such tends to be a more 'advanced' use-case with only specific applicability. Using pools also allows using semaphores (in even more specific use-cases) as well as a simple lock.

If there needs to be a lock per external ID, one approach is to integrate the entities being worked on with a pool, establishing locks across entities:

// Some lock pool. Variations of the strategy:
// - Weak-value hash table
// - Explicit acquire/release lock
// - Explicit acquire/release from ctor and finalizer (or Dispose)
var locks = CreateLockPool();
// When object is created, assign a lock object
var entity = CreateEntity();
// Returns same lock object (instance) for the given ID, and a different
// lock object (instance) for a different ID.
etity.Lock = GetLock(locks, entity.ID);

lock (entity.Lock) {
  // Mutually exclusive per whatever rules are to select the lock
  Process(entity);
}

Another variation is a localized pool, instead of carrying around a lock object per entity itself. It is conceptually the same model as above, just flipped outside-in. Here is a gist. YMMV.

private sealed class Locker { public int Count; }

IDictionary<int, Locker> _locks = new Dictionary<int, Locker>();

void WithLockOnId(int id, Action action) {
  Locker locker;
  lock (_locks) {
     // The _locks might have lots of contention; the work
     // done inside is expected to be FAST in comparison to action().
     if (!_locks.TryGetValue(id, out locker)
        locker = _locks[id] = new Locker();
     ++locker.Count;
  }
  lock (locker) {
     // Runs mutually-exclusive by ID, as established per creation of
     // distinct lock objects.
     action();
  }
  lock (_locks) {
     // Don't forget to take out the garbage..
     // This would be better with try/finally, which is left as an exercise
     // to the reader, along with fixing any other minor errors.
     if (--_locks[id].Count == 0)
       _locks.Remove(id);
  }
}

// And then..
WithLockOnId(x.ID, () => Process(x));

Taking a sideways step, another approach is to 'shard' entities across thread/processing units. Thus each thread is guaranteed to never be processing the same entity as another thread: X,Y,Z always go to #1 and P,D,Q always to #2. (It's a little bit more complicated to optimize throughput..)

var threadIndex = entity.ID % NumThreads;
QueueWorkOnThread(threadIndex, entity); // eg. add to List<ConcurrentQueue>

I would suggest using simple lock (if it is in one area of the code) As it would be processing different objects (meaning .net objects) but having the same value (as it is the same entity) I would rather go with some form of code for entities. If the entity has some form of code I would use it - for example:

But of course, you have to watch out for deadlocks. And String.Intern is tricky, as it Interns the string for as long as application runs.

lock(String.Intern(myEntity.Code))
{
   SaveToDatabaseA(myEntity);
   SaveToDatabaseB(myEntity);
}

But it looks like you want to have some kind of replication mechanism. Then I would rather do it on database level (not on code level)

[UPDATE]

You updated the question with information, that it is being done on multiple servers. And this information is kind a crucial here :) Normal lock wont work.

Of course, you can play with synchronizing the locks across different servers, but is like with distributed transactions. Theoretically speaking you can do it, but most of the persons just avoid it as long as they can, and they play with the architecture of the solution to simplify the process.

[UPDATE 2]

You may also find this interesting: Distributed locking in .NET

:)

Related