Alternatives to temporarily disabling the scheduler in linux for a critical section

Viewed 68

I am porting code written for a real time OS on linux and I have run into a problem.

Context: The code has a number of global variables that can be read and written by two threads. The way these two threads interact with variables is as follows:

  • Thread "A" waits for a "message" on a queue. This thread runs with scheduling policy SCHED_RR and has a priority of "1". Upon receipt of the message and based on the latter, it performs operations on the variables.
  • Thread "B" waits for an event. This thread runs with scheduling policy SCHED_RR and has a priority of "2". Upon receiving the event, it calls a function of an external library, which can read or write these global variables. I have no access to the external library code and no ability to modify its content. I have no knowledge of what is done in it other than reading/writing to these global variables (there may be blocking calls like "sleep"). This function must be therefore considered as a black box function.

The problem is with the synchronization of these threads with regards to accessing global variables. In the original code, synchronization was implemented by temporarily disabling the preemptive thread switch upon receipt of the message on thread "A" (using a feature made available by the real time operating system).

Pseudocode of the original code:

structure_t g_structure;
int g_number;
char* g_string;
bool g_boolean;

void thread_A()
{
    while(true)
    {
        int message = queue.wait();
        OS_DISABLE_PREEMPT();
        switch(message)
        {
            case 1:
                g_number = 100;
                strcpy(g_string, "Message1");
                break;
            
            case 2:
                g_number = 200;
                strcpy(g_string, "Message2");
                g_boolean = true;
                g_structure.field1 = g_number;
                break;
            
            case 3:
                g_number = 200;
                strcpy(g_string, "Message3");
                g_structure.field2 = g_boolean;
                break;
        }
        OS_ENABLE_PREEMPT();
    }
}

void thread_B()
{
    while(true)
    {
        event.get();
        ExternalLibraryFunction();
    }
}

Since this operation is not possible on linux I started looking for solutions and these are the ones that came to my mind:

Solution 1: Using a mutex

structure_t g_structure;
int g_number;
char* g_string;
bool g_boolean;
mutex g_mutex;

void thread_A()
{
    while(true)
    {
        int message = queue.wait();
        g_mutex.lock();
        switch(message)
        {
            case 1:
                g_number = 100;
                strcpy(g_string, "Message1");
                break;
            
            // ... other cases ..
        }
        g_mutex.unlock();
    }
}

void thread_B()
{
    while(true)
    {
        event.get();
        g_mutex.lock();
        ExternalLibraryFunction();
        g_mutex.unlock();
    }
}

This solution involves securing access to global variables through a shared mutex between the two threads. However, this solution has a problem: Since I am not aware of the content of the function on the external library, I cannot exclude that there are blocking calls inside. The problem is that these blocking calls would keep the mutex locked, preventing thread "A" from running even when thread "B" is waiting for something (such as an event). This solution cannot therefore be used..

Solution 2: Temporarily increment thread priority

structure_t g_structure;
int g_number;
char* g_string;
bool g_boolean;
mutex g_mutex;

void enter_cs()
{
    struct sched_param param;
    param.sched_priority = sched_get_priority_max(SCHED_RR);
    pthread_setschedparam(pthread_self(), SCHED_RR, &param);
}

void leave_cs()
{
    struct sched_param param;
    param.sched_priority = RESTORE_OLDER_PRIORITY;
    pthread_setschedparam(pthread_self(), SCHED_RR, &param);
}

void thread_A()
{
    while(true)
    {
        int message = queue.wait();
        enter_cs();
        switch(message)
        {
            case 1:
                g_number = 100;
                strcpy(g_string, "Message1");
                break;
            
            // ... other cases ..
        }
        leave_cs();
    }
}

void thread_B()
{
    while(true)
    {
        event.get();
        ExternalLibraryFunction();
    }
}

This solution foresees to temporarily raise the priority of thread "A" to ensure that its execution cannot be interrupted by thread "B" in the event that it becomes READY. This solution does not have the problem of the previous one which uses mutexes and therefore seems better to me, however I don't know what can be the side effects of dynamically changing thread priorities on linux.

What can be the problems caused by this second solution? Are there any alternatives that I haven't considered?

EDIT: Forgot to mention that this is expected to run on a uniprocessor system, so only one thread at a time can actually run.

EDIT 2: User Aconcagua suggested to use only one thread and wait on both "thread A" queue and "thread B" event by using something like select. This is another solution I hadn't thought of; However, it has the same problem as the solution with the mutex.

Consider the situation below (this is pseudocode):

bool g_boolean;

void unified_loop()
{
    while(true)
    {
        select_result = select();
        if(select_result.who() == thread_A_queue)
        {
            switch(select_result.data)
            {
                case 1:
                    g_boolean = true;
                    break;
            }
        }
        else if(select_result.who() == thread_B_event)
        {
            ExternalLibraryFunction();
        }
    }
}

void ExternalLibraryFunction()
{
    // REMEMBER: I have no control over this code
    while(g_boolean == false)
    {
        sleep_milliseconds(100);
    }
}

In this case, the ExternalLibraryFunction function would block everything as the global variable g_boolean can never be set.

1 Answers

There are two threads running on a single CPU, both waiting for some kind of task (event raising, message arriving), and none of the two threads should interrupt the other one while the latter is processing its respective task.

In such a scenario I do not see the necessity to keep two separate threads at all; a single thread approach appears more suitable to me; such an approach might look as follows:

struct pollfd tasks[] =
{
    { .fd = queue.fd(), .events = POLLIN },
    { .fd = event.fd(), .events = POLLIN },
    // could be extended for yet some other tasks
    // or one for handling clean shutdown, if need be   
};

for(;;)
{
    int n = poll(tasks, sizeof(tasks)/sizeof(*tasks), -1);
    // TODO: add error handling for n <= 0
    // especially: call can get interrupted by signals! -> EINTR
    if(tasks[0].revents)
    {
        // handle incoming message
    }
    if(tasks[1].revents)
    {
        // handle event/call external function
    }
}  

If one of the tasks works based on signals ppoll might be the appropriate alternative, and if one of the events cannot provide a file descriptor you could add one explicitly with eventfd – though that might require some modifications of your event and/or queue facilities.

There are modifications possible, e. g. if you need to read all messages before you restart the event processing you could simply add an else (if(tasks[0].revents) { ... } else if(tasks[1].revents) { ... }) and invert the order if the other way round, apply priorities (if both are set process one of only after you have processed the other one n times already), ...

Related