ansaurus

Question

Of these 3 methods for reading linked lists from shared memory, why is the 3rd fastest?

Answer 1

A:

I've noticed in both method 1 and method 3 you have a line, ACQUIRE_MEMORY_BARRIER, which I assume has something to do with multi-threading/race conditions?

Either way, method 2 doesn't have any sleeps which means the following code...

while(true)
{   
    if(update_cursor->state_ == NOT_FILLED_YET) {
        continue;
    }

is going to hammer the processor. The typical way to do this kind of producer/consumer task is to use some kind of semaphore to signal to the reader that the update list has changed. A search for producer/consumer multi threading should give you a large number of examples. The main idea here is that this allows the thread to go to sleep while it's waiting for the update_cursor->state to change. This prevents this thread from stealing all the cpu cycles.

Pace 2010-03-28 02:59:17

Yes, it's for avoiding race conditions. It was supposed to be in method 2 as well, I've edited my post to fix that (timings are the same). Modern processors (not just the compiler) are free to reorder stores and loads across threads, so one thread might set A then B, but another thread may 'see' that B was set before it sees that A was set. The server fills the node's data, does a RELEASE_MEMORY_BARRIER, then sets state_ to FILLED. Combined with the client doing ACQUIRE_MEMORY_BARRIER, this ensures that the client never perceives state_ as FILLED before the node's data is filled in.

Joseph Garvin 2010-03-28 03:36:38

I should add that this is my first time using memory barriers, and that's how I *think* they're supposed to work ;) If you follow the pastebin link, the first file, list.H defines RELEASE_MEMORY_BARRIER and ACQUIRE_MEMORY_BARRIER to primitives documented here: http://gcc.gnu.org/onlinedocs/gcc-4.1.2/gcc/Atomic-Builtins.html

Joseph Garvin 2010-03-28 03:38:09

Sorry to spam ;p I've edited my post (see the bottom) to address the hammering of the processor.

Joseph Garvin 2010-03-28 03:40:10

Answer 2

+1 A:

Hypothesis: Method 2 is somehow blocking the update from getting written by the server.

One of the things you can hammer, besides the processor cores themselves, is your coherent cache. When you read a value on a given core, the L1 cache on that core has to acquire read access to that cache line, which means it needs to invalidate the write access to that line that any other cache has. And vice versa to write a value. So this means that you're continually ping-ponging the cache line back and forth between a "write" state (on the server-core's cache) and a "read" state (in the caches of all the client cores).

The intricacies of x86 cache performance are not something I am entirely familiar with, but it seems entirely plausible (at least in theory) that what you're doing by having three different threads hammering this one memory location as hard as they can with read-access requests is approximately creating a denial-of-service attack on the server preventing it from writing to that cache line for a few milliseconds on occasion.

You may be able to do an experiment to detect this by looking at how long it takes for the server to actually write the value into the update list, and see if there's a delay there corresponding to the latency.

You might also be able to try an experiment of removing cache from the equation, by running everything on a single core so the client and server threads are pulling things out of the same L1 cache.

Brooks Moses 2010-03-28 05:14:05

I put gettimeofday's around the write in the server, and it always takes 0 or 1 microseconds. But is it possible that the write finishes faster from the server's perspective than it does to the client for the reason you mentioned? i.e. is it plausible the processor is queueing up the write then 'returning' immediately? I'll try running it all on one core tomorrow. Also, I'm only trying 1 method at a time, so there is only the 1 server and 1 client competing.

Joseph Garvin 2010-03-28 05:19:07

That seems possible; I don't know enough about x86 caches to say for certain one way or the other -- or, for that matter, instruction reordering within the processor core. I'm handwaving here, a bit. (Also, somewhere I had the idea that you had multiple clients running in parallel for one server, I guess because you mentioned 4 cores. That doesn't really change things much, though.)

Brooks Moses 2010-03-28 06:04:05

I think Herb Sutter had written something about those multi-cores cache system. Is it possible for you to precise on which core you'd like to execute each thread ?

Matthieu M. 2010-03-28 14:15:01

Answer 3

+1 A:

I don't know if you have ever read the Concurrency columns from Herb Sutter. They are quite interesting, especially when you get into the cache issues.

Indeed the Method2 seems better here because the id being smaller than the data in general would mean that you don't have to do round-trips to the main memory too often (which is taxing).

However, what can actually happen is that you have such a line of cache:

Line of cache = [ID1, ID2, ID3, ID4, ...]
                  ^         ^
            client          server

Which then creates contention.

Here is Herb Sutter's article: Eliminate False Sharing. The basic idea is simply to artificially inflate your ID in the list so that it occupies one line of cache entirely.

Check out the other articles in the serie while you're at it. Perhaps you'll get some ideas. There's a nice lock-free circular buffer I think that could help for your update list :)

Matthieu M. 2010-03-28 14:25:59

Interesting I thought of the cache issue before but figured I'd be somewhat automatically protected from it by virtue of using a linked list as opposed to an array. I haven't actually tried comparing the addresses of the UpdateID objects though, it's possible the allocator is giving me relatively contiguous chunks.

Joseph Garvin 2010-03-28 14:39:10

Answer 4

A:

The answer was tricky to figure out, and to be fair would be hard with the information I presented though if anyone actually compiled the source code I provided they'd have a fighting chance ;) I said that "found via update list" was printed after every latency log message, but this wasn't actually true -- it was only true for as far as I could scrollback in my terminal. At the very beginning there were a slew of updates found without using the update list.

The issue is that between the time when I set my starting point in the update list and my starting point in each of the data lists, there is going to be some lag because these operations take time. Remember, the lists are growing the whole time this is going on. Consider the simplest case where I have 2 data lists, A and B. When I set my starting point in the update list there happen to be 60 elements in it, due to 30 updates on list A and 30 updates on list B. Say they've alternated:

A
B
A
B
A // and I start looking at the list here
B

But then after I set the update list to there, there are a slew of updates to B and no updates to A. Then I set my starting places in each of the data lists. My starting points for the data lists are going to be after that surge of updates, but my starting point in the update list is before that surge, so now I'm going to check for a bunch of updates without finding them. The mixed approach above works best because by iterating over all the elements when it can't find an update, it quickly closes the temporal gap between where the update list is and where the data lists are.

Joseph Garvin 2010-04-14 21:52:21

ansaurus

tags:

views:

answers:

Of these 3 methods for reading linked lists from shared memory, why is the 3rd fastest?

related questions