ansaurus

Question

What's the difference between logical SSE intrinsics?

Answer 1

A:

I think all three are effectively the same, i.e. 128 bit bitwise operations. The reason different forms exist is probably historical, but I'm not certain. I guess it's possible that there may be some additional behaviour in the floating point versions, e.g. when there are NaNs, but this is pure guesswork. For normal inputs the instructions seem to be interchangeable, e.g.

#include <stdio.h>
#include <emmintrin.h>
#include <pmmintrin.h>
#include <xmmintrin.h>

int main(void)
{
    __m128i a = _mm_set1_epi32(1);
    __m128i b = _mm_set1_epi32(2);
    __m128i c = _mm_or_si128(a, b);

    __m128 x = _mm_set1_ps(1.25f);
    __m128 y = _mm_set1_ps(1.5f);
    __m128 z = _mm_or_ps(x, y);

    printf("a = %vld, b = %vld, c = %vld\n", a, b, c);
    printf("x = %vf, y = %vf, z = %vf\n", x, y, z);

    c = (__m128i)_mm_or_ps((__m128)a, (__m128)b);
    z = (__m128)_mm_or_si128((__m128i)x, (__m128i)y);

    printf("a = %vld, b = %vld, c = %vld\n", a, b, c);
    printf("x = %vf, y = %vf, z = %vf\n", x, y, z);

    return 0;
}

$ gcc -Wall -msse3 por.c -o por

$ ./por

a = 1 1 1 1, b = 2 2 2 2, c = 3 3 3 3
x = 1.250000 1.250000 1.250000 1.250000, y = 1.500000 1.500000 1.500000 1.500000, z = 1.750000 1.750000 1.750000 1.750000
a = 1 1 1 1, b = 2 2 2 2, c = 3 3 3 3
x = 1.250000 1.250000 1.250000 1.250000, y = 1.500000 1.500000 1.500000 1.500000, z = 1.750000 1.750000 1.750000 1.750000

Paul R 2010-05-10 18:42:10

ORPD/ORPS are SSE-only, not MMX.

Potatoswatter 2010-05-10 19:10:16

@Potatoswatter: sorry - I meant 64-bit SSE (1) - updated.

Paul R 2010-05-10 20:39:57

But Intel introduced `orps` and later `orpd` *both after* `por`. And the physical basis of SSE has never changed much.

Potatoswatter 2010-05-10 20:53:08

The physical basis of SSE has changed a *lot*, particularly since Woodcrest, when it finally became a full 128 bit unit. However that's probably irrelevant - it sounds like I may be wrong about why there are separate bitwise OR instructions - I thought it was a legacy thing to do with switching context between integer and floating point SSE operations in the old days, but perhaps not.

Paul R 2010-05-10 21:06:41

Answer 2

A:

According to Intel and AMD optimization guidelines mixing op types with data types produces a performance hit as the CPU internally tags 64 bit halves of the register for a particular data type. This seems to mostly effect pipe-lining as the instruction is decoded and the uops are scheduled. Functionally they produce the same result. The newer versions for the integer data types have larger encoding and take up more space in the code segment. So if code size is a problem use the old ops as these have smaller encoding.

Phernost 2010-08-20 19:36:04

ansaurus

tags:

views:

answers:

What's the difference between logical SSE intrinsics?

related questions