ansaurus

Question

Fast sine/cosine for ARMv7+NEON: looking for testers…

Answer 1

+3 A:

Just tested it on my beagleboard.. As said in the comments: Same CPU.

Your code is roughly 15 times faster than the clib.. Well done!

I've measured 82 cycles for each call of your implementation and 1260 for the four c-lib calls. Note that I've compiled with soft-float ABI and my OMAP3 is early silicon, so each call to the the c-lib version has a NEON stall of at least 40 cycles.

I've zipped together the results..

http://torus.untergrund.net/code/sincos.zip

The performance-counter stuff will most likely not work on the iphone.

Hope that's what you've been looking for.

Nils Pipenbrinck 2009-12-06 20:05:46

Thank you very much Nils. I'm a bit surprised it worked out of the box, actually :-) The same method implemented for the VFP11 is only about twice as fast as calling sinf()+cosf() on my iPod Touch, so I used a lookup table instead.

jcayzac 2009-12-07 01:01:47

Ah I see from your test program that you used the double precision variant of the libc function (sin()/cos(), not sinf()/cosf()). That explains why the libc functions performed so poorly I think :-)

jcayzac 2009-12-07 01:22:14

Just compiled and run it with sinf/cosf, and it does not make much of a difference.

Nils Pipenbrinck 2009-12-07 01:31:46

Do you mean, there is almost no difference in results compared to when testing with sin()/cos(), versus sinf()/cosf(), or there is almost no difference in speed between my function and libc's?

jcayzac 2009-12-07 11:29:48

There is not much difference between sin and sinf (same for cos). I measure 1260 cycles/interation for the double-version and 1245 cycles/iteration for the float version. I bet it's the same code and we just save the cast from double to float.

Nils Pipenbrinck 2009-12-07 13:29:25

I use the gnu libc btw..

Nils Pipenbrinck 2009-12-07 13:30:03

Answer 2

A:

Oh - before I forget it: Maybe you can safe yourself a bit of work..

Take a look at these NEON optimized math functions:

http://code.google.com/p/math-neon/

Nils Pipenbrinck 2009-12-06 20:16:05

Yes I know these functions. I wish I had an iPhone 3GS for playing with NEON, which surely is much more fun than working on my iPod's VFP11. With the iPhone a FPU craze started, but before that the rule was: DON'T do floats, DO fixed point. I'm starting to think the guys coding with floats for the iPhone are wrong. ARM is really good with integers, and setting up float registers has a large overhead.

jcayzac 2009-12-07 01:05:35

In fact I know another related project, not this one. The one I had in mind had only matrix/vector functions.Just checked the one you put a link to. It looks like it has all the math.h functions! I'm starring it, thanks! :-)

jcayzac 2009-12-07 01:12:15

ansaurus

tags:

views:

answers:

Fast sine/cosine for ARMv7+NEON: looking for testers…

related questions