ansaurus

Question

how to substitute a for loop in R with an optimized function (lapply?)

Answer 1

+3 A:

The usual solution is to think hard enough about the problem to find something vectorized.

If that fails, I sometimes resort to re-writing the loop in C++; the Rcpp package can helps with the interface.

Dirk Eddelbuettel 2010-09-10 12:01:23

In this case, if type2 always follows type1, you could do some very simple subsetting to get the answer.

hadley 2010-09-10 13:59:14

Answer 2

+2 A:

The *apply suite of functions are not optimized for loops. Further, I've worked on problems where for loops are faster than apply because apply used more memory and caused my machine to swap.

I would suggest fully initializing the myDelay object and avoid using rbind (which must re-allocate memory):

init <- rep(NA, length(dDelay$timeid))
myDelay <- data.frame(time=init, delay=init, jitter=init, nLost=init)

then replace:

myDelay <- rbind(myDelay, c(aRec$timeid, dd, jit, nLost))

with

myDelay[i,] <- c(aRec$timeid, dd, jit, nLost)

Joshua Ulrich 2010-09-10 12:17:38

Could be faster with `data.frame` replaced by `matrix`. Which could be converted to `data.frame` at the end. Inserting to `matrix` is faster then to `data.frame`

Marek 2010-09-10 14:40:10

Answer 3

+1 A:

As Dirk said: vectorization will help. An example of this would be to move the call to as.numeric out of the loop (since this function works with vectors).

dDelay$timeid <- as.numeric(dDelay$timeid)

Other things that may help are

Not bothering with the line aRec <- dDelay[itr,], since you can just access the row of dDelay, without creating a new variable.

Preallocating myDelay, since having it grow within the loop is likely to be a bottleneck. See Joshua's answer for more on this.

Richie Cotton 2010-09-10 14:06:49

Indexing play huge role in optimization. Check for example `ind<-rep(1,1e5);X<-data.frame(a=1,b=2,c=3)` and compare `system.time(for (i in ind) {X[i,1];X[i,2];X[i,3]})` (~15sec) vs `system.time(for (i in ind) {X$a[i];X$b[i];X$b[i]})` (~1sec).

Marek 2010-09-10 14:37:16

Answer 4

A:

Another optimization : If I read your code right, you can easily calculate the vector nLost by using :

nLost <-cumsum(dDelay$typeid==1)

outside the loop. That one you can just add to the dataframe in the end. Saves you a lot of time already. If I use your dataframe, then :

> nLost <-cumsum(dd$typeid==1)
> nLost
 [1] 1 1 2 2 3 3 4 4 5 5

Likewise the times at which the packages were lost can be calculated as:

> dd$timeid[which(dd$typeid==1)]
[1] 18,00035 18,02035 18,04035 18,06035 18,08035

in case you want to report them somewhere too.

For testing, I used :

dd <- structure(list(timeid = structure(1:10, .Label = c("18,00035", 
"18,00528", "18,02035", "18,02116", "18,04035", "18,04116", "18,06035", 
"18,06116", "18,08035", "18,08116"), class = "factor"), valid = structure(c(3L, 
2L, 4L, 1L, 5L, 1L, 6L, 1L, 7L, 1L), .Label = c("0,00081", "0,00493", 
"1,00000", "2,00000", "3,00000", "4,00000", "5,00000"), class = "factor"), 
    typeid = c(1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L)), .Names = c("timeid", 
"valid", "typeid"), class = "data.frame", row.names = c("1", 
"2", "3", "4", "5", "6", "7", "8", "9", "10"))

Joris Meys 2010-09-10 14:35:50

ansaurus

tags:

views:

answers:

how to substitute a for loop in R with an optimized function (lapply?)

related questions