ansaurus

Question

python: what are efficient techniques to deal with deeply nested data in a flexible manner?

Answer 1

+1 A:

You can make your loops look better, by replacing:

for mk in dic.keys():
    for rgk in dic[mk].keys():
        for nk in dic[mk][rgk].keys():
            for ik in dic[mk][rgk][nk].keys():
                for ek in dic[mk][rgk][nk][ik].keys():
                    #do something

with

for mv in dic.values():
    for rgv in mv.values():
        for nv in rgv.values():
            for iv in nv.values():
                for ev in iv.values():
                    #do something

You thus gain access to all the values with a relatively terse code. If you also need some keys, you can do something like:

for (mk, mv) in dic.items():
    # etc.

Depending on your needs, you might also consider creating and then using a single dictionary with tuple keys:

dic[(mk, rgk, nv, ik, ek)]

EOL 2010-03-30 09:45:17

Iterating a dictionary object returns the dictionary's keys. You can't then iterate the key and expect to be accessing the value.

MattH 2010-03-30 09:55:15

I tried the first block, but at for rgv in mv: I see that rgv iterates through the mv string, instead of the mv.keys()?

AlexandreS 2010-03-30 11:07:59

@MattH: you're right, my bad! Fixed. Thanks!

EOL 2010-03-30 20:33:10

Answer 2

A:

You ask: How should I organize the data I'm analyzing, and which tools should I use to manage it?

I suspect that a dictionary, for all its optimization, is not the right answer to that question. I think you'd be better off using XML or, if there is a Python binding for it, HDF5, even NetCDF. Or, as you suggest yourself, a database.

If your project is of sufficient duration and usefulness to warrant learning how to use such technologies, then I think you'll find that learning them now and getting the data structures right is a better path to success than wrestling with the wrong data structures for the entire project. Learning XML, or HDF5, or SQL, or whatever you choose, is building up your general expertise and making you better able to tackle the next project. Sticking with awkward, problem-specific and idiosyncratic data structures leads to the same set of problems next time around.

High Performance Mark 2010-03-30 10:08:35

Thanks for letting me know about HDF5 and NetCDF. Python has bindings for both, and the licenses are usable. I think I will start by reading about HDF5, and if it looks promising I will study how to make use of it with my data.These kinds of pointers are exactly what I was hoping for when I wrote my question.

AlexandreS 2010-03-30 11:40:08

Answer 3

A:

You could write a generator function that allows you to iterate over all elements of a certain level:

def elementsAt(dic, level):
    if not hasattr(dic, 'itervalues'):
        return
    for element in dic.itervalues():
        if level == 0:
            yield element
        else:
            for subelement in elementsAt(element, level - 1):
                yield subelement

Which can then be used as following:

for element in elementsAt(dic, 4):
    # Do something with element

If you also need to filter elements, you could first get all elements that need to be filtered (say, the 'rgk' level):

for rgk in getElementsAt(dic, 1):
    if isValid(rgk):
        for ek in getElementsAt(rgk, 2):
            # Do something with ek

At least that'll make it a little easier to work with a hierarchy of dictionaries. Using more descriptive names would help, too.

Pieter Witvoet 2010-03-30 10:24:17

thanks for showing me the use of dic.itervalues(), which I had not appreciated so far. The problem is that it does not really address the issue of deeply nested and inflexible for chains.

AlexandreS 2010-03-30 11:26:58

Answer 4

+6 A:

"I stored it in a deeply nested dictionary"

And, as you've seen, it doesn't work out well.

What's the alternative?

Composite keys and a shallow dictionary. You have an 8-part key: ( individual, imaging session, Region imaged, timestamp of file, properties of file, regions of interest in image, format of data, channel of acquisition ) which maps to an array of values.
```
{ ('AS091209M02', '100113', 'R16', '1263399103', 'Responses', 'N01', 'Sequential', 'Ch1' ): array, 
...
```
The issue with this is search.
Proper class structures. Actually, a full-up Class definition may be overkill.

"The type of operations I perform is for instance to compute properties of the arrays (listed under Ch1, Ch2), pick up arrays to make a new collection, for instance analyze responses of N01 from region 16 (R16) of a given individual at different time points, etc."

Recommendation

First, use a namedtuple for your ultimate object.

Array = namedtuple( 'Array', 'individual, session, region, timestamp, properties, roi, format, channel, data' )

Or something like that. Build a simple list of these named tuple objects. You can then simply iterate over them.

Second, use many simple map-reduce operations on this master list of the array objects.

Filtering:

for a in theMasterArrrayList:
    if a.region = 'R16' and interest = 'N01':
        # do something on these items only.

Reducing by Common Key:

individual_dict = defaultdict(list)
for a in theMasterArrayList:
    individual_dict[ a.individual ].append( a )

This will create a subset in the map that has exactly the items you want.

You can then do indiidual_dict['AS091209M02'] and have all of their data. You can do this for any (or all) of the available keys.

region_dict = defaultdict(list)
for a in theMasterArrayList:
    region_dict[ a.region ].append( a )

This does not copy any data. It's fast and relatively compact in memory.

Mapping (or transforming) the array:

for a in theMasterArrayList:
    someTransformationFunction( a.data )

If the array is itself a list, you're can update that list without breaking the tuple as a whole. If you need to create a new array from an existing array, you're creating a new tuple. There's nothing wrong with this, but it is a new tuple. You wind up with programs like this.

def region_filter( array_list, region_set ):
    for a in array_list:
        if a.region in region_set:
            yield a

def array_map( array_list, someConstant ):
    for a in array_list:
        yield Array( *(a[:8] + (someTranformation( a.data, someConstant ),) )

def some_result( array_list, region, someConstant ):
    for a in array_map( region_filter( array_list, region ), someConstant ):
        yield a

You can build up transformations, reductions, mappings into more elaborate things.

The most important thing is creating only the dictionaries you need from the master list so you don't do any more filtering than is minimally necessary.

BTW. This can be mapped to a relational database trivially. It will be slower, but you can have multiple concurrent update operations. Except for multiple concurrent updates, a relational database doesn't offer any features above this.

S.Lott 2010-03-30 10:26:35

This is very helpful. I had already resorted to flattening the data dictionary or a (subsets of it) + list comprehensions to filter the data, but your description of the use of named tuple enables a much cleaner syntax and code than what I have. Plus it involves minimal learning to get the desired results, so I think I'll go for that, at least for the moment.

AlexandreS 2010-03-30 11:20:41

Answer 5

A:

I will share some thoughts about this. Instead of this function:

for mk in dic.keys():
    for rgk in dic[mk].keys():
        for nk in dic[mk][rgk].keys():
            for ik in dic[mk][rgk][nk].keys():
                for ek in dic[mk][rgk][nk][ik].keys():
                    #do something

Which you would like to simply write as:

for ek in deep_loop(dic):
    do_something

There are 2 ways. One is functional, second is generator-like. The second one is:

def deep_loop(dic):
    for mk in dic.keys():
        for rgk in dic[mk].keys():
            for nk in dic[mk][rgk].keys():
                for ik in dic[mk][rgk][nk].keys():
                    for ek in dic[mk][rgk][nk][ik].keys():
                        yield ek

This allows you to capture the logic of going through the dictionary. It is very easy to modify this function to support different ways of going through the structure. It depends on the way your structure changes, if it is just a depth of the loop or something different. Could you post some more advanced examples on what requirements on going through the tree you have? Like filtering, search etc.? The depth would look like this (untested) - it will yield a pair of (tuple of keys), (value):

def deep_loop(dic, depth):
    if depth == 0:
        yield (), dic
    for subkey, subval in dic.items():
        for ktuple, value in deep_loop(subval, depth-1):
            yield (subkey,)+ktuple, value

Now it gets easier:

for (k1,k2,k3,k4), value in deep_loop(dic, 4):
    # do something

There are other ways to customize this, you could add a named tuple type as a parameter of deep_loop. Deep_loop could autodetect the depth from the named tuple and return the named tuple.

ondra 2010-03-30 10:58:35

I posted additional examples of what I managed to implement as an addendum, hope that helps. Your first solution packs the search process in a function, but does not solve most of my problems, as searches are often variable, and I would have to define many functions to describe each search. Your second option addresses this point, but uses recursion, which frightens me a bit after the bad experiences I had with it (although I managed to write some working recursive functions). But thanks for showing me how you coupled dic.items() with generator expressions, that's useful to me.

AlexandreS 2010-03-30 12:35:46

ansaurus

tags:

views:

answers:

python: what are efficient techniques to deal with deeply nested data in a flexible manner?

related questions