ansaurus

Question

Don't more than a few dozen partitions make sense?

Answer 1

+2 A:

It's probably unwise to have that many partitions, yes. The main reason to have partitions at all is not to make indexed queries faster (which they are not, for the most part), but to improve performance for queries that have to sequentially scan the table based on constraints that can be proved to not hold for some of the partitions; and to improve maintenance operations (like vacuum, or deleting large batches of old data which can be achieved by truncating a partition in certain setups, and such).

Maybe instead of using ranges of simulation_id (which means you need more and more partitions all the time), you could partition using a hash of it. That way all partitions grow at a similar rate, and there's a fixed number of partitions.

The problem with too many partitions is that the system is not prepared to deal with locking too many objects, for example. Maybe 200 work fine, but it won't scale well when you reach a thousand and beyond (which doesn't sound that unlikely given your description).

There's no problem with having billions of rows per partition.

All that said, there are obviously particular concerns that apply to each scenario. It all depends on the queries you're going to run, and what you plan to do with the data long-term (i.e. are you going to keep it all, archive it, delete the oldest, ...?)

alvherre 2010-08-18 18:17:12

thnx so much alvherre. I should keep all the simulation results to query various historical statistics. And the reason i used ranges of simulation_id for partition is that i guessed it would be good to save results of adjacent simulations together in one partition bcz i usually query group of adjacent simulations together. Anyway, you helped me to release all the concerns. I'll follow the way to use hash partition with limited number of partitions.

tk 2010-08-18 18:44:05

If you are going to query adjacent simulations together, then it's probably a good idea to choose a mapping that puts several adjacents simulations in the same partition. (So don't use plain modulo arithmetic for the hash). On the other hand, make sure you use constraints that can be proved true or false for each partition for whatever query you're going to use, so that partitions that don't contain any simulation in your result set can be quickly discarded (as with range partitioning).

alvherre 2010-08-18 19:50:18

alvherre, it seems that hash partition is not supported yet. http://wiki.postgresql.org/wiki/Table_partitioningDo you know how to work around for hash partitioning? Thnx.

tk 2010-08-23 21:26:15

ansaurus

tags:

views:

answers:

Don't more than a few dozen partitions make sense?

related questions