From: Brian Candler Date: 2007-05-03T23:04:35+09:00 Subject: Re: Beyond YAML? (scaling) On Thu, May 03, 2007 at 10:50:06PM +0900, Bil Kleb wrote: > I've been using YAML files to store hashes of numbers, e.g., > > { ["O_Kc_01"] => [ 0.01232, 0.01212, 0.03222, ... ], ... } > > This has worked wonderfully for portability and visibility > into the system as I've been creating it. > > Recently, however, I've increased my problem size by orders > of magnitude in both the number of variables and the number > of associated values. The resulting YAML files are prohibitive: > 10s of MBs big and requiring 10s of minutes to dump/load. > > Where should I go from here? Use a SQL database? It all depends what sort of processing you're doing. If you're adding to a dataset (rather than starting with an entirely fresh data set each time), having a database makes sense. If you're doing searches across the data, and/or if the data is larger than the available amount of RAM, then a database makes sense. If you're only touching small subsets of the data at any one time, then a database makes sense. Put it another way, does your processing really require you to read the entire collection of objects into RAM before you can perform any processing? If it does, and your serialisation needs are as simple as it appears above, then maybe something like CSV would be better. O_Kc_01,0.01232,0.01212,0.03222,... If the source of the data is another Ruby program, then Marshal will be much faster than YAML (but unfortunately binary). You could consider using something like Madeleine: http://madeleine.rubyforge.org/ This snapshots your object tree to disk (using Marshal by default I think, but can also use YAML). You can then make incremental changes and occasionally rewrite the snapshot. B.