From: Robert Klemme Date: 2005-11-09T19:57:12+09:00 Subject: Re: Question about symbols Matthew Smillie wrote: > I wonder if I could trouble the list a bit further on this one: I've > got a collection of newswire articles, and I was thinking of using > symbols to represent the words in each. I'd likely only ever be > using a half-dozen articles at any one time, but I'd be using on the > order of tens of thousands of articles over the course of the > program. My hunch is that to avoid stuffing memory full of unused > symbols, I should deal with each subset of articles in a forked > process - I'm assuming that since threads run in the interpreter, > they'll share the same symbol ids as the parent. I would not use Symbols for that and I'd also not change the application architecture (i.e. using fork) because forking essentially would try to fix a problem introduced by using Symbols. Also, this smells a bit like premature optimization. Do you actually need to have those words separately? What exactly are you doing with your articles? > This raises a couple of quick questions: firstly, how big is a > "limited set", in a hand-waving, give or take an order-of-magnitude > sort of estimate? IMHO typically these are only few per use case. > Secondly, I note that the FAQ says that fork "is slow". Am I right > to think that's simply the overhead of starting up the thread, and > then things run smoothly thereafter, or is there some other penalty > of which I'm not aware? fork doesn't start a thread but it creates a new process. This means it increases resource consumption on the system (i.e. there's a slight management overhead for the additional process and you need more memory - even with copy on write). Kind regards robert