From: Robert Klemme Date: 2009-10-20T00:10:06+09:00 Subject: Re: deep cloning, how? On 10/19/2009 10:32 AM, Brian Candler wrote: > Robert Klemme wrote: >> 2009/10/16 lith : >>>> But I think the spirit of dup >>>> described above is that dup defined in a subclass should initialize it >>>> using its constructor. >> Brian, I disagree. The proper way is to implement #initialize_copy. >> That way you can make sure you do not get aliasing effects even if >> source and copy are frozen because in #initialize_copy frozen state is >> not applied. > > I don't understand what you mean by that. If #dup calls self.class.new > then you obviously get a new and hence unfrozen object. > > It is certainly true that the *default* implementation of both #dup and > #clone (defined in Object) calls initialize_copy. A generic #dup must > behave this way; it doesn't know what the new() method arguments are in > any particular subclass of Object. I don't think this should be taken as > necessarily implying that you are expected to leave #dup alone in your > own classes, and only override #initialize_copy instead. > > The way I read the documentation implies to me that #dup in user defined > classes *should* call new. Silly example: > > class NewsReader > def initialize(url, state_filename) > @url = url > @http_client = HTTPClient.new(@url) > @state_filename = state_filename > @state_file = File.open(@state_filename) > end > def dup > self.class.new(@url, @state_filename.dup) > end > end > > Here the logic of how to build a NewsReader, including building all the > associated helper objects, is built into the #initialize method. Brian, the approach shown above does not work well with subclasses. The code attempts to be safe with regard to inheritance (by doing self.class.new instead of NewsReader.new) but it will fail miserably as soon as a sub class constructor has a different argument list (which is not too uncommon). I completely agree with Rick here: the comment in Object#dup is probably outdated. The most reasonable way to customize object cloning *and* dupping is to implement #initialize_copy in a way to at least ensure no aliasing of unfrozen members takes place. > I don't > think you would want to duplicate all this logic in #initialize_copy. You would not duplicate the logic from #initialize in #initialize_copy because #initialize_copy does a completely different job: it copies state of an instance which is known to be consistent and just needs to ensure that aliasing of object references does not break your class invariants later accidentally. This is the reason why in #initialize_copy different logic should be applied - even for shallow copies! Method #initialize OTOH needs to work with its arguments which were provided from the outside (outside of this class that is) and may not meet expectations or valid ranges. > Furthermore, I think I would expect #clone only to copy the top object, > and leave all the instance variables aliased. As far as I can see both #clone and #dup are meant to do shallow copies but I may be wrong here. At least this is what the contract ob Object promises and I tend to be cautious about changing such things. Even if you redefine semantics to being deep copy for certain classes then implementing it in #initialize_copy is superior to other approaches IMHO. > Obviously there are no hard-and-fast rules here, and with Ruby there are > many ways to achieve the same goal. That's true. But I would say at least when considering inheritance some ways are better than others. In fact I have been doing self.class.new most of the time in #dup because I completely forgot about #initialize_copy. But I will certainly change that habit from now on. > I'd certainly agree this is an area where Ruby's documentation falls > short. Right. > Taking another example: I don't think you'll disagree that 99% of the > time you are expected to leave Object.new alone and instead define > #initialize in your own classes. But you wouldn't find that out from the > documentation: > > $ ri Object.new > ------------------------------------------------------------ Object::new > Object::new() > ------------------------------------------------------------------------ > Not documented > > $ ri Object#initialize > Nothing known about Object#initialize Funny that you mention it: #new and #initialize on one side and #dup / #clone and #initialize_copy on the other side have one thing in common: object allocation is separated from initialization. I believe this was a wise decision because that way allocation policies can be implemented easier than in languages like C++ and Java where both are inseparable. For example, you can add your own #deep_dup to the language: class Object def deep_dup cp = self.class.allocate instance_variables.each do |var| cp.instance_variable_set(instance_variable_get(var)) end cp.initialize_deep_copy(self) cp end def initialize_deep_copy(source) # nothing to do here end end class String def initialize_deep_copy(source) replace source end end # note this implementation is not robust against # circles in the object graph! class Array def initialize_deep_copy(source) source.each do |y| self << y.deep_dup end end end a = %w{foo bar baz} b = a.dup b[2].replace "CHANGED" p a, b a = %w{foo bar baz} b = a.deep_dup b[2].replace "CHANGED" p a, b Kind regards robert -- remember.guy do |as, often| as.you_can - without end http://blog.rubybestpractices.com/