From: Giuseppe Bilotta Date: 2010-09-05T08:10:59+09:00 Subject: [ruby-core:32060] Re: merging nokogiri to ext/ On Sat, Sep 4, 2010 at 11:51 PM, Ryan Davis wrote: > > On Sep 4, 2010, at 10:11 , Yusuke ENDOH wrote: > >> I really, really hate separate releases. �It leads to synchronization >> problem. �See dismal states of Rdoc and Rubygems: >> >> �- [ruby-core:26679] >> �- [ruby-core:31499] >> �- [ruby-core:31503] > > That isn't a problem with separate releases. It is a problem with dual VCs. This is why I don't allow changes to minitest in ruby's repo and insist on patches being sent instead. If rubygems and rdoc did the same thing then we wouldn't have those problems. We'd still have problems, they'd just be easier to manage. This might be a little OT because it's not specific to the nokogiri in/out of stdlib/ext, but I believe there are actually two problems in the current situation. The first problem, as Ryan mentions here, is the "dual VC": the same (sub)tree is tracked in two separate repositories, stand-alone in one case and as part of the main distribution in the other. This means for example that bugs can get fixed in one repo and not in the other. The second problem occurs user-side when the user might want to update a single component of stdlib without installing a new version of ruby altogether (e.g. because that particular part of stdlib had a minor update that fixes a bug that is important to him and no new release is avaiable). Note that this second problem is disjoint from the dual VC one because it would occur even if you had disjoint repositories for the interprenter and for each component of the stdlib. I will start with the second issue by suggesting that a solution would be packaging every single component of stdlib as a gem: each Ruby release would then have the interpreter and a given set of packages at given versions (the stdlib components), and the user would still have the possibility to update single components individually (via gem) whenever a new version comes out. The problem of the dual VC can probably be solved as well, although in this case the way to do it is obviously dependent on the revision control system that is being used. For example, the repository hosting the git scm includes two components (gitk and git-gui) that are _also_ tracked separately in stand-alone repositories. Periodically, the git repository imports any updates from the stand-alone gitk and git-gui projects, and conversely the gitk and git-gui project can merge changes that were only submitted to the git project. Of course, all of this criss-cross updating is possible and easy because git allows these kinds of crazy stuff, but I gather that something similar can also be achieved in SVN, possibly with appropriate use of externals, or even just by exploiting the ease with which single subdirectories can be isolated from a svn repository. In a way, this fits also in-between the "big ruby" and "small ruby" debate: by clearly separating each stdlib component _and_ at the source level (e.g. by periodically merging the corresponding stand-alone repository, and conversely having the stand-alone repository pull in changes from the ruby repo) _and_ at the installation level (everything is a rubygem so it can be updated independently from the main ruby package if necessary) we would have the benefits of being able to include whatever components are deemded necessary in stdlib without forcing any kind of "version coupling" between the interpreter and the stdlib components. It would also make it possible to manage "different sizes of ruby": it would be possible to define a "minimal ruby" that only has the interpreter and a very limited set of selected packages (if any at all), a medium ruby with a decent-size stdlib and a "full ruby" with a bunch of extras, and the only diffference between them would be the (now modular) components included in the packaged stdlib. Do these ideas make sense? -- Giuseppe "Oblomov" Bilotta