From: Martin Duerst Date: 2008-10-24T15:52:17+09:00 Subject: [ruby-core:19471] Re: Default source encoding (Was: [Bug #680] csv.rb: CSV.parse is toolate when encoding is mismatch) A default for the source encoding has been discussed quite a long time ago (in some Japanese meetings or on ruby-dev, I don't remember), and the conclusion was that the source encoding has to be given (with a majic comment) in the file itself (unless the file is all ascii). The reason for this is that the source encoding is a property of the source, and nothing else. On very simple scripts, it might occasionally be slightly easier if it were the same as default_external or default_internal, but this is only the case as long as you stay in exactly the same environment, and don't move the script. But scripts grow and move, so it's better to get the settings right at the start. However, as far as I remember, the idea was that for -e, default_external should be used, because that's what one is using in a shell. I'm not sure why this doesn't work below. (assuming Takeyuki is working in a Shift_JIS environment, which isn't completely sure). Regards, Martin. At 12:12 08/10/24, Michael Selig wrote: >Hi, > >This bug actually brings up an interesting issue - should the source >encoding default to something other than UTF-8 (ie: if it is not specified >in the "magic comment")? > >Perhaps it should default to the encoding specified by the user's locale? >Or perhaps it should default to the value of "default_internal" if it is >set? Or even default_external? > >I suggest that it should default to "default_internal" if that is set, and >then to the locale encoding if not. > >What do others think? >Having it default to the locale in this case would probably avoid the >encoding mismatch entirely (and the resulting confusion). > >Cheers >Mike > >On Fri, 24 Oct 2008 11:58:33 +1100, Takeyuki Fujioka > wrote: > >> Bug #680: csv.rb: CSV.parse is too late when encoding is mismatch >> http://redmine.ruby-lang.org/issues/show/680 >> >> Author: Takeyuki Fujioka >> Status: Open, Priority: Normal >> Category: lib, Target version: 1.9.x >> >> I think this result is true, but encoding mismatch raise is too late. >> >> see: >> % time ruby19 -rcsv -e >> 'CSV.parse(("\x82\xA0,\x82\xA2\n"*10000).force_encoding("shift_jis"))' >> ruby19 -rcsv -e 0.30s user 0.02s system 96% cpu 0.330 total >> >> % time ruby19 -rcsv -e 'CSV.parse(("\x82\xA0,\x82\xA2\n"*10000))' >> /Users/fujioka/local/lib/ruby/1.9.0/csv.rb:1981:in `=~': broken UTF-8 >> string (ArgumentError) >> from /Users/fujioka/local/lib/ruby/1.9.0/csv.rb:1981:in >> `init_separators' >> from /Users/fujioka/local/lib/ruby/1.9.0/csv.rb:1563:in `initialize' >> from /Users/fujioka/local/lib/ruby/1.9.0/csv.rb:1350:in `new' >> from /Users/fujioka/local/lib/ruby/1.9.0/csv.rb:1350:in `parse' >> from -e:1:in `
' >> ruby19 -rcsv -e 'CSV.parse(("\x82\xA0,\x82\xA2\n"*10000))' 1.55s user >> 2.57s system 90% cpu 4.530 total >> >> >> ---------------------------------------- >> http://redmine.ruby-lang.org > > > #-#-# Martin J. Du"rst, Assoc. Professor, Aoyama Gakuin University #-#-# http://www.sw.it.aoyama.ac.jp mailto:duerst@it.aoyama.ac.jp