From: Martin Duerst Date: 2009-02-25T18:25:34+09:00 Subject: [ruby-core:22461] Re: [Question]utf-8 data contains BOM - by intention or by accident? I think this has been asked earlier, and I can only repeat it here: I don't think there are any plans to build in BOM recognition or removal into file opening. As for forward compatibility for the case that we would add BOM removal later, I suggest that instead of doing something like File.open(..., 'r:utf-8') do |f| f.getc # remove BOM # rest of processing end which won't work anyway for cases where there is no BOM, you do something like File.open(..., 'r:utf-8') do |f| f.ungetc c unless (c = f.getc)=="\uFEFF" # remove BOM, if present # rest of processing end Regards, Martin. At 01:36 09/02/25, Wolfgang N疆asi-Donner wrote: >Hi! > >One short question. When I'm reading an utf-8 encoded file (Windows, Ruby 1.9.1-p0, built using MinGW/MSYS) the Data contains the BOM if present. > >File.open('text.txt','r:utf-8') do |f| > s = f.read > p s.encoding # => # > p s.length # => 7 > p s.bytesize # => 9 > puts s.codepoints.to_a[0].to_s(16) # => feff > p s # => "�ソHelau!" >end > >The BOM is not part of data as namend in Unicode standard, so I mus remove it before processing the data. > >Is this the planned and stable behaviour for Ruby 1.9.1 an future Ruby 1.9 versions? > >I don't have any problems with this, but it should not change later on because it the "remove BOM is there" code parts must be present in nearly every Ruby program. > >Wolfgang N疆asi-Donner > #-#-# Martin J. Du"rst, Assoc. Professor, Aoyama Gakuin University #-#-# http://www.sw.it.aoyama.ac.jp mailto:duerst@it.aoyama.ac.jp