From: Stanley Xu Date: 2011-03-23T00:16:49+09:00 Subject: Re: How could I make the Ruby 1.9 string ignore the invalid utf-8 byte sequence in split? --90e6ba6e8af28b1eb8049f13bb43 Content-Type: text/plain; charset=ISO-8859-1 Anyway, let me resend it again. Dear buddies, I am using ruby to run some map reduce job in hadoop streaming. Unfortunately, we have some dirty data which have invalid byte sequence as the input. So while running things like line.chomp.split("\t") I will get errors like :in `split': invalid byte sequence in UTF-8 (ArgumentError) I searched a little bit and try to use iconv to ignore the invalid sequence by if !line.valid_encoding? ic = Iconv.new('UTF-8//IGNORE', 'UTF-8') line = ic.iconv(line) end It resolve most of the invalid lines but will still a couple of line will have the same error. I am wondering if there is a way I could let the string.split() worked in ruby1.9 with invalid character sequences? Thanks in advance Best wishes, Stanley Xu On Tue, Mar 22, 2011 at 11:09 PM, Robert Klemme wrote: > On Tue, Mar 22, 2011 at 3:30 PM, Stanley Xu wrote: > > Sorry, I just mis-sent the half-typed mail by a short-cut in gmail. > > > > I just resent a mail to described the problem. > > Did you? I can't seem to find it. > > Cheers > > robert > > -- > remember.guy do |as, often| as.you_can - without end > http://blog.rubybestpractices.com/ > > --90e6ba6e8af28b1eb8049f13bb43--