From: James Gray Date: 2009-07-24T08:16:36+09:00 Subject: Re: Byte–stream parsing in Ruby On Jul 23, 2009, at 12:02 PM, Brian Candler wrote: > Where does it say that String#<< can now raise an exception, and under > what circumstances? Quoting from the page I linked to in my last message: It's probably worth mentioning that it is possible for a transcoding operation to fail with an error. For example: $ cat transcode.rb # encoding: UTF-8 utf8 = "Résumé…" latin1 = utf8.encode("ISO-8859-1") $ ruby transcode.rb transcode.rb:3:in `encode': "\xE2\x80\xA6" from UTF-8 to ISO-8859-1 (Encoding::UndefinedConversionError) from transcode.rb:3:in `
' Naturally this fails because "…" is not a valid character in Latin-1. > Ah, I finally found it, right at the end of the > *comments* at the bottom of section 8, added a month after initial > publication. (+) What does how long it took me to write the content have to do with anything? I added that comment to cover some items you had mentioned I had overlooked. Now it's invalid because it took me a while??? > Where does it say that the encoding of a String can change when you > concatenate another string onto it? Quoting from the same page: One thing that my help a little in normalizing your data is Ruby's concept of compatibleEncodings. Here's an example of checking and taking advantage of compatible Encodings: # data in two different Encodings p ascii_my # >> "My " puts ascii_my.encoding.name # >> US-ASCII p utf8_resume # >> "Résumé" puts utf8_resume.encoding.name # >> UTF-8 # check compatibility p Encoding.compatible?(ascii_my, utf8_resume) # >> # # combine compatible data my_resume = ascii_my + utf8_resume p my_resume # >> "My Résumé" puts my_resume.encoding.name # >> UTF-8 In this example I had data in two different Encodings, US-ASCII and UTF-8. I asked Ruby if the two pieces of data were compatible?(). Ruby can respond to that question in one of two ways. If it returns false, the data is not compatible and you will probably need to transcode at least one piece of it to work with the other. If an Encoding is returned, the data is compatible and can be concatenated resulting in data with the returned Encoding. You can see how that played out when I combined these Strings. > (+) There is a warning in the string *comparisons* section saying > that, > basically, the rules are too complicated to understand, so you should > always ensure that two strings are in the same encoding before > comparing > them. Arguably you could say the same applies to any other operation > which takes two strings. > > But this to me shows the whole exercise is futile. But you should be doing the exact same thing in Ruby 1.8, which I understand you believe to be a superior system. If you are going to have two pieces of data interact, it just makes sense that they will pretty much always need to be the same kinds of data. > If, in order to write a valid program, you need to ensure that all > strings are in the same encoding, then there should be a global flag > which sets the encoding. Like -E and -U in Ruby 1.9? > And in many apps, the encoding information is carried "out of band" > anyway: for example: in HTTP or MIME, the encoding info is in a > Content-Type: header. Yeah, that's why a global switch won't really save you from doing your job. You need to read that header, and treat the content accordingly. James Edward Gray II