From: Phillip Gawlowski Date: 2010-11-25T00:47:00+09:00 Subject: Re: Ruby 1.8 vs 1.9 On Wed, Nov 24, 2010 at 4:15 PM, Brian Candler wrote: > For example, say I have s1 tagged UTF-8 and s2 tagged ISO-8859-1. If > "s1+s2" always raised an exception, it would be easy to find, and easy > to fix. > > However the 'compatibility' rules mean that this is data-sensitive. In > many cases s1+s2 will work, if either s1 contains non-ASCII characters > but s2 doesn't, or vice-versa. It's really hard to get test coverage of > all the possible cases - rcov won't help you - or you just cross your > fingers and hope. Convert your strings to UTF-8 at all times, and you are done. You have to check for data integrity anyway, so you can do that in one go. >> I hear nobody complain that 1 / 0 raises but 1.0 / 0.0 gives Infinity, > > Well, IEEE floating point is a well-established standard that has been > around for donkeys years, so I think it's reasonable to follow it. Every natural number is an element of the set of rational numbers. For all intents and purposes, 0 == 0.0 in mathematics (unless you limit the set of numbers you are working on to natural numbers only, and let's just ignore irrational numbers for now). And since the 0 is around for a bit longer than the IEEE, and the rules of math are taught in elementary school (including "you must not and cannot divide by zero"), Ruby exhibits inconsistent behavior for pretty much anyone who has a little education in maths. The IEEE standards deal with representing floating point numbers in an inherently integer-based numerical system, but they don't supersede the rules of maths. Ruby's behavior of returning *infinity* is the proverbial icing on the cake, since dividing something large by something infinitely small results in something large (so, x / 0.000000...[ad infinitum]...1 = x ; a trick used in integrals, too). Thus, you have to exercise due diligence in this area if you want to keep your results in the sphere of what's possible and sane. > And yes, if I see code like "c = a / b", I do think to myself "what if b > is zero?" It's easy to decide if it's expected, and whether I need to do > something other than the default behaviour. Then I move onto the next > line. It's easy? Take a look at integrals, and infinitesimal[0] numbers. Infinitesimal are at the same time zero and *not* zero. > For "s3 = s1 + s2" in 1.9 I need to think to myself: "what if s1 has a > different encoding to s2, and s1 is not empty or s2 is not empty and > s1's encoding is not ASCII-compatible or s2's encoding is not > ASCII-compatible or s1 contains non-ASCII characters or s2 contains > non-ASCII characters? And what does that give as the encoding for s3 in > all those possible cases?" And then I have to carry the possible > encodings for s3 forward to the next point where it is used. Then, as I suggested above, enforce a standard encoding in your code. Convert everything into UTF-8, and you are pretty much done. [0] http://en.wikipedia.org/wiki/Infinitesimal -- Phillip Gawlowski Though the folk I have met, (Ah, how soon!) they forget When I've moved on to some other place, There may be one or two, When I've played and passed through, Who'll remember my song or my face.