From: Robert Klemme Date: 2010-11-26T02:30:43+09:00 Subject: Re: Ruby 1.8 vs 1.9 On Thu, Nov 25, 2010 at 6:05 PM, Oliver Schad wrote: > Robert Klemme wrote: > >> On Thu, Nov 25, 2010 at 11:12 AM, Phillip Gawlowski >> wrote: >>> On Thu, Nov 25, 2010 at 10:45 AM, Robert Klemme >>> wrote: >>>> >>>> This may be true for the western world but I believe I remember one >>>> of our Japanese friends state that Unicode does not cover all Asian >>>> character sets completely; it could have been a remark about Java's >>>> implementation of Unicode though, I am not 100% sure. >>> >>> Since UTF-8 is a subset of UTF-16, which in turn is a subset of >>> UTF-32, >> >> I tried to find more precise statement about this but did not really >> succeed.  I thought all UTF-x were just different encoding forms of >> the same universe of code points. > > Yes this is correct. Many people don't get the difference between a > charset and the corresponding encoding. Btw, this happens all the time: for example, people often do not grasp the difference between "point in time" and "representation of a point in time in a particular time zone and locale". This becomes an issue if you want to calculate with timestamps at or near the DST change time. :-) > Unicode is a charset not with one encoding but with many encodings. So > we talk about the same characters and different mappings of this > characters to bits and bytes. This mapping is a simple table which you > can write down to a paper sheet and the side with the characters will > always be the same with UTF-8, UTF-16 and UTF-32. > > The encodings UTF-8, UTF-16 and UTF-32 were build for different > purposes. The number after UTF says nothing about the maximum length in > first place it says something about the shortest length (and often about > the usual length if you use this encoding in that situation which it was > build for). More precisely the number indicates the "encoding unit" (see my quote in an earlier posting). One could think up an encoding with encoding unit of 1 octet (8 bits, 1 byte) where the shortest length would be 2 octets. Example 1st octet: number of octets to follow 2nd and subsequent octets: encoded character The shortest length would be 2 octets, but the length would increase by 1 octet so the encoding unit is 1 octet. Cheers robert -- remember.guy do |as, often| as.you_can - without end http://blog.rubybestpractices.com/