From: dda Date: 2006-06-29T05:03:35+09:00 Subject: Re: Unicode roadmap? Byte arrays – memory blocks, whatever – *do* have their uses, althoughmainly not for string ops. I know I've used memory blocks a lot, forimage processing or other exotic tasks. But never, far as I can tell,for strings. However, byte-level ops can be useful on strings. I cansee two uses: One is for 1-byte encodings. If you know that char==byte, byte-levelops will speed up processing of the strings, since no second guessinghas to be done. Another is because sometimes you have to rip multi-byte chars open andlook at their entrails. Say I want to decompose a hangul syllable intoits primary letters. Unless there is a function provided for that –fat chance considering the lack of interest in Unicode from the BDFL –I'll have to do my own cooking at the byte-level.Example:irb(main):001:0> "한글".length==> 2irb(main):002:0> "かたかな".lengthB==> 12[assuming utf-8 here of course.] I don't really care about a memory block [using this term instead ofbytearray so that I don't get classified in any camp :)], but ifStrings go encodings-aware [hooray], we'll need both types ofoperations... However, I think it is a bit psychotic to base the fundations of animportant feature of the language on the whims and needs[?] of a*small* percentage of the user base. Unicode is an international,*working* standard, whereas this m17n thing has little to show so far,both in terms of production and acceptance [who uses m17n outside afew agitated fellows inside Japan?]. Besides, while some variants of sinograms, aka kanji, and other exoticchars, may not be in the Unicode project *yet* [including the firstsinogram of my wife's given name, which is not to be found in anydictionary listing less than 50,000 sinograms; yeah, blame myfather-in-law...], what's in there for CJKV covers day-to-day needs ofmost people. Seriously, how many times have you seen transcripts ofbone inscriptions on web sites or e-docs? Or arcane kanji pulled outof the Morohashi? Or chu nom chars? Or Jürchen script? Sure, somepeople do work with this stuff. I studied this stuff, and probablywould have liked a way to input/display them. But how many? And howmany use^H^H^H know of Ruby? Let's not lose focus on who's usingwhat... my 0.02€ -- Didier On 6/28/06, Austin Ziegler wrote:> Argue for first-class Unicode support. But you should do so within the> framework which Matz has said he prefers (m17n String and no separate> byte array). Think about API changes that can make this valuable. I> think that Matz *has* settled on the basic data structure, though, and> it's a fight you probably won't win with him. Since, as he pointed out> to Charles Nutter, he's in the percentage of humanity which needs to> deal with non-Unicode more than it needs to deal with Unicode.>> -austin