From: matz@... (Yukihiro Matsumoto) Date: 2002-08-01T14:07:29+09:00 Subject: Re: Unicode in Ruby now? Hi, In message "Re: Unicode in Ruby now?" on 02/08/01, Clifford Heath writes: |Is Mojikyo a superset of Unicode? If not, how hard is the translation to |UCS-4? Mojikyo character set contains all the CJK characters in ISO 10646. But codepoint number for each character is different, so that we can only do Unicode -> Mojikyo conversion using table lookup. The translation to UCS-4 from Mojikyo is nearly impossible. (How can one assign non-Unicode codepoint number in UCS-4?) http://www.mojikyo.org/ |I designed the UCS-4 string class we use here in C++, with a UTF-8 storage |format (up to 31-bit with a six-byte UTF-8 sequence). The string class |remembers which character you last accessed and at what byte offset it |started, so that when you ask for another character, it can decide whether |heuristically to search forward from the start, forward or backward from |the remembered point (most common), or if it has ever counted the |characters, backward from the end. This minimises the search cost since |most string processing is largely sequential. Remembered point technique is very interesting. When we meet performance problem with Ruby I18N, I will use it. matz.