From: Tim Bray Date: 2006-06-18T13:44:08+09:00 Subject: Re: Unicode roadmap? On Jun 17, 2006, at 10:34 AM, Stefan Lang wrote: > Or do you mean Ruby should determine the encoding > automatically? IMO, that would be bad magic and error-prone. Not possible in the general case. There are a few data formats including XML and ASN.1, which make it possible to reliably infer the encoding from the instance, but a lot of Web processing these days is best-guess, and often fails. > How does the regular expression engine work then? The two sane options are (a) have a fixed encoding for Strings and compile the regex in such a way that it runs directly on the encoding. This has been done for both UTF-8 and UTF-16 and is insanely efficient, but it locks you into the fixed encoding. (b) have an iterator which produces abstract characters from whatever encoding is in use and run the regex over the characters, not the bytes of the representation. The implementation is trickier and performance is an issue, but you're not locked to an encoding. -Tim