From: Michal Suchanek Date: 2006-06-15T22:11:29+09:00 Subject: Re: Unicode roadmap? On 6/15/06, Julian 'Julik' Tarkhanov wrote: > > On 15-jun-2006, at 13:21, Michal Suchanek wrote: > > > On 6/15/06, Julian 'Julik' Tarkhanov wrote: > >> > > > >> 5. Preferably separate (and strictly purposed) Bytestring that you > >> get out of Sockets and use in Servers etc. - or the ability to > >> "force" all strings recieved from external resources to be flagged > >> uniformly as being of a certain encoding in _your_ program, not > >> somewhere in someone's library. If flags have to be set by libraries, > >> they won't be set because most developers sadly don't care: > >> > >> http://www.zackvision.com/weblog/2005/11/mt-unicode-mysql.html > >> http://thraxil.org/users/anders/posts/2005/11/01/unicodification/ > > > > Where else should the strings be flagged? > They should nog be flagged, because some strings will be flagged and > some won't and exactly > in the wrong places at the wrong time. See _is_uf_8_ in Perl to > witness the terrible ugliness of this. You can certainly get the things wrong. But if you get a string that is wrongly flagged you have the choice to fix the code where the string originates or work arond it by flagging it right. If you have a code that gets the encoding wrong, and it tries to convert the string to some 'universal' encoding you want to use everywhere in your application, you get a broken string. > > > If you get a web page > > through http request, and the library parses the response for you, it > > should set enconding on the web page. You would never know since you > > only received the page, not the header. > > That's why you should distinguish between a ByteArray and a String. How does it help you here? > > > > >> > >> 6. Unicode-aware strip dealing with weirdo whitespaces (hair space, > >> thin space etc.) > >> 7. And no, as I mentioned - it doesn't handle it properly because > >> the /i modifier is broken, and to deal without it you need to > >> downcase BOTH the regexp and the string itself. Closed circle - you > >> go and get the Unicode gem with tables. > >> > >> All of this can be controlled either per String (then 99 out of 100 > >> libraries I use will be getting it wrong - see above) or by a global > >> setting such as $KCODE. > > > > I do not see why libraries should be always wrong. After all, you can > > always fix them. And setting the encoding globally is a bad thing. You > > cannot have strings encoded in different encodings in one process > > then. It looks quite limiting. For one, the web pages that you get > > from various servers (and even the same server) can be in varoius > > encodings. > > Of course they can (and will). When I have to approach this I usually > just snif the encoding of the strings I recieved and then feed them > to iconv and friends before doing any processing. A library that > downloads stuff off the Internet should be (IMO) aware of > the charset madness and decode the strings for me. If it can decode them, it can flag them. It has to be aware - that's it. > > Trust me, when multibyte/Unicode handling is optional, 80% of > libraries do it wrong. Re-read the links above if you don't believe. But they get the very foundation wrong. In Python functions that take multiple strings can only thake them in one encoding. It is impossible to concatenate differently encoded strings. Of course, this is bound to fail. In the other case they use a database with poor support for unicode, and mysql that does exactly the same thing ruby does right now - works with strings as arrays of bytes. Of course, this is going to break. Neither is the case when the strings carry information about their encoding, and the string functions can handle strings encoded differently. The fact that there are libraries and languages with poor unicode support does not mean it must be always poor. Thanks Michal