From: Shot - Piotr Szotkowski Date: 2006-01-14T01:37:47+09:00 Subject: Re: Ruby, Unicode - ever? Hello. Austin Ziegler: > On 13/01/06, Shot - Piotr Szotkowski wrote: >> Do what other languages do, ask the underlying operating >> system for LC_COLLATE (or ask it to sort the data by itself). > Which is not really appropriate for all operating > systems, and is one of the *dumbest* things about POSIX. I'm not a language developer, but a programmer (doing a lot of i18n and l10n work lately) and user, and to me using LC_COLLATE when available seems much better than defaulting to binary collation 'just because'. BTW: PostgreSQL defaults to system's LC_COLLATE if not told explicitely on cluster init as well. >> No, Ruby could implement a counterpart of MySQL's >> utf8_general_ci and/or utf8_unicode_ci collations: >> http://dev.mysql.com/doc/refman/5.0/en/charset-unicode-sets.html > Given how much else MySQL gets wrong, why should > this be trusted to be anything close to right? Have you actually followed the link? Quoting: The utf8_unicode_ci collation is implemented according to the Unicode Collation Algorithm (UCA) described at http://www.unicode.org/reports/tr10/ The collation uses the version-4.0.0 UCA weight keys: http://www.unicode.org/Public/UCA/4.0.0/allkeys-4.0.0.txt I won't start discussing MySQL's quality[1], but I don't see a reason why Unicode's own collation algorithm shouldn't at least be considered. [1] Question of the Week: Who in their right mind documents 'CREATE USER' in a shell for version 4.1, but does not actually introduce the command until 5.0? Cheers, -- Shot -- To a database person, every nail looks like a thumb. Or something like that. -- Jamie Zawinski ====================== http://shot.pl/hovercraft/ === http://shot.pl/1/125/ ===