From: Dmitrii Dimandt Date: 2006-06-26T16:37:23+09:00 Subject: Re: Unicode roadmap? On 6/26/06, Daniel DeLorme wrote: > Logan Capaldo wrote: > > > > Regular expressions are a very powerful tool, but they do not describe > > the entire set of operations one would reasonably want to perform on a > > string. Or perhaps they do but in a needlessly complex way. I want to > > get the first letter (character?) of a sentence, in pure regexp terms > > I'd do this: str.match(/\A./)[0] It's needlessly cryptic. Note that I'm > > not trying to make a commentary on whether or not character string/byte > > string should be separate, just trying to point out that "use regular > > expressions" shouldn't always be the answer. > > It's funny, maybe I'm just dumb but I can't think of a single *real-world* > example where you'd want to access particular characters of a string. Why do you > want the first char? In the context of a byte string there might be something > special at position n (e.g. exif header), but in the context of a human-readable > string what is there? For example, if you want that first char in order to check > if it's a space or not, you should use str =~ /^ /, etc, etc. I honestly can't > think of any real-world examples where regular expressions are less appropriate > than pointer arithmetic. Can you illuminate me with some? Substrings? Finding occurence of a string in a nother string? Why shouldn't str[0..3] work on characters (for a string with encoding set)? Maybe I want to do something like str[0] = Unicode::upcase(str[0])? :) Isn't that what Humane Interface Design (http://www.martinfowler.com/bliki/HumaneInterface.html) is all about ;-) Regular expressions _are_ cryptic. They are powerful, but do I need a sledgehammer when I need a paperclip?