From: Robert Dober Date: 2009-07-22T17:55:40+09:00 Subject: Re: String spliting and inclusion On 7/22/09, Robert Dober wrote: > On 7/22/09, David A. Black wrote: > > > Hi -- > > > > > > On Wed, 22 Jul 2009, Stuart Clarke wrote: > > > > > > > Hi all, > > > > > > I am having trouble working out some logic for my problem. I basically > > > have a long string (320 characters) and I want to split into smaller > > > strings no longer than 50 characters in length. At present I have the > > > following regex: > > > > > > data = "big long string" > > > > > > puts data.scan(/{50}/) > > > > > > This nicely breaks up the string however there are a few problems with > > > it, including: > > > > > > It only outputs 50 character chunks, therefore when it gets to the end > > > and only 20 characters remain it misses them off the output (it outputs > > > 6 50 characters strings and ignores the remaining 20) > > > > > > This regex also splits up words, which is something I don't want. I want > > > a script to count to 50 and when it gets there, go backwards to find > > > some white space and split it at that point, therefore not breaking up a > > > word. As a result a number of sub strings of various sizes will be > > > created all less than 50 chars. > > > > > > I hope this makes sense, to summarise I want to break up a string into a > > > max of 50 characters without breaking up words. > > > > > > > Try this. I don't guarantee robustness. > > > > str.scan(/\b.{0,50}(?:$|\b)/m) > > Hmm my \b at the end of my solution might have been a problem in some > edge cases, however I would suggest the usage of \z instead of $ and > the m switch. I fail to see why you put a \b at the beginning David, > would you mind to explain? > > In Ruby 1.9 (or Oniguruma that is) the negative lookahead assertion > might lead to the most elegant solution: > > /.{,50}(?!\B)/ Nahh that leaves us with spaces at the beginning of the line, of course we could do scan(...).map( &:lstrip ) but that hurts my regex pride ;) This seems to work (but does not really): s = "Some words are made of letters! Some are not!" puts s.scan( /.{,10}\p{Graph}(?:\P{Graph}|\z)/ ) Replace the puts with p and you will see trailing whitespace now :(. This is a little bastard of a problem indeed. Simplest I could come up with so far: s = "Some words are made of letters! Some are not!" p s.scan( /.{,10}\p{Graph}(?:\P{Graph}|\z)/ ).map( &:strip ) HTH Robert > > BTW it seems that {n,m} does not have a "non greedy" and "possessive" > variant, or did I miss it? Yes I did, they are there {n,m}? and {n,m}+, sorry. > > Cheers > Robert > > -- Toutes les grandes personnes ont d’abord été des enfants, mais peu d’entre elles s’en souviennent. All adults have been children first, but not many remember. [Antoine de Saint-Exupéry]