From: "David A. Black" Date: 2008-12-05T04:09:13+09:00 Subject: Re: reg exp --1926193751-979513573-1228418104=:24059 Content-Type: MULTIPART/MIXED; BOUNDARY="1926193751-979513573-1228418104=:24059" This message is in MIME format. The first part should be readable text, while the remaining parts are likely unreadable without MIME-aware tools. --1926193751-979513573-1228418104=:24059 Content-Type: TEXT/PLAIN; charset=X-UNKNOWN; format=flowed Content-Transfer-Encoding: QUOTED-PRINTABLE Hi -- On Fri, 5 Dec 2008, Jes=FAs Gabriel y Gal=E1n wrote: > On Thu, Dec 4, 2008 at 3:34 AM, Hai anh Le wrote: >> I have a problem with regexp. I have some document like : >> >> A method for detecting a post-translationally modified protein with a >> glycosyl group comprising contacting the protein with a glycosyl >> transferase enzyme and a labeling agent, wherein the labeling agent >> comprises a chemical handle and a transferable glycosyl group. >> >> I want to divide it to some string follow a rule that string start with >> "a, an, the" like : >> >> "A method for detecting " >> "a post-translationally modified protein with " >> "a glycosyl group comprising contacting " >> "the protein with " >> "a glycosyl transferase enzyme and " >> "a labeling agent, wherein " >> "the labeling agent comprises " >> "a chemical handle and " >> "a transferable glycosyl group." > > I though of split, but then you get an array entry for the "separator" > and the following part, so you would need to paste them again > yourself: > > irb(main):008:0> a =3D "A method for detecting a post-translationally > modified protein with a glycosyl group comprising contacting the > protein with a glycosyl transferase enzyme and a labeling agent, > wherein the labeling agent comprises a chemical handle and a > transferable glycosyl group." > =3D> "A method for detecting a post-translationally modified protein > with a glycosyl group comprising contacting the protein with a > glycosyl transferase enzyme and a labeling agent, wherein the labeling > agent comprises a chemical handle and a transferable glycosyl group." > irb(main):011:0> require 'enumerator' > =3D> true > irb(main):012:0> result =3D [] > =3D> [] > irb(main):015:0> a.split(/\b(a|an|the)\b/i)[1..-1].each_slice(2) {|a, > b| result << (a+b)} > =3D> nil > irb(main):016:0> result > =3D> ["A method for detecting ", "a post-translationally modified > protein with ", "a glycosyl group comprising contacting ", "the > protein with ", "a glycosyl transferase enzyme and ", "a labeling > agent, wherein ", "the labeling agent comprises ", "a chemical handle > and ", "a transferable glycosyl group."] > > Maybe someone can come with a split solution easier to join? Try this: a.split(/(?=3Dan?\b|the\b)/i) David --=20 Rails training from David A. Black and Ruby Power and Light: INTRO TO RAILS (Jan 12-15), Fort Lauderdale, FL See http://www.rubypal.com for details Coming in 2009: The Well-Grounded Rubyist (http://manning.com/black2) --1926193751-979513573-1228418104=:24059-- --1926193751-979513573-1228418104=:24059--