From: "Mauricio Fernández" Date: 2003-06-12T15:43:43+09:00 Subject: Re: HTML -> list of sentences? (semi-impossible task) --WIyZ46R2i8wDzkSu Content-Type: text/plain; charset=us-ascii Content-Disposition: inline On Thu, Jun 12, 2003 at 11:38:13AM +0900, Hal E. Fulton wrote: > Here's an idea I'm toying with. Suggestions > are welcome. > > I want to take an HTML document (reasonably > well-formed, but not guaranteed) and remove > all the tags from it... > > ...and get a list of the *sentences* in the > document. Attached is a 30 mins. hack of mine that does something like that. The scanning part is really a kludge but I've been using it w/ acceptable results in a proxy that add hints to webpages on the fly. > There are, of course, several things that make > this difficult: > - need to distinguish between end-of-sentence > and embedded punctuation, including both > abbreviations and textual references to > Ruby methods such as eof? and split! They only way I can think of to do that is having a list of methods and abbreviations to ignore. > - need to treat sentence fragments as sentences > - need to ignore blocks of code Kind of doable if you have a dictionary (/usr/share/dict/words should be enough). For each candidate sentence, you see how many words are there and take it if the percentage is above some threshold. > - etc. > > My current approach is to start with htmlsplit > from the RAA. This is fairly simplistic, but > at least it doesn't have any dependencies. > > Not sure whether to do it in two steps or not: > 1. Convert to text > 2. Process > > Might be just as easy to do it in one step if > I knew what I was doing. IMHO it can be done in one pass. ---- Given http://www.rubygarden.org/ruby?ClassMethodsTutorial, my hack returns ["This", "is", "simply", "an", "extract", "from", "a", "post", "to", "ruby-talk", "by", "DavidBlack", "on", "the", "topic", "of", "class", "methods."] ["It", "is", "stored", "here", "in", "the", "hope", "that", "it", "will", "be", "useful!"] ["It", "actually", "goes", "beyond", "the", "surface", "of", "class", "methods", "to", "describe", "the", "nature", "of", "classes", "as", "objects", "so", "is", "interesting", "reading", "for", "anyone", "progressing", "to", "intermediate-level", "Ruby."] ["See", "also", "ClassMethods", "for", "an", "overview", "of", "the", "options", "available", "in", "Ruby", "for", "creating", "class", "methods", "and", "SingletonTutorial", "for", "a", "detailed", "explanation", "of", "singleton", "methods."] ["Every", "object", "responds", "to", "certain", "messages", "i.e.", "can", "call", "methods", "with", "certain", "names", "."] ["Usually", "those", "methods", "are", "the", "instance", "methods", "defined", "by", "the", "object's", "class", "However", "it's", "also", "possible", "to", "add", "methods", "to", "individual", "objects", "Now", "c", "will", "respond", "to", "speak", "--", "but", "other", "instances", "of", "class", "C", "will", "not", "This", "means", "that", "speak", "is", "a", "singleton", "method", "of", "c."] ["now", "look", "at", "this", "Notice", "the", "similarity", "between", "the", "syntax", "involved", "in", "creating", "a", "new", "singleton", "method", "for", "c", "and", "creating", "a", "class", "method", "of", "class", "D", "In", "fact", "these", "are", "essentially", "the", "same", "thing."] ["In", "both", "cases", "what's", "happening", "is", "that", "a", "singleton", "method", "is", "being", "added", "to", "a", "particular", "object."] ["It", "just", "happens", "to", "be", "that", "in", "the", "second", "case", "the", "object", "getting", "the", "new", "method", "is", "a", "Class", "object", "as", "opposed", "to", "a", "String", "an", "Array", "an", "instance", "of", "MyClass", "?"] ["So", "now", "D", "responds", "to", "greet", "just", "as", "c", "responds", "to", "speak", "."] ["In", "other", "words", "the", "term", "class", "method", "is", "just", "a", "special", "term", "for", "something", "which", "you", "can", "do", "with", "any", "mutable", "object", "namely", "add", "a", "singleton", "method", "to", "it."] ["It", "has", "a", "special", "name", "because", "in", "actual", "program", "design", "class", "methods", "have", "a", "special", "role", "to", "play."] ["But", "what", "they", "are", "at", "heart", "is", "singleton", "methods", "defined", "on", "objects", "where", "those", "objects", "happen", "to", "be", "instances", "of", "a", "class", "called", "Class."] ["The", "use", "of", "uppercase", "names", "constants", "for", "classes", "can", "obscure", "the", "fact", "that", "classes", "are", "just", "objects."] ["Also", "the", "usual", "style", "is", "to", "put", "class", "method", "definitions", "inside", "the", "class", "definition", "which", "makes", "it", "look", "like", "they", "have", "some", "special", "status."] ["But", "look", "at", "this", "etc."] ["You", "can", "see", "that", "some", "of", "the", "special", "treatment", "of", "classes", "--", "constants", "as", "names", "the", "separate", "notion", "of", "class", "method", "for", "their", "singleton", "methods", "--", "is", "just", "that", "special", "treatment."] ["Underneath", "a", "class", "is", "indeed", "an", "object."] ["CategoryDocumentation", "CategoryTutorial", "HomePage", "RecentChanges", "Preferences", "RubyGarden", "Edit", "text", "of", "this", "page", "View", "other", "revisions", "Last", "edited", "May", "am", "diff", "Search"] note that "i.e." was recognized :) However several problems are yet to be solved: * how to get rid of meaningful lone words? (last line) * what happens to things like As seen here: CODE bla bla bla. * etc However solving that would transform the 30mins. hack into a 1H kludge, better stay this way :) -- _ _ | |__ __ _| |_ ___ _ __ ___ __ _ _ __ | '_ \ / _` | __/ __| '_ ` _ \ / _` | '_ \ | |_) | (_| | |_\__ \ | | | | | (_| | | | | |_.__/ \__,_|\__|___/_| |_| |_|\__,_|_| |_| Running Debian GNU/Linux Sid (unstable) batsman dot geo at yahoo dot com Because I don't need to worry about finances I can ignore Microsoft and take over the (computing) world from the grassroots. -- Linus Torvalds --WIyZ46R2i8wDzkSu Content-Type: text/x-csrc; charset=us-ascii Content-Disposition: attachment; filename="bloom.c" #include static VALUE mBloom; static VALUE cBloom; static VALUE cMD5 = 0; static ID mHash; typedef struct { unsigned int capa; unsigned int size; unsigned char *ptr; } BloomData; /* list of primes 2^n+a, 10<=n<=30 plus selected primes in between ;) */ static long primes[] = { 1024 + 9, 2048 + 5, 4096 + 3, 8192 + 27, 16384 + 43, 25111, 32768 + 3, 48311, 65536 + 45, 92311, 131072 + 29, 192113, 262144 + 3, 368111, 524288 + 21, 756011, 1048576 + 7, 1500101, 2097152 + 17, 3000017, 4194304 + 15, 6000011, 8388608 + 9, 12000017, 16777216 + 43, 24000001, 33554432 + 35, 67108864 + 15, 134217728 + 29, 268435456 + 3, 536870912 + 11, 1073741824 + 85, 0 }; static void bl_free(void *p) { BloomData *ptr = (BloomData *)p; free(ptr->ptr); } static void set_bit(unsigned char *ptr, unsigned int position) { int bpos; int rest; bpos = position / 8; rest = position - bpos*8; ptr[bpos] |= 1 << rest; } static unsigned char is_bit_set(unsigned char *ptr, unsigned int position) { int bpos; int rest; bpos = position / 8; rest = position - bpos*8; return ptr[bpos] & (1 << rest); } static int compute_size(int size) { int i; for (i = 0; i < sizeof(primes)/sizeof(primes[0]); i++) if (primes[i] > size) return primes[i]; /* Ran out of polynomials */ rb_raise(rb_eRuntimeError, "BloomFilter would be too big."); } static VALUE bl_new(VALUE class, VALUE size) { BloomData *ptr; VALUE ret; int nsize; nsize = compute_size(NUM2INT(size)); ret = Data_Make_Struct(cBloom, BloomData, 0, bl_free, ptr); ptr->capa = nsize; ptr->size = 0; /* not really needed */ ptr->ptr = (unsigned char *)malloc(nsize/8+1); return ret; } static VALUE bl_add(VALUE self, VALUE obj) { BloomData *ptr; char* hash_ptr; VALUE hash; unsigned long key = 0; unsigned int i; Data_Get_Struct(self, BloomData, ptr); if ( rb_respond_to(obj, mHash ) ) hash = rb_funcall(obj, mHash, 0); else { if (!cMD5) { rb_require("digest/md5"); cMD5 = rb_eval_string("Digest::MD5"); } hash = rb_funcall(obj, rb_intern("to_s"), 0); hash = rb_funcall(cMD5, rb_intern("digest"), 1, hash); } hash_ptr = RSTRING(hash)->ptr; for(i = 0; i < (16/sizeof(unsigned long)); i++) { key = *((unsigned long *)(hash_ptr + i*sizeof(unsigned long))); set_bit(ptr->ptr, key % ptr->capa); } ptr->size++; return self; } static VALUE bl_contains(VALUE self, VALUE obj) { BloomData *ptr; char* hash_ptr; VALUE hash; unsigned long key = 0; unsigned int i; Data_Get_Struct(self, BloomData, ptr); if ( rb_respond_to(obj, mHash ) ) hash = rb_funcall(obj, mHash, 0); else { if (!cMD5) { rb_require("digest/md5"); cMD5 = rb_eval_string("Digest::MD5"); } hash = rb_funcall(obj, rb_intern("to_s"), 0); hash = rb_funcall(cMD5, rb_intern("digest"), 1, hash); } hash_ptr = RSTRING(hash)->ptr; for(i = 0; i < (16/sizeof(unsigned long)); i++) { key = *((unsigned long *)(hash_ptr + i*sizeof(unsigned long))); if ( !is_bit_set(ptr->ptr, key % ptr->capa) ) return Qfalse; } return Qtrue; } static VALUE bl_capacity(VALUE self) { BloomData *ptr; Data_Get_Struct(self, BloomData, ptr); return INT2FIX(ptr->capa); } static VALUE bl_size(VALUE self) { BloomData *ptr; Data_Get_Struct(self, BloomData, ptr); return INT2FIX(ptr->size); } void Init_bloom() { cBloom = rb_define_class("BloomFilter", rb_cObject); mHash = rb_intern("bloom_hash"); rb_define_singleton_method(cBloom, "new", bl_new, 1); rb_define_method(cBloom, "add", bl_add, 1); rb_define_method(cBloom, "include", bl_contains, 1); rb_define_method(cBloom, "size", bl_size, 0); rb_define_method(cBloom, "capacity", bl_capacity, 0); rb_define_alias(cBloom, "[]", "include"); rb_define_alias(cBloom, "length", "size"); rb_define_alias(cBloom, "capa", "capacity"); } --WIyZ46R2i8wDzkSu Content-Type: text/plain; charset=us-ascii Content-Disposition: attachment; filename="extconf.rb" require 'mkmf' create_makefile("bloom") --WIyZ46R2i8wDzkSu Content-Type: text/plain; charset=iso-8859-1 Content-Disposition: attachment; filename="scanhtml.rb" Content-Transfer-Encoding: quoted-printable X-MIME-Autoconverted: from 8bit to quoted-printable by funfun.nagaokaut.ac.jp id h5C6i5265988 require 'strscan' require 'bloom' class HtmlScanner attr_accessor :on_word, :on_not_word =09 def initialize @on_word =3D proc {|x| } @on_not_word =3D proc {|x| } end =09 def scan(aString) ow =3D @on_word onw =3D @on_not_word # make them method calls to increase speed class << self self end.instance_eval do define_method(:_on_word, ow) define_method(:_on_not_word, onw) end =09 s =3D StringScanner.new aString nolt =3D /([^>\"\']|\"([^"]|[^\\]\\\")*?\"|\'([^']|[^\\]\\\')*?\')*?/m re1 =3D /\s*<\s*(?:script|option|title|style|pre)\s*#{nolt}>(.*?)<\/(?:= script|option|title|style|pre)>\s*/im=20 re1a =3D /\s*\s*/m re2 =3D /\s*<#{nolt}>\s*/xm re3 =3D /[^a-zA-Z=E1=E9=ED=F3=FA=F1=D1=E4=EB=EF=FC=F6=DF=E7=C1=C9=CD=D3= =DA=E0=E8=EC=F2=F9=C0=C8=CC=D2=D9=C4=CB=CF=D6=DC&;<>\.\!\?'-]+/im aWord =3D /([a-zA-Z=E1=E9=ED=F3=FA=F1=D1=E4=EB=EF=FC=F6=DF=E7=C1=C9=CD=D3= =DA=E0=E8=EC=F2=F9=C0=C8=CC=D2=D9=C4=CB=CF=D6=DC.?!'-]|\&([^;<]+?);)+/i while s.rest? if s.scan re1 # ignore cause inside unwanted tags _on_not_word s.matched else=20 if s.scan re1a _on_not_word s.matched else # if here, not inside