From: Trans Date: 2007-03-23T17:20:36+09:00 Subject: Re: Just for fun... On Mar 22, 11:45 pm, Clifford Heath wrote: > Trans wrote: > > On Mar 22, 6:00 am, Clifford Heath wrote: > >> Your analysis is well-meaning, but flawed. > > General compression is all about removing repetition. > > This is a very concise statement about what's flawed in your analysis. > It's not just about repetition, it's about removing *whatever* is > predictable. All's fair when it comes to improving predictability. Of course. When I used the term "repetition" I didn't mean to be so narrow as to to suggest nothing more than "a a a a a". A better term would have been "patterns". > If I cared enough to write a compressor that could recognise > Somerset Maugham's writing style and vocabulary, it could adapt > its general compression algorithm to that style and compress it > better. Sure the decompressor has to know his style as well, but > that's allowed - it can be a general compressor with specific modes > (like video, i386 opcodes, etc, as I mentioned). > > > Try assign a > > number to every word in the dictionary and "compress" a piece of text > > with it. It won't help. > > Funny you should say that because I did that once, building a dictionary > of the King James Bible (10.5Kwords) and Huffman coding the words. It > worked very well, thanks very much. It was better than deflate for the > purpose because you can start decompressing anywhere (though that's not > relevant to overall compression goodness). Right. Size is of first relevance. So was you file smaller than deflates or bzip2s? > > "How it does it" is the interesting part. How does BWT improve > > deflate? Ie. What's that tell us about deflate? > > It tells us that there's more to the compressing typical data sets > than simply identifying repetition - there is data that encodes > higher order patterns that are not directly visible as repetition, > but are visible to a better predictor. That's a little belittling, don't you think? I mean, you've taken one connotation of one word I used and seemingly made it the centerpiece of all my thoughts on the matter, while missing the larger picture of my point. I'm not saying your analysis is completely wrong. I'm just saying I think you are missing a subtle point about what makes compression work. Predictability is not the whole center piece of compression either. If it were, you couldn't get much better a compression of Kings James than a file with this in it: "http://patriot.net/~bmcgin/kjv12.txt" Pretty predictable. That may seems silly, but there's a point. This is not general compression. It doesn't really matter that the process of decompression involves a remote machine -- locality is relative. But even better, if I sent you a new algorithm for compressing King James, and along with it the file I sent you to uncompress had only one bit of data in it, "1", then how can it work? Of course, the algorithm itself had the entire content of King James within it. Pretty predictable again. But not general compression. When looking at general compression its important to take the size of all the elements involved as well, including the compression algorithm itself. I find this a very fascinating subject. I think ultimately it will turn out to be very important, not just for compressing data, but for understanding physics too. I've done some deep meditation on the topic. It is interesting to consider that all conceivable data can be found somewhere in any transcendental number. You might think, that being the case, an excellent compression algorithm would be to find the starting digit, say in Pi, for the data one is seeking. But curiously it would do you know good. On average the number for the starting digit would be just as long as the data itself. But I digress, if you think you are so sure about what you are saying, and that I am completely off-base, I think you should try to solve this puzzle: http://prize.hutter1.net/ I believe what they are trying to do is essentially impossible. At best they can shave a few bytes off here and there by removing some "higher-order" patterns, as you say. Their premise is flawed. Knowledge is no more compressible b/c it's knowledge, than anything else. They should realize either the human brain is just that vast or that human knowledge uses lossy compression. They'll gain nothing of great significance until they do. T.