From: Niko Schwarz Date: 2002-01-22T07:10:49+09:00 Subject: Re: ruby binary handling hi matz, Yukihiro Matsumoto wrote: > |> I've done some simple binary handling in Ruby > |> (http://www.ruby-lang.org/) - Find it ok for me ! :) > | > |dunno, binary streams were the only thing i _really_ had to disagree with > |about ruby. > | > |not only that i dont see any way to handle streams (which means for me: a > |pointer to some far resource, so that i dont need to keep my 2mg-picture > |entirely in memory when i just want to send it over the net. > | > |and then: was matz drunk when he decided to represent binary arrays as > |strings? > > I don't drink, so it never be the truth. ;-) > Anyway, do you have any idea about "better binary handling" API? Let me think aloud? We have two problems here, imho: a) characters should be all general, i.e. not just ascii or unicode but being subclassable to all the existing keycodes. that, in turn, would require a completely different string->character translation, because the char-code depends on the charset. im not sure how to model this, of course, but its true that byte[] and char[] are somewhat similar, so a common superclass, like with the numerics, would be useful. however, for both clobs (character large objects in rdbms-speech) and blobs (binary large objects) would appreciate some streaming mechanisms. im not sure how to achieve that. i could imagine that it might be a wise technique to check on reading of a string from a repeatable resource (such as the harddrive) if the free memory compared to the object size exceeds some determined ratio. if it does, then you could assume the user only wants to stream the data, and then you could handle them like a java-stream. maybe this would requir some change in the api, too, to determine better whether you have streaming or "normal" field access. though i dont think so, cos u use it as a field basically when either of two conditions are met: a) the content is modified b) the content is used heavily. b wouldnt be a big problem, cos even on pure streaming the OS buffers would keep this fast -- remember were talking bout *huge* blobs/clobs here. the real problem is that, with such a way, ud keep quite extensive reading-locks on files. that doesnt need to be a problem, but it would require some change in programmers behaviour, no more: close-as-soon-as-possible but keep-open-as-long-as-possible. in any way, there should be two different classes for strings and binary data, for a few reasons: - binary data doesnt need methods like capitalize! - string data doesnt need methods like xor - for string data its not certain at all that one character will consist of one byte. java solves the ladder with character-iterators, however, i myself think that in a clean language string[2] should return the second _character_ and nothing else. so, summed up, my ideas would be: - a common superclass for strings and binaries - different intern handling of those, depending on the size and the type theyre used maybe a string-buffer has some justification, too, a lil test-script will certainly clear this in no time. ive got another thing in my back head about ruby, bout nil, ill post this when i can clear it in my head. regards, nick -- Whenever people agree with me I always feel I must be wrong. -- Oscar Wilde