From: Charles Oliver Nutter Date: 2010-03-08T00:34:31+09:00 Subject: [ruby-core:28549] Re: [Feature #905] Add String.new(fixnum) to preallocate large buffer On Sun, Mar 7, 2010 at 4:58 AM, Yusuke ENDOH wrote: > Ko1 told me that GC makes the second benchmark slower than JRuby. > In MRI, a string literal is duplicated whenever evaluated. > I moved the literals out of the loop: JRuby behaves the same, since literal strings are still separate objects and mutable. > ��results.report "'' <<" do > �� ��s = '' > �� ��s1, s2 = '.', 'word' > �� ��N.times { s << s1 << s2 } > ��end > > ��ruby19 > �� �� �� �� �� �� �� �� �� �� �� �� �� ��user �� �� system �� �� ��total �� �� �� ��real > ��'' << �� �� �� �� �� �� �� �� 6.810000 �� 0.040000 �� 6.850000 ( ��6.851979) > > ��jruby > �� �� �� �� �� �� �� �� �� �� �� �� �� ��user �� �� system �� �� ��total �� �� �� ��real > ��'' << �� �� �� �� �� �� �� �� 7.159000 �� 0.000000 �� 7.159000 ( ��7.126000) > > Indeed, there is room for optimization in MRI, but in this case, > it is not in string concatenation, I guess. My numbers came out somewhat differently. Make sure you're running with the JVM's "server" mode if you run on Hotspot (Sun/OpenJDK): ~/projects/jruby ��� jruby --server string_bench.rb user system total real loop 0.572000 0.000000 0.572000 ( 0.523000) '' << 1.470000 0.000000 1.470000 ( 1.470000) ~/projects/jruby ��� ruby1.9 string_bench.rb user system total real loop 0.810000 0.000000 0.810000 ( 0.838414) '' << 2.670000 0.040000 2.710000 ( 2.733041) Here's numbers with a prototypical String.buffer implementation: ~/projects/jruby ��� jruby --server string_bench.rb user system total real loop 0.655000 0.000000 0.655000 ( 0.606000) '' << 1.390000 0.000000 1.390000 ( 1.390000) user system total real loop 0.321000 0.000000 0.321000 ( 0.321000) '' << 1.241000 0.000000 1.241000 ( 1.241000) user system total real loop 0.314000 0.000000 0.314000 ( 0.314000) '' << 1.229000 0.000000 1.229000 ( 1.229000) Of course, this 10-15% improvement could simply be because the JVM does not provide a "realloc" for its arrays (for various reasons, some of them presumably because it moves objects around in memory a lot). In order to grow a string, we have to allocate a new array and copy its contents. Under those circumstances, String.buffer makes a lot of sense, since the copying can get expensive at large sizes. I don't know enough about MRI internals to implement an equivalent String.buffer, but here's the patch to JRuby: diff --git a/src/org/jruby/RubyString.java b/src/org/jruby/RubyString.java index 71e6b63..e618ec8 100644 --- a/src/org/jruby/RubyString.java +++ b/src/org/jruby/RubyString.java @@ -451,6 +451,11 @@ public class RubyString extends RubyObject implements EncodingCapable { public static RubyString newStringLight(Ruby runtime, int size) { return new RubyString(runtime, runtime.getString(), new ByteList(size), false); } + + @JRubyMethod(meta = true) + public static IRubyObject buffer(ThreadContext context, IRubyObject self, IRubyObject size) { + return newStringLight(context.getRuntime(), (int)size.convertToInteger().getLongValue()); + } public static RubyString newString(Ruby runtime, CharSequence str) { return new RubyString(runtime, runtime.getString(), str);