*edit* Sorry – jumped the gun with my original test code here – need to close the IndexWriter after the optimize! The gains are only with multi segment indexes. Corrected entry follows:
Lets do a little test. We will load up a FieldCache with 5,000,000 unique strings and see how long it takes Lucene 2.4 in comparison to Lucene 2.9.
Lets use my quad core laptop and the following test code:
public class ContrivedFCTest extends TestCase { public void testLoadTime() throws Exception { Directory dir = FSDirectory.getDirectory(System.getProperty("java.io.tmpdir") + File.separator + "test"); IndexWriter writer = new IndexWriter (dir, new SimpleAnalyzer(), true, IndexWriter.MaxFieldLength.LIMITED); writer.setMergeFactor(37); writer.setUseCompoundFile(false); for(int i = 0; i < 5000000; i++) { Document doc = new Document(); doc.add (new Field ("field", "String" + i, Field.Store.NO, Field.Index.NOT_ANALYZED)); writer.addDocument(doc); } writer.close(); IndexReader reader = IndexReader.open(dir); long start = System.currentTimeMillis(); FieldCache.DEFAULT.getStrings(reader, "field"); long end = System.currentTimeMillis(); System.out.println("load time:" + (end - start)/1000.0f + "s"); } }
The results?
Lucene 2.4: 150.726s
Lucene 2.9: 9.695s
We discovered early this year that in the past, Lucene has been terribly inefficient when loading FieldCaches over multiple segments. Lucene 2.9 addresses this at the MultiReader level (thank you Yonik!). Also, internal FieldCache usage is now per segment, which sidesteps loading FieldCaches over mutiple segments all together – each segment has its own FieldCache.
This was our biggest issue by far. Its taken significant load off our servers when installing a new snapshot.
Thanks indeed Yonik.
September 22, 2009 16:16 — Jim Murphy - PostRank
Why did you use merge factor of 37?
September 23, 2009 19:46 — Anonymous
mergeFactor=37 — presumably in order to avoid any segments to be merged during indexing, thus making it possible to show off the new and faster segment reloading.
September 24, 2009 10:15 — Otis Gospodnetic
Partially – since I am just timing the loading of the FieldCache (and not doing it per segment). It’s just to make sure I have a bunch of segments – its only faster over multiple segments – its the same speed on an optimized Index. The more segments, the faster it is.
The reason that its also faster when you do it per segment (how Lucene works internally now), is that it avoids the speed trap that was in MultiTermEnum, and uses SegmentTermEnum – Yonik fixed that as well though, and this test shows the fruits of that. So essentially, it was both fixed and side stepped at the same time
September 24, 2009 11:10 — Mark Miller
[...] 在保证了正确性之后,要关注的便是性能了。根据我的推测,由于IKVM需要在Java生成的.NET程序集和BCL之间加上一层Runtime和JDK,因此其性能几乎一定会比Java原有的程序要差。不过,对于Lucene这种项目来说,算法才是性能的关键。例如,有人测试Lucene 2.9.0在某些情况下会比2.4有15倍左右的性能提升。不过由于没有很好的测试数据和场景,目前我只进行了最最简单的,不涉及磁盘IO的性能比较。 [...]
September 24, 2010 07:49 — 尝试使用IKVM运行Lucene 2.9.0版 | 彭旭赣州seo优化
[...] Imagination的Mark Miller运行了一个简单的性能测试,表明在5,000,000个不同字符串下的情况下,Lucene [...]
September 25, 2010 03:37 — Apache Lucene 2.9的改进 | 彭旭赣州seo优化