Changing Bits: Eating dog food with Lucene

Monday, May 13, 2013

Eating dog food with Lucene

Eating your own dog food is important in all walks of life: if you are a chef you should taste your own food; if you are a doctor you should treat yourself when you are sick; if you build houses for a living you should live in a house you built; if you are a parent then try living by the rules that you set for your kids (most parents would fail miserably at this!); and if you build software you should constantly use your own software.

So, for the past few weeks I've been doing exactly that: building a simple Lucene search application, searching all Lucene and Solr Jira issues, and using it instead of Jira's search whenever I need to go find an issue.

It's currently running at jirasearch.mikemccandless.com and it's still quite rough (feedback welcome!).

It's a good showcase of a number of Lucene features:

Highlighting using the new PostingsHighlighter; for example, try searching for fuzzy query.
Autosuggest, using the not-yet-committed AnalyzingInfixSuggester (LUCENE-4845).
Sorting by various fields, including a blended recency and relevance sort.
A few synonym examples, for example try searching for oome.
Near-real-time searching, and controlled searcher versions: the server uses NRTManager, SearcherLifetimeManager and SearcherManager.
ToParentBlockJoinQuery: each issue is indexed as a parent document, and then each comment on the issue is indexed as a separate child document. This allows the server to know which specific comment, along with its metadata, was a match for the query, and if you click on that comment (in the highlighted results) it will take you to that comment in Jira. This is very helpful for mega-issues!
Okapi BM25 for ranking.

The drill-downs on the left also show a number of features from Lucene's facet module:

Drill sideways for all fields, so that the field does not disappear when you drill down on it.
Dynamic range faceting: the Updated drill-down is computed dynamically, e.g. all issues updated in the past week.
Hierarchical fields, which are simple since the Lucene facet module supports hierarchy natively. Only the Component dimension is hierarchical, e.g. look at the Component drill down for all Lucene core issues.
Multi-select faceting (hold down the shift key when clicking on a value), e.g. all improvements and new features.
Multi-valued fields (e.g. User, Fix version, Label).

This is really eating two different dog foods: first, as a developer I see what one must go through to build a search server on top of Lucene's APIs, but second, as an end user, I experience the resulting search user interface whenever I need to find a Lucene or Solr issue. It's like having to eat both wet and dry dog food at once, and both kinds of dog food have uncovered numerous issues!

The issues ranged from outright bugs such as PostingsHighlighter picking the worst snippets instead of the best (LUCENE-4826), to missing features like dynamic numeric range facets (LUCENE-4965), to issues that make consuming Lucene's APIs awkward, especially when mixing different features, such as the difficulty of mixing non-range and range facets with DrillSideways (LUCENE-4980) and the difficulty of using NRTManager with both a taxonomy index and a search index (LUCENE-4967), or finally just inefficient, such as the inability to customize how PostingsHighlighter loads its field values (LUCENE-4846).

The process is far from done! There are still a number of issues that need fixing. For example, it's not easy to mix ToParentBlockJoinQuery and grouping, which is frustrating because fields like who reported an issue, severity, issue status, component would all be natural group-by fields. Some issues, such as the inability of PostingsFormatter to render directly to a JSONObject (LUCENE-4896) are still open because they are challenging to fix cleanly. Others, such as the infix suggester (LUCENE-4845) are in limbo because of disagreements on the best approach, while still others, like BlendedComparator used to sort by mixed relevance and recency, I just haven't pushed back into Lucene yet.

There are plenty of ways to provoke an error from the server; here are two fun examples: try fielded search such as summary:python, or a multi-select drilldown on the Updated field.

Much work remains and I'll keep on eating both wet and dry dog food: it's a very productive way to find problems in Lucene!

38 comments:

Ivan BrusicMay 13, 2013 at 4:20 PM
The synonym example is using an (internal?) IP address.

Are you using PyLucene to connect with a Lucene instance?
ReplyDelete
Replies
Michael McCandlessMay 13, 2013 at 4:38 PM
Hi Ivan,

Woops: I fixed the synonym example link ... thanks.

I'm made a simple HTTP server (using Netty) to wrap Lucene APIs as JSON, and I access Lucene through that currently.
ReplyDelete
Replies
SudarshanMay 13, 2013 at 8:33 PM
It would be great if the source code of this app were available. I think lots of Lucene users would benefit from having this code to dig in when looking for examples of how to accomplish certain search features using Lucene.
ReplyDelete
Replies
UnknownMay 14, 2013 at 4:17 AM
Hey Mike, this is very cool and super fast! I'm curious about the suggester, how/when do you rebuild it? Do you do it every X hours or maybe something more sophisticated?
ReplyDelete
Replies
Mark HarwoodMay 14, 2013 at 4:43 AM
Very, very nice. I'm a big supporter of the "eat your own dog food" mindset. It's the best approach to uncovering issues or finding potential new features.

Presumably the source in patches or related commit logs could provide a way of tying the Java package name structures in as a hierarchical facet?
ReplyDelete
Replies
UnknownMay 14, 2013 at 11:12 PM
Excellent stuff
ReplyDelete
Replies
Alexandre RafalovitchMay 15, 2013 at 9:15 AM
Very nice and useful too.

Would be nice to have Author field in the search. I frequently do "my issues that mention X". Putting in my name and issue keyword would be great. Currently, it picks up my name but only in comments.
ReplyDelete
Replies
varunthackerMay 15, 2013 at 11:08 AM
Hi,

Awesome Idea. I'm inspired to make an application myself :) Could you point me to a good resource where I could use Netty to accept JSON, parse it to consume the Lucene API's. Till you publish your code I could get started.

Also I had a couple of questions.
I wanted to know how did you build weights for the AnalyzingInfixSuggester? Did you use issue priority /open closed?

You have used ToParentBlockJoinQuery for indexing comments. If instead we used a multivalued field what would be different?

ReplyDelete
Replies
Han JiangMay 17, 2013 at 7:58 AM
The speed is amazing! Quite interesting if we can use this to search other apache issues :)
ReplyDelete
Replies
AnonymousJune 19, 2013 at 1:25 PM
Very cool! Waiting for a day this would be open sourced. If not the UI, at least the netty wrapper around lucene.

-Ramesh
ReplyDelete
Replies
UnknownJuly 31, 2013 at 5:03 PM
I suggest you sell it to http://www.thoughtworks-studios.com/mingle-agile-project-management it will be really good business.
ReplyDelete
Replies
UnknownAugust 19, 2013 at 5:54 PM
Hi-- This is great stuff. I've made an adapted version that glues the suggestion onto the end of the search passed in (after truncating whatever suffix of the search is a prefix of the suggestion).

I was wondering about the way you open your readers in the build() method.

In particular, starting at line 257 (of the code from lucene 4.4.0), we have this:

r = new SlowCompositeReaderWrapper(DirectoryReader.open(w, false));
//long t1 = System.nanoTime();
w.rollback();

final int maxDoc = r.maxDoc();

In my naivete, it would have tried something along the lines of

w.close();
r = new SlowCompositeReaderWrapper(DirectoryReader.open(dirTmp));

Where dirTmp is the directory defined above on line 189.

Down below, after you create the second indexwriter that you fill from this sorted reader, you do a similar thing when you open the searcher.

Is this opening a reader from a writer then closing the writer something I should be doing? I tried going through the code a bit, but I still haven't gotten comfortable enough in the innards to get very far very quickly.

ReplyDelete
Replies
UnknownSeptember 18, 2013 at 7:57 AM
awesome, is there a way to have a look to the source code?
learning by studying the source code would help me a lot.
is that on github or something
ReplyDelete
Replies
ShyamsunderOctober 31, 2014 at 12:51 AM
Mike, how did you retrieve key value pair kind of results in the auto complete response. Do you add more than one field to the index when you build the suggestion dictionary?
http://jirasearch.mikemccandless.com/suggest.py?term=6025&source=titles_infix&index=jira&contexts=

[{"volLink": "http://issues.apache.org/jira/browse/LUCENE-6025", "label": "LUCENE-6025: Add BitSet.prevSetBit"}, {"volLink": "http://issues.apache.org/jira/browse/SOLR-6025", "label": "SOLR-6025: StreamingUpdateSolrServer is mentioned in various schema.xml files"}]
ReplyDelete
Replies
ShyamsunderOctober 31, 2014 at 1:11 AM
Another example is
http://jirasearch.mikemccandless.com/suggest.py?term=candles&source=titles_infix&index=jira&contexts=

[{"user": "Michael McCandless", "label": "user: \"Michael McCandless\""}]
ReplyDelete
Replies
UnknownJanuary 16, 2017 at 11:49 PM
Thanks
ReplyDelete
Replies

Add comment