Lucene 7.5.0 how to set lowercase expanded terms to true

Viewed 452

I have implemented my own Analyzer, QueryParser and PerFieldAnalyzerWrapper to implement the ElasticSearch ${field}.raw feature. Everything seems to be working ok, except for when I test using wildcards, etc on StringField types.

I understand this is because these queries don't use the analyzer at all.

In previous versions of lucene, there was a config option to enable the lowercasing of these queries.

I can't find how to do this in the latest version 7.5.0. Can anyone shed some light on this?

2 Answers

Expanded terms are processed by Analyzer.normalize. Since you have implemented your own Analyzer, add an implementation of the normalize method which runs the tokenStream through a LowerCaseFilter.

It can be as simple as:

public class MyAnalyzer extends Analyzer {
    protected TokenStreamComponents createComponents(String fieldName) {
        //Your createComponents implementation
    }

    protected TokenStream normalize(String fieldName, TokenStream in) {
        return new LowerCaseFilter(in);
    }
}

You can set up an analyzer like this for more details you can check out this link Git link for CJK Bigram Plugin

    @BeforeClass
public static void setUp() throws Exception {
    analyzer = new Analyzer() {
        @Override
        protected TokenStreamComponents createComponents(String fieldName) {
            Tokenizer source = new IcuTokenizer(AttributeFactory.DEFAULT_ATTRIBUTE_FACTORY,
                    new DefaultIcuTokenizerConfig(false, true));
            TokenStream result = new CJKBigramFilter(source);
            return new TokenStreamComponents(source, new StopFilter(result, CharArraySet.EMPTY_SET));
        }
    };
    analyzer2 = new Analyzer() {
        @Override
        protected TokenStreamComponents createComponents(String fieldName) {
            Tokenizer source = new IcuTokenizer(AttributeFactory.DEFAULT_ATTRIBUTE_FACTORY,
                    new DefaultIcuTokenizerConfig(false, true));
            TokenStream result = new IcuNormalizerFilter(source,
                    Normalizer2.getInstance(null, "nfkc_cf", Normalizer2.Mode.COMPOSE));
            result = new CJKBigramFilter(result);
            return new TokenStreamComponents(source, new StopFilter(result, CharArraySet.EMPTY_SET));
        }
    };
Related