Search in sources :

Example 1 with TokenStreamComponents

use of org.apache.lucene.analysis.Analyzer.TokenStreamComponents in project Krill by KorAP.

the class TextAnalyzer method createComponents.

@Override
protected TokenStreamComponents createComponents(final String fieldName) {
    final Tokenizer source = new StandardTokenizer();
    TokenStream sink = new LowerCaseFilter(source);
    return new TokenStreamComponents(source, sink);
}
Also used : TokenStream(org.apache.lucene.analysis.TokenStream) StandardTokenizer(org.apache.lucene.analysis.standard.StandardTokenizer) Tokenizer(org.apache.lucene.analysis.Tokenizer) StandardTokenizer(org.apache.lucene.analysis.standard.StandardTokenizer) LowerCaseFilter(org.apache.lucene.analysis.core.LowerCaseFilter) TokenStreamComponents(org.apache.lucene.analysis.Analyzer.TokenStreamComponents)

Example 2 with TokenStreamComponents

use of org.apache.lucene.analysis.Analyzer.TokenStreamComponents in project Krill by KorAP.

the class KeywordAnalyzer method createComponents.

@Override
protected TokenStreamComponents createComponents(final String fieldName) {
    final Tokenizer source = new WhitespaceTokenizer();
    TokenStream sink = new LowerCaseFilter(source);
    return new TokenStreamComponents(source, sink);
}
Also used : WhitespaceTokenizer(org.apache.lucene.analysis.core.WhitespaceTokenizer) TokenStream(org.apache.lucene.analysis.TokenStream) WhitespaceTokenizer(org.apache.lucene.analysis.core.WhitespaceTokenizer) Tokenizer(org.apache.lucene.analysis.Tokenizer) LowerCaseFilter(org.apache.lucene.analysis.core.LowerCaseFilter) TokenStreamComponents(org.apache.lucene.analysis.Analyzer.TokenStreamComponents)

Example 3 with TokenStreamComponents

use of org.apache.lucene.analysis.Analyzer.TokenStreamComponents in project epadd by ePADD.

the class EnglishNumberAnalyzer method createComponents.

/*
 * Creates a {@link org.apache.lucene.analysis.Analyzer.TokenStreamComponents} which
 * tokenizes all the text in the reader that is set after the instantiation of this analyzer .
 *
 * @return A {@link org.apache.lucene.analysis.Analyzer.TokenStreamComponents} built from an {@link StandardTokenizer} filtered with {@link StandardFilter}, {@link EnglishPossessiveFilter}, {@link LowerCaseFilter}, {@link StopFilter} , {@link SetKeywordMarkerFilter} if a stem exclusion set is
 *         provided and {@link PorterStemFilter}.
 **/
protected TokenStreamComponents createComponents(String fieldName) {
    final Tokenizer source = new StandardNumberTokenizer();
    TokenStream result = new StandardFilter(source);
    // @TODO document what each of these filters do
    result = new EnglishPossessiveFilter(result);
    result = new LowerCaseFilter(result);
    result = new StopFilter(result, this.stopwords);
    if (!this.stemExclusionSet.isEmpty()) {
        result = new SetKeywordMarkerFilter((TokenStream) result, this.stemExclusionSet);
    }
    result = new PorterStemFilter((TokenStream) result);
    return new TokenStreamComponents(source, result);
}
Also used : TokenStream(org.apache.lucene.analysis.TokenStream) EnglishPossessiveFilter(org.apache.lucene.analysis.en.EnglishPossessiveFilter) StopFilter(org.apache.lucene.analysis.StopFilter) SetKeywordMarkerFilter(org.apache.lucene.analysis.miscellaneous.SetKeywordMarkerFilter) StandardFilter(org.apache.lucene.analysis.standard.StandardFilter) PorterStemFilter(org.apache.lucene.analysis.en.PorterStemFilter) Tokenizer(org.apache.lucene.analysis.Tokenizer) LowerCaseFilter(org.apache.lucene.analysis.LowerCaseFilter) TokenStreamComponents(org.apache.lucene.analysis.Analyzer.TokenStreamComponents)

Aggregations

TokenStreamComponents (org.apache.lucene.analysis.Analyzer.TokenStreamComponents)3 TokenStream (org.apache.lucene.analysis.TokenStream)3 Tokenizer (org.apache.lucene.analysis.Tokenizer)3 LowerCaseFilter (org.apache.lucene.analysis.core.LowerCaseFilter)2 LowerCaseFilter (org.apache.lucene.analysis.LowerCaseFilter)1 StopFilter (org.apache.lucene.analysis.StopFilter)1 WhitespaceTokenizer (org.apache.lucene.analysis.core.WhitespaceTokenizer)1 EnglishPossessiveFilter (org.apache.lucene.analysis.en.EnglishPossessiveFilter)1 PorterStemFilter (org.apache.lucene.analysis.en.PorterStemFilter)1 SetKeywordMarkerFilter (org.apache.lucene.analysis.miscellaneous.SetKeywordMarkerFilter)1 StandardFilter (org.apache.lucene.analysis.standard.StandardFilter)1 StandardTokenizer (org.apache.lucene.analysis.standard.StandardTokenizer)1