Package org.apache.lucene.analysis.uk
Class UkrainianMorfologikAnalyzer
java.lang.Object
org.apache.lucene.analysis.Analyzer
org.apache.lucene.analysis.StopwordAnalyzerBase
org.apache.lucene.analysis.uk.UkrainianMorfologikAnalyzer
- All Implemented Interfaces:
Closeable,AutoCloseable
A dictionary-based
Analyzer for Ukrainian.- Since:
- 6.2.0
-
Nested Class Summary
Nested classes/interfaces inherited from class org.apache.lucene.analysis.Analyzer
Analyzer.ReuseStrategy, Analyzer.TokenStreamComponents -
Field Summary
Fields inherited from class org.apache.lucene.analysis.StopwordAnalyzerBase
stopwordsFields inherited from class org.apache.lucene.analysis.Analyzer
GLOBAL_REUSE_STRATEGY, PER_FIELD_REUSE_STRATEGY -
Constructor Summary
ConstructorsConstructorDescriptionBuilds an analyzer with the default stop words.UkrainianMorfologikAnalyzer(CharArraySet stopwords) Builds an analyzer with the given stop words.UkrainianMorfologikAnalyzer(CharArraySet stopwords, CharArraySet stemExclusionSet) Builds an analyzer with the given stop words. -
Method Summary
Modifier and TypeMethodDescriptionprotected Analyzer.TokenStreamComponentscreateComponents(String fieldName) Creates aAnalyzer.TokenStreamComponentswhich tokenizes all the text in the providedReader.static CharArraySetReturns the default stopword set for this analyzerprotected ReaderinitReader(String fieldName, Reader reader) Methods inherited from class org.apache.lucene.analysis.StopwordAnalyzerBase
getStopwordSet, loadStopwordSet, loadStopwordSetMethods inherited from class org.apache.lucene.analysis.Analyzer
attributeFactory, close, getOffsetGap, getPositionIncrementGap, getReuseStrategy, initReaderForNormalization, normalize, normalize, tokenStream, tokenStream
-
Constructor Details
-
UkrainianMorfologikAnalyzer
public UkrainianMorfologikAnalyzer()Builds an analyzer with the default stop words. -
UkrainianMorfologikAnalyzer
Builds an analyzer with the given stop words.- Parameters:
stopwords- a stopword set
-
UkrainianMorfologikAnalyzer
Builds an analyzer with the given stop words. If a non-empty stem exclusion set is provided this analyzer will add aSetKeywordMarkerFilterbefore stemming.- Parameters:
stopwords- a stopword setstemExclusionSet- a set of terms not to be stemmed
-
-
Method Details
-
getDefaultStopwords
Returns the default stopword set for this analyzer -
initReader
- Overrides:
initReaderin classAnalyzer
-
createComponents
Creates aAnalyzer.TokenStreamComponentswhich tokenizes all the text in the providedReader.- Specified by:
createComponentsin classAnalyzer- Returns:
- A
Analyzer.TokenStreamComponentsbuilt from anStandardTokenizerfiltered withLowerCaseFilter,StopFilter,SetKeywordMarkerFilterif a stem exclusion set is provided andMorfologikFilteron the Ukrainian dictionary.
-