Rowles.LeanCorpus.Analysis.Filters
Classes
AccentFoldingFilter
Normalises accented/diacritic characters to their ASCII base form (e.g., é→e, ñ→n, ü→u) for language-neutral matching. Uses Unicode canonical decomposition followed by stripping combining marks.
CachingTokenFilter
Captures tokens as they pass through the filter chain, enabling multiple passes over the same token stream without re-running the analysis pipeline.
ClassicFilter
Removes possessive endings and periods from acronyms, replicating Lucene's ClassicFilter.
CommonGramsFilter
Produces bigrams of consecutive common words to improve phrase query recall.
DecimalDigitFilter
Normalises Unicode decimal digits to ASCII digits.
ElisionFilter
Removes configured elided articles before straight or curly apostrophes.
FlattenGraphFilter
Normalises token position increments so same-position alternates remain explicit and the stream stays consumable by LeanCorpus's linear postings model.
HtmlStripCharFilter
Strips HTML/XML tags and HTML entities from input text, leaving only text content. Uses a span-based scanner to avoid regex allocations.
HunspellDictionary
Lightweight Hunspell dictionary that supports short-form flags and simple prefix and suffix rules.
HunspellStemFilter
Stems tokens using a pre-parsed HunspellDictionary.
HyphenatedWordsFilter
Joins consecutive tokens that share the same position into a single hyphenated token.
KeepWordFilter
Removes any token whose text is not present in the configured keep-word set.
KeywordMarkerFilter
Identifies tokens that should be treated as keywords by compatible analysers.
LengthFilter
Removes tokens whose text length falls outside an inclusive range.
LimitTokenCountFilter
Truncates the token stream after a fixed number of emitted tokens.
LowercaseFilter
Performs an in-place lowercase transformation on tokens or a character buffer.
MappingCharFilter
Maps specific characters or strings to replacements using a lookup table. Useful for normalising special characters (e.g., smart quotes → straight quotes).
MetaphoneFilter
Emits Metaphone encodings for tokens.
PatternReplaceCharFilter
Replaces text matching a regex pattern with a replacement string.
PatternReplaceFilter
Applies a regex replacement to each token's text.
PhoneticAlternatesFilter
Emits bounded Latin-name phonetic alternates at the same token position.
PorterStemmerFilter
Porter Stemming Algorithm implementation as an ISpanTokenFilter. Based on the Porter 1980 specification for English stemming. Operates on tokens in-place, replacing text with stemmed form.
ReverseStringFilter
Reverses the characters in each token.
ShingleFilter
Emits contiguous token shingles for phrase-oriented analysis.
StemTokenFilter
Applies an ISpanStemmer to each token in the list. Useful as a drop-in filter in the composable Analyser pipeline.
StopWordFilter
Removes common English stop words from a token list using a frozen set for fast, allocation-free lookups.
SynonymGraphFilter
Token filter that supports multi-token synonym expansion using a trie-based SynonymMap. Uses longest-match lookahead for multi-word synonyms and inserts replacement tokens at the same position offsets.
SynonymMap
Trie-based synonym map supporting multi-token source phrases. Used by SynonymGraphFilter for longest-match multi-token synonym expansion.
TruncateTokenFilter
Truncates token text to a maximum character length.
TypeTokenFilter
Filters tokens by extensible token type.
UniqueTokenFilter
Removes duplicate tokens at the same position, keeping the first occurrence.
WordDelimiterFilter
Splits compound tokens on word delimiters, case transitions, and letter-digit boundaries.
Interfaces
ICharFilter
Interface for character-level filters that transform raw text before tokenisation. Char filters run before the tokeniser, operating on the entire input string.