We are switching to Tantivy, but the problem hasn’t been solved yet. We are still trying to figure out how to implement hieroglyph encodings.
We need to provide an option to choose the preferred language for FTS (Full-Text Search) because each language is different.
For Japanese, Chinese, Korean, etc., only a dictionary-based approach seems to work effectively. Therefore, we must include all these dictionaries (5 MB for the Chinese language, e.g., GitHub - messense/jieba-rs: The Jieba Chinese Word Segmentation Implemented in Rust).
The same applies to all other hieroglyphic languages. Alternatively, we could implement a mechanism that allows clients to choose the dictionaries themselves.
But it takes a lot of time.