Incomplete search when using Chinese

What’s The Bug?

Dear Anytype Team,
After installing and using version 0.40.21, I’ve come across two issues that I would like to bring to your attention.
Search Bar Display Issue: The search bar is not displaying completely, as shown in the attached image. This does not appear to be a screenshot issue.
Chinese Search Result Issue: The search results for Chinese queries seem inaccurate. For instance, when I search for “处理AI” (Process AI), the results suggest that Anytype is searching for “处,” “理,” and “AI” separately and then ranking results based on the number of matches. However, Anytype fails to prioritize results containing the full phrase “处理.”
As highlighted in the attached image, the more accurate search result for the Chinese phrase “处理” is not displayed at the top. Furthermore, only “处” is highlighted, while “理” is not, which seems like another potential bug.
[Attach Image Here]
Thank you for your time and attention to these matters. I appreciate your efforts in continually improving Anytype.
Sincerely,
itez

How To Reproduce It

Ctrl + S

Image or Video

The Expected Behavior

I hope Anytype can search Chinese more perfectly.

Device

ThinkPad E14

OS

Windows 11

Anytype Version

v0.40.21

Network Mode

AnySync

Technical Information

操作系统版本: win32 x64 10.0.22631
应用版本: 0.40.21-beta
构建版本: build on 2024-05-24 09:27:24 +0000 UTC at #ca126ef233de85a42cf54dcb53e060f304f2b123 (dirty)
库版本: v0.34.0-rc4
Anytype 身份标识: AAfozvNJ4mFLYftgfMABui9RXSVGvMiG822PKvk8KVEjGXLx
分析 ID: bac6e8a0-075f-4f9d-bf98-2ebe8236cee5
设备 ID: 12D3KooWCEKBEkCcaMSofTxjbfsz4KU3we6rqcdCTcSLf6uHM4Ao

@itez Please create a separate bug report for the search bar display issue.

This report has been added to our issue tracker and received by the Development Team.

Hi @itez !

Our team is aware of the issue & working on a solution to optimize the search engine and provide better results for non-English languages. We hope to implement it within the next few weeks :sparkles:

I have also done some research on this issue today. The problem may lie in the word segmentation. Since I don’t understand Go, I can’t personally verify its feasibility.

Changes needed: anytype-heart/pkg/lib/localstore/ftsearch/ftsearch.go at main · anyproto/anytype-heart · GitHub

Reference code: bleve结合 jieba 分词实现中文分词 · GitHub

Hi, I also have the same question when searching Chinese words.
It seems that the question wasn’t be solved in the version 0.42.4.
I really hope this question can be solved in the upcoming update version. Thank you very much.

We are switching to Tantivy, but the problem hasn’t been solved yet. We are still trying to figure out how to implement hieroglyph encodings.

We need to provide an option to choose the preferred language for FTS (Full-Text Search) because each language is different.

For Japanese, Chinese, Korean, etc., only a dictionary-based approach seems to work effectively. Therefore, we must include all these dictionaries (5 MB for the Chinese language, e.g., GitHub - messense/jieba-rs: The Jieba Chinese Word Segmentation Implemented in Rust).

The same applies to all other hieroglyphic languages. Alternatively, we could implement a mechanism that allows clients to choose the dictionaries themselves.
But it takes a lot of time.

Using word segmentation might have introduced a more serious problem: I’m having a hard time finding my notes.

For instance, with a note titled “远程办公” searching for either “远程” or “办公” separately yields no results; I must input the complete term “远程办公” to find it. This is really frustrating.

The word segmentation is too conservative, making it difficult to search. Can future versions allow for setting more aggressive word segmentation?

My Anytype version is 0.43.2.

18ca824e46adb390305c99b7ad5017d2

0b10d6734d12dae04e3298d3decf0cac

I have encounter same problem with chinese search today,if I have document with name “水龍” I have to search them with whole word “水龍” insteadof single word “龍”. I can’t search all document related to “龍” because it only give me empty result since it always combined with other word, its annoying and made the search useless. I remember I didn’t encounter this behavior in older version but not sure which version. Im current at verison 0.43.4.

I recommend going back to version 0.42.8 for now, It’s really frustrating

please don’t paste downloads from unverified sources and use the github releases links instead. all the releases are still there.

and as a rule of thumb: always do backups before downgrading. the old client might be incompatible with the data structure of newer versions. this should not be recommended to average users.

Version 0.44.0 still has this issue. Is there an expected timeline for its resolution? Searching is a critical function for PKM.

If implementing word segmentation for graphical languages involves a large workload, could you first provide a full-word matching option? I believe this is a simple and less error-prone approach.