Hello everyone Thanks for rising this up. I will add this idea in our product backlog.
For now Anytype is not really optimised for working with visual information like pdf and images. And we want address this request as a part of integral initiative “visual content as a first class citizen” to provide solid experience.
For now we are actively preparing for public release and we could start this initiative after public release.
11/8/23
This was probably mentioned before, but OneNote also has this hability, and it seems to process the images once at their creation, and keep the text assosiated on the background, so the search doesn’t have to trigger the OCR each time. Also, getting the word highlighted is a must. I know this is a big feature, but I think it’s one that most people would benefit from.
Edit 3/3/25:
Just came back to see if there was any news regarding this request. I’ll add that this other request would complement OCR capability quite nicely: Mobile document scanner
I also can’t believe this request is +3 years old. They grow up so quickly
Implementing this is definitely challenging because Tesseract is quite outdated and fails about 99% of the time. There are much better vision models available here: Models - Hugging Face. However, integrating a high-quality, compact vision model into Anytype seems almost impossible. While it would be a game-changer and incredibly impressive, I’m not very optimistic about seeing a reliable OCR feature implemented anytime soon.