Should a privacy-focused note-taking app adopt AI?

This post explores the topic: “Should a privacy-focused note-taking app adopt AI?” Everyone is welcome to join the discussion.

As a powerful productivity-boosting technology, AI is being embraced across all industries. For a note-taking app whose core value is data privacy, whether or not to introduce AI features is a topic of great interest.

Personally, there are a few points I’m particularly concerned about:

  1. Will Anytype introduce AI features?
  2. The value proposition of AI seems to contradict that of Web3.
  3. Assuming Anytype does adopt AI — as we all know, AI needs to understand content in order to assist — once users feed their data into the AI, no matter what the tech provider promises, in reality, data privacy is essentially gone.

P.S. My personal prediction: as AI makes data easier to interpret and understand, once AI technology reaches a certain point, people will start to feel the privacy intrusion more deeply — and as a result, data security will become even more valued.

First of all there is local and online AI and your concerns are only about online ones.

  1. This is listed in the Backlogs, so probably.
  2. Please explain why.
  3. You might give AI a very limited context window, there is so many path available to implement AI inside Anytype. Speculation doesn’t help much.

Besides, I don’t see a scenario where Anytype will force you to let AI have access to your data if you don’t want it to.

Disclaimer: I don’t want AI, at least for the time being. As far as I’m concerned, we’d have to improve Anytype’s capabilities first, to encourage use, and only then would AI intervene.

That said, AI can be useful.
And AI is not necessarily a risk for data.
AI is just an application.

A lot of companies (including large firms at the cutting edge of security) use AI. It’s like any tool: either you use a tool provided free of charge by a big foreign company: easy, powerful, free, but your data is exploited.
Or we use an in-house or open-source tool. At work, I’ve set up an image analysis AI: developed in-house, the data has never left an offline server and the tool is also only accessible locally.

  1. Why? If IA is local, nothing to do with the web, whatever the version. And there are AIs that use web3 (not being centralized doesn’t preclude any functionality, it’s “just” the way it works that needs to be different).
  2. To draw a parallel: when you use the search function in Anytype, the search engine also needs to access your data (and at least understand it if it’s to be fast and efficient). Is this a bad thing, or does it mean you lose the privacy of your data?

Here is the answer provided by GPT.

The difference between my brother and me is: I need artificial intelligence, but I don’t need Anytype to have AI.
I use Anytype to solve data security issues (otherwise, I would just use Notion).

As previously said, there is a path where Web3 is align with AI implementation.

But you’ll just have to not use it. What’s the issue here?

No problem.

I have previously deployed a local large model, and I found that it has certain requirements for computer configuration, and it isn’t as smart as the online version. If an AI is running locally but isn’t powerful, what is the point of adding it?

Of course, it’s also possible that I don’t understand the relevant technology, and I look forward to someone providing a detailed explanation.

Local LLMs use a smaller scale of parameters, they are not so smart than online ones, but they can be helpful too. For example, AI can be used as an automation tool, which previously might have required specialized plugins or cumbersome operations to accomplish.
Like if you create a new template and want to apply it to existing notes, you could say, “Help me change all notes using the xxx template to this new template.”
Another example could be, “Find notes from xx to xx date, filter out the content related to XXX, and create a new collection.”

you can also ask chatgpt to compare oranges and apples and you will get a similar “sophisticated” answer. but if it adds something to the matter of if it even makes much sense is another question.

There are many other improvements that I would prefer to see accomplished. There are numerous opportunities to use AI that can augment a user now both locally and online. Stay the course build the best local first privacy oriented note taking app and you may find the OS itself ends up bringing the AI to the table in a possibly superior way.

For those who are worried about privacy using AI.
Perhaps there is a compromise in the form of a functional limitation: explicitly set the context for the AI.
For example: use AI to record a specific tag, for a specific object, for a specific collection, and so on.

This way, the user will be able to decide what information they can share with the AI, leaving sensitive information out of the AI’s field of view.

I second your questioning and I think that it would be highly contradictory if Anytype, which brands itself with “trust”, “autonomy”, “security”, and privacy, would implement any kind of tools based on generative AI/LLMs. I see e.g. that a so-called “AI assistant” is listed as “highly requested feature” and that there’s already an AI tool in the research within the documentation website. This is quite concerning to me.

Even if it would be “local” only, generative AI has been trained through highly unethical ways (both in terms of labor conditions and of stealing intellectual property) and has tremendous environmental impacts. It is also proven to be often inaccurate (so-called hallucinations), and thus ineffective. Basically, implementing generative AI is part of the trend towards enshittification we’ve been seeing in IT products in the last few years.

For me it is incredibly difficult to trust a developer that implements a generative AI tool in their products.

If Anytype developers would ever decide to add such a feature (which I would strongly oppose), it needs to be optional and opt-in only (not turned on by default).

I agree with much of what you say.

On the other hand there are also other ways to implement LLMs. Simple tasks could get automated (like auto suggesting tags, types based on the content, asking a model to summarize journal entries of the last week…). This all could run locally with a model that is trained reasonably. Not every hardware would support it.
With such tedious tasks LLMs are quite efficient if the corpus of the data is restricted to just what you create.

The Anytype team couldn’t even be bothered to respond to customers’ concerns in this post.

as long as no one knows if and how they would integrate it there is no reason to be concerned.
when it’s based on a local mcp with ollama and a model of your choice - is there even something to worry about?

Hi Kadiz, welcome to the community! Glad to have you with us. With respect towards you, I’m trying to be really objective and rational and I just can’t seem to see how this makes sense…

You said it’s concerning that AI is potentially coming to Anytype but the quoted sentence (and your whole comment) doesn’t tell me why you think having AI is contradictory to Anytype’s branding.

I know that it being “local” is not your concern as to why Anytype, ‘which brands itself with “trust”, “autonomy”, “security”, and privacy’ shouldn’t onboard AI. Unless it is please let me know.

I don’t see how OpenAI hiring Kenyan workers has anything to do with Anytype. You don’t even know if Anytype is planning to use ChatGPT yet). Even if they are, how does this have anything to do with Anytype?
The roads I drive on were built by people getting paid way less than me for working way harder than me under the hot sun while I drive in my air conditioned car, with that logic I shouldn’t use roads then.

Again, not sure how this has anything to do with Anytype. Also, “stealing” is not the right word here but that would go into a whole different discussion.

With this logic, we shouldn’t be using our computers because it needs to be charged and that electricity is coming from somewhere that is negative to the environment (unless you are self sustaining via wind turbine, solar panels etc.). Anytype shouldn’t exist because it’s using servers under AWS and those servers are a net negative directly to the environment…

Just want to make sure that we know AI right now is only in its amoeba stage, right? It’s like someone complaining about the “world wide web” in 1994. Let’s be real, AI is inevitable. Not wanting to use AI down the line will be like those uncontacted tribes that rejecting contact with civilization. Do you think 30 years from now you still wouldn’t be using AI?

Would love to get some examples of this. All the enshitification that I can think of are done by human decision…

To wrap it up, I don’t see how any of the backup points you made has got to do with convincing us that local AI/LLM would go against Anytype’s privacy…

I have the same notion with @Shampra and I voiced it out half a year ago. AI is incredibly useful for a PKM, especially for a PKM that I’ll be using for the rest of my life. Imagine all the information you’ve gathered in 40 years (2065), and you trying to remember something you wrote into your Anytype second brain 30 years ago in 2035, are you gonna spend 30 minutes looking for it or are you gonna get AI to pull it up for you and help you remember the context as to why it was written down 30 years ago? I willing to bet it’s the latter.
But as of right now, AI/LLM shouldn’t be the focus in my opinion, especially not when there are so many other features that Anytype has yet to implement that are much more needed. This has been my stance since last year and I still stand by that.

I really can’t understand the panic that some users have. although I value privacy a lot!

You can buy for about 50,- Euro a microcontroller board without any internet connection, that can do for example face recognition and more.
It has the AI local on it.

If even a cheap microcontroller board is able to do some AI functions locally, why should it be a problem if a PC App has such a 100 percent local solution integrated?

What bothers me more is the fact that Anytype does internet connections on it’s own and there is no switch to disable that all completely if there is need for it.

  • There is the syncing
  • There is telemetry
  • There is the connection to websites if the mouse hovers over a web link.

As long as I’m able to disallow AI to do any internet connection, I have no problem with it.

Depends on how far they want to play the game. If it’s just auto suggesting tags or relations, that should be doable.
The micro computer example is using a highly specialized algorithm while generative AI, like a large language model, needs a lot of tokens if it should be able to generate content from all kinds of text content that it needs to “understand” first. Then there is also the context window that can be kind of small even with commercial models. It would be very hard to give an LLM the whole Anytype vault all the time.

For tedious tasks there are possible solutions with local AI. But also not many computers these days meet those requirements to run those.

Have a look what this AI can do:

It generates videos.
It runs local and offline.
And it needs only 6 GB graphic RAM.

If this is possible, then there should be less problems to do something with texts.

But I would like to have an AI tool that can not only do stuff with text, but can also recognize things on the images in my Space, so that I can ask it:
“Find each image in which you recognize a mushroom, tag these images with “Mushroom” and then create a Query for them.”