n8n's “We Don’t Train on Your Data”, With a Footnote
Worth Knowing
n8n kept its headline promise: it will not use your content or the AI outputs to train machine learning models. What changed, across two documents updated the same week, is what n8n can do with your data short of training on it. The AI Terms now let n8n use your prompts, in aggregated and de-identified form, to improve its AI services, and the rewritten Self-Serve terms let n8n derive de-identified datasets from your content to run and develop the platform.
In human terms: You chose n8n partly on the strength of a clean no training promise. That promise holds. But the prompts you send can now feed product analytics once de-identified and aggregated, and the content you run through the platform can be turned into de-identified datasets that help n8n build its product. Training the model on your data is still off the table. Mining around it is not.
Why this matters: This is the clearest illustration of the week of why a “we don’t train on your data” badge is not the end of a review. Training a model and mining prompts and content for de-identified analytics are different activities, and a policy can forbid the first while permitting the second. The questions that actually matter are whether prompts are retained, when they are de-identified, and how far a de-identified dataset built from your content can travel.
In a future deep dive, we'll review the competitive moat that companies like n8n have built, and predict how they may use it: how far down the chain a platform can reach before it is competing with the customers who built on it. For now, read below for the mechanics of this current change:
The mechanics: The AI Terms now provide that “n8n may collect and internally use Usage Data and Customer prompts to improve AI Services, provided such Usage Data and prompts remain in an aggregated and de-identified form.” The prior version permitted only Usage Data that “does not contain Customer Data or identify Customer or its users.” The Self-Serve terms add that “n8n may derive de-identified data sets from Customer Content and may use such derived data to operate, enhance, improve and develop the Cloud Services.”
The no training commitment is retained throughout: “n8n will not use Customer Content or Output to train machine learning models.” One clarification, because it is easy to misread: n8n’s broad, sublicensable license over publicly posted community workflows, and its provision that you have no claim if it independently builds something similar, are carried over from the prior terms. They are not new this week. The new material is the de-identified use of prompts and content described above.