ⒾMachine translations by Deepl

GPT-NL: demonstrating that privacy-friendly AI is indeed possible

Then GPT-NL The aim of entering the Dutch Privacy Awards was not simply to gain recognition for a technically innovative project. Far more important was the opportunity to demonstrate that privacy need not be an obstacle to innovation, but can, in fact, serve as a starting point for better technology. That message is proving to be more relevant than ever.

Privacy as a starting point, not just a box to tick

For the organisers of GPT-NL, entering the Dutch Privacy Awards came at just the right time. The project had originated from a desire to investigate how generative AI can be developed within the framework of privacy legislation, copyright law and European values.

“Privacy is often seen as a separate issue”, says the team. “But in reality, it affects virtually everything: innovation, trust, digital autonomy and the way we use technology.”

The Privacy Awards provided a platform to bring that story to a wider audience – not only within the privacy community, but also beyond it. It was precisely that dialogue between developers, lawyers, policymakers and regulators that proved valuable.

What is actually allowed?

One of the key lessons from GPT-NL is that the debate surrounding AI all too often focuses on what is not permitted. During the development process, the team regularly came up against guidelines that are difficult to apply in practice to innovative AI projects.

Rather than concluding from the outset that certain applications are not possible under the GDPR, GPT-NL opted for a different approach: to investigate what is actually possible within the legal framework. This involved a thorough examination of data selection, filtering and risk mitigation.

This approach also caught the attention of regulators. For example, the Dutch Data Protection Authority cited GPT-NL as an example in discussions about the possibilities of generative AI within the framework of privacy legislation. The French data protection authority, the CNIL, also showed interest and adapted its guidelines in light of the practical experience gained from the GPT-NL project.

Every dataset is a trade-off

Much of the work lay not in building the model itself, but in carefully compiling the training data.

Time and again, the team asked the same question: what data is really necessary? For example, a conscious decision was made to remove email addresses and telephone numbers from the training data. After all, these add little to a language model’s actual knowledge.

Conscious decisions were also made regarding public figures. GPT-NL developed special filters to determine which information could and could not be included. In doing so, a balance was sought at all times between usability, quality and the protection of privacy.

Moreover, that process is never-ending. As new insights emerge, filters are refined and datasets are reassessed. Sometimes data turns out to have been unintentionally excluded because a filter was set too strictly. According to GPT-NL, it is precisely this ongoing evaluation that makes responsible AI development possible.

The temptation to make concessions

As with many AI projects, there were times when it seemed tempting to abandon the original principles. After all, more data often means better performance.

Nevertheless, the project team remained true to its core principles. According to those involved, it helped that GPT-NL was developed within TNO. This meant there was scope to make decisions based on more than just commercial considerations.

It is striking that there was never any serious discussion within the project about lowering the privacy standard. However, there were ongoing discussions about the best way to apply that standard in practice. This involved active collaboration with experts and regulators.

More than just privacy

Although privacy was the main focus, GPT-NL ultimately revolved around a broader ambition: digital sovereignty.

Europe wants to reduce its dependence on technology from a handful of major international suppliers. GPT-NL demonstrates that it is possible to build up in-house knowledge and expertise in advanced AI systems.

The organisers emphasise that GPT-NL is not intended to compete with the largest international frontier models. Its strength lies precisely in specialised applications for organisations that wish to retain control over their data, processes and AI infrastructure.

Transparency as a strength

A notable aspect of the project is the openness with which it was carried out. For example, a Data Protection Impact Assessment (DPIA) was carried out and published for the entire system.

That makes an organisation vulnerable, but according to GPT-NL, that is precisely the intention. Transparency leads to better discussions, greater trust and, ultimately, better technology.

The recognition The Dutch Privacy Awards played a part in this. The award demonstrates that privacy-friendly innovation is not only theoretically possible, but can actually be achieved.

Call for future participants

The stakeholders have a clear message for future nominees for the Dutch Privacy Awards: look beyond privacy alone.

A successful privacy innovation is not just about technology or legal compliance. The real question is what kind of social impact you want to make and how you can inspire others.

“It doesn’t necessarily have to be a technical innovation. Think about how you can achieve more than the sum of its parts. That’s precisely where real innovation arises.”

Looking ahead

GPT-NL is now focusing on further collaboration with businesses, public authorities and other partners. The aim is not to let an interesting experiment gather dust, but to build a complete ecosystem in which privacy, innovation and European digital autonomy come together.

After all, it’s not just about a language model. It’s about creating practical solutions that genuinely help organisations move forward. And according to GPT-NL, that is precisely where the future of responsible AI in the Netherlands and Europe lies.