Purpose Limited Openness: Normative Guardrails for Open Source AI
The hype surrounding artificial intelligence (AI) seems omnipresent. While AI regulation is at the forefront of policy efforts, data protection law often seems to be taking a back seat. This is evident in the Commission’s current efforts to undermine the fundamental decisions of the GDPR through the “Omnibus Package”: not only is the scope of personal data to be modified, but the draft also departs from the GDPR’s technology-neutral approach.
Specifically, a training privilege for AI systems is to be introduced regarding the processing of special categories of personal data. The proposal establishes a new legal basis for training AI systems with particularly sensitive data without restricting their purposes in any way. However, we argue the opposite direction, namely that purposes should be clearly defined and regulatory guidelines should be established. Powerful AI systems pose significant societal risks, ranging from privacy violations and discrimination to the monopolisation of corporate power and the exploitation of labour. These systems rely on vast quantities of data extracted from millions of individuals—including surveillance, transaction, and location data—as well as the “hidden” labour required to annotate and moderate digital content. The production of AI models has created a profound power imbalance. Because these systems can be easily scaled across diverse sectors—from healthcare to hiring—a tiny number of dominant models now influence vast areas of human life.
There is no doubt that data protection law is reaching its limits in its regulatory focus on individualised data processing and specific rights of data subjects in mass data processing, such as AI. Nevertheless, data protection law provides suitable instruments for effectively regulating AI systems. These include the principle of purpose limitation, which we argue, is more important than ever.
Purpose limitation as a fundamental principle of data protection
Purpose limitation is a cornerstone of data protection. The GDPR’s purpose limitation principle does more than protect individual privacy; it promotes systemic transparency and fairness. By including specific exceptions (Articles 5 (1b) and 89 (1) GDPR), the law explicitly prioritises public interest activities—
like scientific research and statistics—over commercial data processing. The purpose limitation principle prevents the—potentially endless—unauthorised reuse of collected data, which would otherwise lead to repeated privacy violations, erosion of autonomy and control. Purpose limitation ensures that processing remains transparent to the individual and verifiable by regulators, additionally reflecting reasonable expectations of privacy of the data subjects.
We have argued elsewhere, that purpose limitation can function as a powerful lever in AI governance. While traditional ‚purpose limitation‘ focuses on protecting the individual who provided the data, we argue for a broader application: purpose limitation for AI models. Our goal is to prevent the data of one person from being used to build tools that harm others, entire groups, or society as a whole.
We believe that data subjects provide their information with the implicit expectation that it will not be weaponised against the public good.
Problems of secondary misuse of data sets and AI models
Unfortunately, the purpose limitation principle is widely ignored in practice. In the case of AI and GDPR compliance, massive datasets and complex AI supply chains make purpose limitation practically unenforceable; the more hands the data passes through, the harder it is to track its original purpose. Data-intensive technologies such as AI provide unprecedented opportunities for the secondary use of data, in particular for special categories of personal data. The case of the UK Biobank illustrates that the GDPR provisions addressing the processing of personal data are alone not sufficient: thousands of people donated their health data for medical research in the public interest. Despite promises not to share this data, it was anonymised and shared with commercial insurance companies.
Or see another example: An AI model designed for medical research—such as diagnosing psychiatric disorders from behaviour—becomes a tool for harm when repurposed to screen job applicants.1 When data is anonymised for the training of AI models, which is often the case, the data processing falls outside the scope of the GDPR and therefore the bond of purpose limitation is cut.
These cases of secondary re-use or misuse of datasets and AI models undermine the regulatory decision to privilege only certain types of secondary use of data—namely, those for the public good under Article 89 of the GDPR. Furthermore, data subjects have no possibility of preventing their data from being used for other purposes after anonymisation. This also erodes trust in public institutions and fundamental research.
Open Source at odds with purpose limitation?
The concept of openness seems inherently at odds with the principle of purpose limitation, which is meant to restrict secondary use rather than open it further. Open source software and “open” AI have been discussed as opportunities to “democratising AI”, to promote scientific innovation, public control and sovereignty. Open source and open data communities have historically championed unrestricted accessibility as a service to the common good.
However, with the emergence of powerful AI technologies, the legal and ethical landscape has changed. Widder et al. have argued that openness alone fails to challenge AI’s power imbalance. Much like traditional open-source software was absorbed by Big Tech, the rhetoric of “open” AI is often used to strengthen, rather than dismantle, corporate monopolies.2 Furthermore, open AI models can undermine the controllability of data processing and the use of AI systems. Huang et al.3 demonstrated the risks of AI “function creep” by adapting Meta’s general-purpose speech model, wave2vec, into a high-precision psychiatric diagnostic tool. By fine-tuning the model on a dataset of therapy sessions, they created a tool that identifies depressive symptoms from just 60 seconds of audio. While this has clinical potential, it also opens the door for “emotional surveillance” if companies use similar tech on sales calls or support hotlines to secretly filter out individuals. Using a publicly licensed, open-source dataset the researchers achieved significant prediction opportunities.
These scenarios illustrate a “secondary misuse” of opensource AI: tools and datasets created for a beneficial purpose (Context A, like healthcare) are redirected toward harmful or ethically suspect commercial uses (Context B). Strikingly, because the models are open and the data is often anonymised, this type of repurposing usually bypasses current AI and data protection laws entirely.
In their interdisciplinary collaboration, Prof. Hannah Ruschemeier (law) and Prof. Rainer Mühlhoff (ethics and philosophy of technology) focus on the structural and societal implications of data-driven technologies. Since 2021 they have been developing legal frameworks and ethical theories such as Predictive Privacy and Purpose Limitation for AI Models, see https://purposelimitation.ai.
In their project “Purpose-Limited Openness”, funded by the Volkswagen Stiftung, Germany, they develop purpose-limited open data and open AI licenses: a framework that seeks to reconcile the ideals of open science and opensource AI with safeguards against harmful or unintended reuse. Inspired by Creative Commons licensing, these licenses would allow open access while embedding clearly defined purpose limitations designed to prevent harmful secondary uses. Combining ethical analysis, legal design, and collaboration with civil society actors, the project explores whether such licensing frameworks could foster a culture of responsible openness in the AI ecosystem.
Merging purpose limitation and openness in AI Governance
Neither the GDPR nor the AI Act (Regulation (EU) 2024/1689) address these problems adequately. On the contrary, the AI Act provides regulatory exemptions for open-source models (Article 2(12)) and introduces the category of “general purpose AI systems” (Article 3(67)). Models for general use are a problematic category from a regulatory perspective: on the one hand, these models always serve a purpose, e.g., aligning with the economic interests of their developers; on the other hand, this category leads to difficulties in assigning responsibility and enforcing and verifying compliance.
In response to concerns about the secondary utilisation of AI models and open-access datasets when transitioning from their foundational development environments, we propose the concept of purpose-limited openness. Conceptualising the collective societal risks that arise from the secondary misuse of trained AI models and anonymised data sets requires a shift beyond the individualistic paradigm that dominates contemporary data ethics, data protection law and digital regulation towards public and societal implications. While notions such as group privacy4 or the right to reasonable inferences5 have attempted to move toward collective concerns, they remain insufficient for addressing the full ethical and political implications of AI in open data environments. Our project starts from the insight that the circulation and reuse of anonymised models represent not just a privacy issue, but a broader socio-political problem involving power asymmetries, bias reinforcement, and informational exploitation.
We aim to establish the theoretical underpinnings for mitigating the systemic power imbalances arising from the unmonitored exploitation of AI architectures and datasets initially conceived to advance the public interest. A primary challenge identified is the phenomenon wherein AI models and data assets, originally curated for a high-integrity, beneficial environment—characterised as Context A (such as clinical medical research)—are subsequently co-opted and deployed within a disparate, frequently profit-driven Context B. In these secondary environments, the application of such technologies may lead to ethically ambiguous outcomes or result in direct societal harms.
Drawing on the normative frameworks of law and ethics, our aim is to construct rigorous theoretical basis for constraining the asymmetries of power arising from the non-discretionary application of public-interest AI tools across external domains. Furthermore, we plan to align this normative approach with the ethically grounded principles of the Open Source Software (OSS) by operationalising these values into robust, scalable governance architectures. Through a rigorous, multi-stakeholder collaboration involving the opensource and communities, as well as academic the objective is to engineer a pragmatic implementation strategy in the form of a specialised licensing framework for open-source AI. This licensing mechanism is designed to serve as a functional safeguard, ensuring that the evolution and deployment of these models remains strictly tethered to the promotion of the public good.
Outlook
The current direction of the Omnibus proposal—specifically the introduction of a general privilege for AI training—is a significant step backward for the protetion of fundamental rights and societal interests. By granting broad exemptions for AI development, the Commission risks codifying a “technology-
first” approach that hollows out the GDPR’s core protections exactly when they are most needed. Rather than treating AI as a special category exempt from scrutiny, regulation should establish a framework that regulates AI based on its specific purposes and its demonstrable contribution to the public good. Only by tethering the reuse of models and datasets to high-integrity, beneficial contexts can we prevent “openness” from becoming a tool for corporate power concentration and secondary misuse.
The authors are researching these issues as part of a joint project at the University of Osnabrück. The project, titled ‚The Dangers of Open AI: Towards Ethical Data Sharing Through Purpose Limitation‘, is funded by the Volkswagen Foundation as part of the Aufbruch funding programme.
1 See for a detailed discussion of this example: Mühlhoff/Ruschemeier, International Journal of Law and Technology, Volume 33, 2025, eaaf003
2 Widder, D.G., Whittaker, M. & West, S.M. Why ‘open’ AI systems are actually closed, and why this matters. Nature 635, 827–833, 2024
3 Huang, X., Wang, F., Gao, Y. et al. Depression recognition using voice-based pre-training model. Sci Rep 14, 12734, 2024
4 Mittelstadt 2017, From individual to group privacy in big data analytics
5 Wachter/Mittelstadt 2019, A Right to Reasonable Inferences: Re-Thinking Data Protection Law in the Age of Big Data and AI.