Trusted AI

AI & Privacy: How to protect your data?

Data protection has been at the heart of the digital sphere for years. While the GDPR has helped regulate personal data, its collection, and the security of its storage, the arrival of AI is reshuffling the deck. This technology expands the "attack surface" and creates new privacy challenges.

⚡️ TLDR

  • Privacy vs Confidentialité dans l'IA : Distinguer la protection des données personnelles (privacy, encadrée par le RGPD) de la restriction d'accès aux informations sensibles (confidentialité, requise par l'AI Act pour les modèles à haut risque).
  • Impact des fuites de données : Au-delà des sanctions financières et réglementaires pouvant représenter un pourcentage significatif du chiffre d'affaires, la conséquence la plus critique est la perte de confiance des utilisateurs.
  • Les 4 piliers de la sécurisation data :
    1. Privacy by Design : Intégrer la protection de la vie privée dès le cadrage du projet.
    2. Minimisation des données : Ne collecter que les informations strictement nécessaires à l'usage métier.
    3. Differential Privacy : Injecter du « bruit » mathématique pour empêcher les attaques par réidentification.
    4. Souveraineté : Héberger la donnée sur des serveurs européens.
  • Responsabilité partagée : La sécurité exige une collecte raisonnée et transparente de la part des entreprises, associée à une stricte hygiène numérique des utilisateurs (ne jamais saisir de données confidentielles dans les prompts des IA génératives).

To discuss this topic, we were pleased to host a roundtable featuring: 

  • ​Paul Mangold - Assistant Professor - École Polytechnique
  • ​Matthieu BOUSSARD - VP AI - Craft AI
  • ​Henri CLAUDOT - Director, Innovation & Collaborative Programs - IDEMIA Public Security
  • ​Isabelle LANDREAU - Chief Privacy Officer, Group Data Protection Officer - IDEMIA Public Security
  • ​Jean-Philippe Clement - Deputy Director General of Services in charge of public space and facilities - City of Paris 

But what exactly is data? 

We often hear about the need to protect our data, data breaches, or the GDPR... But what is data, really? 

Data is simply information (raw or a collection of information) that can be stored and then processed to meet various objectives.  

Today, each of us provides our information to hundreds of major players who store and process it (data only has value if it can be processed or refined). 

New challenges in data protection

In recent years, artificial intelligence has continued to advance with increasingly powerful models.
But this technology doesn't progress by "magic"; AI models aggregate data and are trained on it. Furthermore, generative AI and its varying levels of controlled adoption reinforce the need for data security. 

Data protection, privacy, confidentiality... What’s the difference? 

Today, when we talk about data, many terms and regulations come to mind, but they don't all mean the same thing: 

  • Privacy privacy refers to personaldata. Be aware that not all countries define personal data in the same way (for example, some countries do not consider religion to be personal data, while others do), and data is only no longer considered personal if it is truly anonymized (meaning it is impossible to re-identify the person later).
  • Confidentiality confidentiality is the practice of ensuring that information is accessible only to authorized individuals (it is not made public). However, this concept is not defined in the AI Act (the European regulation designed to govern the use of AI); only a few articles are devoted to it, and they concern high-risk regimes (patents, trademarks, sensitive information) for generative AI. This highlights the importance of privacy by design , which involves integrating privacy protection from the very first stages of a project's creation, rather than as an afterthought. However, organizations that "create" AI solutions are required to inform users about how their data is used in the interest of transparency (regardless of whether the AI is high-risk or not), in accordance with GDPR regulations. 

The importance of data in AI models and its key applications 

Securing data is even more critical in the context of an AI project because, without it, AI would not function. Data is essential in many scenarios: 

  • For example, data from the company "Pay by Phone" (a mobile parking payment solution) is used to indicate available parking spaces. In fact, 25% of traffic in these areas is actually caused by people searching for parking... This is an effective way to reduce and streamline traffic. 
  • For Idemia's Entry/Exit system, installed at 150 border crossing points, biometric and personal data are used for security purposes to implement a system that identifies who is authorized to leave the Schengen Area. 

What are the risks in the event of a data breach? 

A database containing millions of personal records is a prime target for cyberattacks... But what are the consequences for the organizations that hold this data? 

A data breach can have serious consequences for a solution's users (identity theft, alteration of information, etc.). The AI Act provides for sanctions against organizations that fail to protect their users' data. Beyond economic sanctions (which can reach a significant percentage of the offending company's turnover) and administrative penalties, the primary—and most severe—consequence for these organizations is the damage to their reputation among users (a loss of trust and perceived security). 

What is the best way to secure data?

To prevent data leaks and malicious use, every company that aggregates data must secure it in several ways: 

  • Privacy by design: by placing data protection at the heart of system design 
  • Using only the data that is truly necessary for the solution (for example, there is no need for a high-definition camera just to know if a parking space is full or empty) and finding a compromise between privacy and the solution's capabilities. This is more precisely known as data "minimization." 
  • Hosting collected data on European servers. 
  • Differential privacy: This mathematical technique involves injecting slight "noise" (adding random false information to mask an individual's real data) into the dataset to prevent "re-identification" attacks. Even by cross-referencing multiple databases, a hacker will not be able to isolate a specific user's profile within an AI model.

In conclusion 

Data is useful for AI agents in many ways, but both organizations and users must follow best practices to ensure the secure use of a solution. 

For companies, robust system security, "reasoned" data collection, and clear communication with the public regarding data usage are essential. Meanwhile, users must practice good "digital hygiene" and use these solutions (especially generative AI) mindfully by not openly sharing personal information in their prompts. 

To launch your secure AI project contact our experts!