Research
Data for AI
AI is only as good as the data it learns from and retrieves. This area covers data quality and readiness, synthetic data, governance and lineage for AI, and privacy and sovereignty obligations.
Topics in Data for AI
Privacy & Data Sovereignty in AI
Using personal data lawfully and keeping it in the right jurisdiction.
Overview
Many AI projects slow down once teams discover that the data they need is incomplete, inconsistent, poorly labeled or not approved for the intended use. Generative AI adds unstructured content such as documents, emails and tickets, which most governance programs were not designed to handle.
Data for AI is therefore both a technical and a governance problem. Teams need to know what data exists, whether it is fit for purpose, where it came from and whether they are allowed to use it.
What leaders need to get right
- Readiness. Assess data quality against the needs of specific use cases.
- Permissions. Make sure AI tools respect existing access controls.
- Lineage. Trace training and retrieval data back to its source.
- Lawful use. Check privacy, consent and residency before data enters an AI system.
Questions leaders should ask
- Which data is approved for use in AI, and is that enforced?
- Can we trace the data behind an AI output back to its source?
- Would our AI tools expose documents to people who should not see them?
- Where is the data we send to AI services processed and stored?
Other research areas
Latest insights on Data for AI
New articles in this area are on the way. In the meantime, explore the Research library.