AI is only as good as the data it can safely use. GovTech's Data Practice enables responsible data use across the Whole of Government, letting agencies share and build on sensitive data for AI without putting citizen trust at risk. We do this by advancing Privacy-Enhancing Technologies (PETs) such as advanced anonymisation, synthetic data generation, and federated computing, and by tackling the new privacy risks that emerge as AI systems become agentic.
You will work on real problems in AI and data privacy, on projects such as:
Multi-turn privacy leakage in AI systems. As agencies deploy LLM assistants and agents that accumulate context through conversation history, retrieval, and memory, sensitive information can leak across interactions in ways single-turn safeguards miss. This is an open problem with no established solution in industry. You would quantify how the risk builds up and evaluate mitigations such as memory controls, context isolation, aggregation thresholds, and runtime guardrails.
Image anonymisation. There is strong demand, particularly from healthcare, to use image data safely. OCR is mature, but a hard gap remains in accurately determining what in an image should be anonymised and cleanly replacing it. You would develop and evaluate detection and anonymisation approaches, with potential to work directly with a real pilot user.
Learning Outcomes
- AI privacy in practice: Understand how privacy risks arise in modern AI and agentic systems, and help tackle problems at the edge of current industry practice.
- Applied research and experimentation: Design and run experiments to measure privacy risk and evaluate mitigation or anonymisation approaches against real constraints.
- Hands-on PET development: Build working knowledge of privacy-enhancing technologies and how they enable responsible data use across government.
- Prototype to pilot: See how a privacy solution moves from research to real deployment, with potential exposure to a live pilot user.
- Technical communication: Translate findings and trade-offs clearly for engineering, product, and policy audiences.
Prerequisites
- Currently pursuing a degree or diploma in Computer Science, Data Science, or a related field
- Proficiency in Python
- Foundation in machine learning, with familiarity with LLMs or computer vision
- Ability to read and synthesise technical research in a new domain
- Interest or prior exposure to AI privacy, data privacy, or PETs is a plus but not required