Overview
Large Language Models (LLMs) have transformed natural language processing, but their reliance on vast, diverse datasets raises critical concerns around privacy, ethics, and compliance with regulations like GDPR and CCPA. As AI adoption grows, so does the need for machine unlearning—the ability to systematically erase specific data from models to uphold privacy, mitigate biases, and ensure ethical AI deployment.
This research explores the challenges, techniques, and implications of machine unlearning in pre-trained LLMs, an area that remains largely unaddressed compared to fine-tuned models. Through an extensive literature review, the study identifies and evaluates existing unlearning methodologies, such as data-centric, model-centric, graph-based, and federated unlearning approaches, analyzing their effectiveness, limitations, and trade-offs.
A key focus of the project is proposing hybrid unlearning techniques that balance efficiency, interpretability, and computational feasibility. Additionally, it outlines practical frameworks for operationalizing unlearning, including standardized evaluation metrics, governance strategies, and adaptive mechanisms to support AI systems that can forget responsibly.
By bridging technical research with real-world applications, this study contributes to the growing body of knowledge on AI safety, responsible model development, and compliance with evolving data protection laws. It serves as a foundation for future advancements in AI governance, ensuring that LLMs remain transparent, adaptable, and aligned with ethical principles.
Motivations
My passion for machine learning and artificial intelligence has driven me to explore the frontiers of responsible AI development, where innovation must be balanced with ethical considerations. The rapid advancement of large language models (LLMs) has unlocked groundbreaking possibilities but also raised pressing concerns around data privacy, bias mitigation, and regulatory compliance. As someone deeply invested in both research and industry, I am particularly drawn to the challenge of ensuring that AI systems not only learn effectively but also possess the ability to forget responsibly—a capability that is essential for building ethical and transparent AI.
Throughout my academic and professional journey, I have sought the intersection between research-driven innovation and real-world applications. While I thrive in fast-paced environments, I have always been drawn to deep, analytical problem-solving—the kind that requires careful thought, rigorous evaluation, and a commitment to pushing boundaries. However, I have also witnessed firsthand how industry demands often prioritize rapid deployment over critical reflection, making it difficult to tackle fundamental AI safety issues like machine unlearning. This research project, therefore, represents my commitment to bridging this gap—to contribute meaningful solutions that address these concerns while maintaining the efficiency and scalability required in industry settings.
Through this study, I aim to explore effective unlearning techniques that preserve model utility while mitigating risks associated with data retention. By investigating a combination of data-centric, model-centric, and hybrid unlearning methods, I seek to develop practical frameworks that organizations can leverage to implement responsible AI governance. More broadly, this work aligns with my long-term aspiration to be at the forefront of AI research, collaborating with institutions and industry labs that are pioneering efforts in ethical AI and advanced machine learning techniques.
This thesis is not just a research endeavor; it is a stepping stone toward my larger goal—to contribute to the evolving conversation on responsible AI and establish myself as a researcher who is not only technically proficient but also attuned to the broader societal implications of AI systems. By integrating machine unlearning into LLM development, I hope to push AI forward in a way that is both innovative and accountable, ensuring that the future of AI remains as adaptable and ethical as it is intelligent.
Reflections and Learnings
Completing my thesis on Machine Unlearning in Large Language Models has been an immensely rewarding and intellectually stimulating journey. This research allowed me to explore the intersection of privacy, AI ethics, and model governance, areas that are becoming increasingly critical as AI systems continue to expand in influence and responsibility. Throughout this process, I deepened my understanding of large language models (LLMs) and their inherent challenges, particularly in the realm of data removal, compliance, and technical feasibility.
Receiving a perfect score (100/100) and high praise from my supervisor and academic committee was a deeply gratifying moment. The feedback emphasized the strengths of my work—its structure, clarity, depth of research, and practical applications—reinforcing my confidence in my ability to contribute meaningfully to the field of AI research. I was particularly proud of how my thesis was recognized for connecting complex, multi-faceted challenges and offering pragmatic solutions that could guide future research and industry applications.
However, the feedback also presented valuable insights for future improvement. A key takeaway was the suggestion to expand on the taxonomy of machine unlearning, categorizing different types of data and refining accountability mechanisms for each case. This pushed me to reflect on how machine unlearning is not a one-size-fits-all solution—rather, it requires tailored approaches depending on the data type, industry, and regulatory context. This is an area I plan to explore further, whether through academic research, industry applications, or collaborations.
Additionally, the thesis defense was a pivotal learning experience. It reinforced the importance of articulating complex concepts clearly and confidently, adapting explanations based on the audience's background. The positive feedback from both my thesis supervisor and the Associate Dean reassured me that I had not only developed strong research but also demonstrated mastery over the subject matter in real-time discussions. This experience has strengthened my ability to defend and expand on my work under critical scrutiny, a skill that will be invaluable in future research, public speaking, and professional engagements.
Key Learnings and Next Steps
Bridging Research with Real-World Applications – While this thesis was deeply theoretical, I realized the importance of practical, industry-relevant implementations. Future work should focus on applying unlearning techniques in live AI systems and assessing their scalability, efficiency, and ethical implications.
Refining the Taxonomy of Machine Unlearning – Given the complexity of this issue, a more granular classification of different unlearning scenarios—covering types of data, regulatory constraints, and technical trade-offs—could help standardize the field and guide both policymakers and AI practitioners.
Developing Accountability Mechanisms – One of the most important open challenges is ensuring that unlearning is verifiable and auditable. My research touched on governance strategies, but further work is needed to develop robust accountability frameworks that align with AI ethics, corporate compliance, and legal regulations.
Strengthening Thought Leadership and Research Contributions – This thesis reinforced my desire to continue contributing to AI ethics and governance research. I plan to expand on this work through publications, conference talks, and collaborations with academia and industry.
This research experience has been a defining milestone in my academic and professional growth. It has equipped me with a strong foundation in technical research, critical problem-solving, and thought leadership in AI governance—all of which will shape my next steps as I continue advancing in the field of AI.
Feedback Received
Grade: 100/100
Feedback on Thesis Submission
Fantastic work. Very serious scholarship on this topic.
Some thoughts:
- Very well written and organized
- Well sourced and researched
- Practical, pragmatic suggestions for people continuing on this research
- This problem is soooooo complex. You did a great job connecting a bunch of the dots.
- It's an incredibly important question as LLMs do more important work for businesses, organizations and individuals in the future
- I'd suggest digging deeper into the various use cases and perhaps building a more thorough taxonomy of types of unlearning - what types of data, what unlearning would technically mean in each case, accountability mechanisms, etc
I'm proud to be associated with this high caliber work. Thanks.- Geoff Nesnow, Supervisor
Feedback on Thesis Defense
Camila clearly understands the subject well. She did an excellent job explaining the key concepts, the challenges in the field, its overall significance and new areas for additional exploration.
Overall, she did an outstanding job defending and expanding on her work.- Geoff Nesnow, Supervisor
A very impressive piece of work, and she did a great job in answering any questions, explaining her conclusions, and considering areas to expand on in the future.
- Karl Petrick, Associate Dean


