I'm a PhD Student at MBZUAI, working with Prof. Monojit Choudhury. I like working on problems with heavy social impact. I am currently interested in AI Alignment, with particular interests around representation geometry and issues related to inner alignment, this includes interpretability based audits, science of evals, and RL algorithms less prone to goal misgeneralization and reward hacking. I've also recently started working on human simulators, A.K.A Cognitive Digital Twins (CDTs)!
Here's a running, non-exhaustive list of my current research questions. I also frequently share opinions and updates on X and Substack. I also love music, philosophy and sociology. If anything here interests you, feel free to reach out!
News
- Aug 2026- Our paper on Cognitive Digital Twins is accepted to AIES 2026!
- July 2026- Attending ICML 2026 and The Alignment Workshop by FAR.AI!
- Jun 2026- New arXiv preprints on efficient safety benchmarking and cognitive digital twins.
- Jun 2026- Two mentored papers accepted to AI4GOOD at ICML 2026!
- Jan 2026- Started mentoring two AI Safety projects in SPAR.
- Jan 2026- Gave a guest lecture at the course Responsible and Safe AI Systems, IIIT Hyderabad.
- Dec 2025- Attending IndoML 2025, catch me there!
- Nov 2025- Gave a talk on AI Safety at NIMHANS Bangalore (LinkedIn post).
- Oct 2025- Reached 100+ citations on Google Scholar.
- Sep 2025- Published Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety at NeurIPS 2025 (Reliable ML and Regulatable ML Workshops).
- Aug 2025- Started my PhD!
- Jun 2025- Finished my internship at the Center for Human-Compatible AI (CHAI), UC Berkeley.
- May 2025- Journal paper accepted to TALLIP.
Education
-
CS PhD (2025-Present)
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, UAE -
BTech + MS in Computer Science and Computational Linguistics (2021-2025)
International Institute of Information Technology, Hyderabad