Algorithmic Alignment Group

Researching frameworks for human-aligned AI @ MIT CSAIL.

Team

Principal Investigator

Dylan Hadfield-Menell Dylan Hadfield-Menell, dhm[at]csail[dot]mit[dot]edu, Website
Dylan is an assistant professor on the faculty of Artificial Intelligence and Decision-Making in the EECS Department and Computer Science and Artificial Intelligence Laboratory (CSAIL) at the Massachusetts Institute of Technology (MIT). His research focuses on the problem of agent alignment: the challenge of identifying behaviors that are consistent with the goals of another actor or group of actors. His work aims to identify algorithmic solutions to alignment problems that arise from groups of AI systems, principal-agent pairs (i.e., human-robot teams), and societal oversight of ML systems.

Postdoctoral Researchers

Jakob Stenseke Jakob Stenseke, stenseke[at]mit[dot]edu
Dr. Jakob Stenseke is a postdoctoral fellow and a SERC group leader whose research revolves minds, machines, and morality. He earned his PhD in philosophy from Lund University in 2025, where he examined artificial approaches to morality through the lenses of virtue ethics, game theory, and computational complexity. At MIT, Jakob seeks to combine insights from moral philosophy and cognitive science to inform the design and training of AI systems that better capture the richness of human values and the intricacies of human thought.

Ph.D Students

Aruna Sankaranarayanan Aruna Sankaranarayanan, arunas[at]mit[dot]edu, Linkedin
Aruna is a PhD student interested in using interpretability methods to explain model and human behaviours. Her work utilizes causal interpretability methods to encourage model safety. Her past work has included understanding how humans interface with AI generated content, as well as audits of black box social media algorithms. She loves free and open source software, Tamil Bhakti poetry, and giant trees.

Neil Kale Neil Kale, nkale[at]mit[dot]edu, Website
Neil is a PhD student interested in post-trainability and algorithms for continual learning: teaching ML models to adapt to distribution shift and learn new skills over time. Before his PhD, he completed his Master’s degree at Carnegie Mellon University, advised by Prof. Virginia Smith and Prof. Aditi Raghunathan, and spent a year in industry working on continual learning. Prior to that, he studied CS and Mathematics at WPI, supervised by Randy Paffenroth. He enjoys urban sketching, cafe-hopping, and vibe-coding cool math demos.

Phillip Christoffersen Phillip Christoffersen, philljkc[at]mit[dot]edu, Website
Phillip is broadly interested in reinforcement learning topics including AI alignment, neurosymbolic AI, and multi-agent RL. Before the Algorithmic Alignment Group, Phillip was an undergraduate researcher at the University of Toronto, advised by Prof. Sheila McIlraith. His main hobbies include reading, playing piano, and composing music.

Rachel Ma Rachel Ma, rachelm8[at]mit[dot]edu, Website
Rachel is interested in AI agent alignment to human preferences in open-ended settings, collaborative teaming and interactions between humans and autonomous agents/systems, and decision making under uncertainty. She works within a mix of NLP, vision, and robotics. Before her PhD, she double majored in Computer Science and Music at Brown University and did robotics+NLP research advised by George Konidaris and Stefanie Tellex. Her hobbies include playing piano, composing, watching movies, and spending time with friends.

Yanchen Liu Yanchen Liu, ychenliu[at]mit[dot]edu, Website
Yanchen is a PhD student interested in scalable supervision and scalable oversight, i.e., how to train and oversee superhuman-level AI systems. Before his PhD, he obtained his Master's degree from Harvard University, advised by Prof. Hima Lakkaraju, and his Bachelor's degree in CS + Computational Linguistics from TUM/LMU, supervised by Prof. Hinrich Schütze. He was also a visiting researcher at the Stanford NLP Group. He loves soccer, music, and traveling.

Masters Students

Emaan Khan Emaan Khan, emaan[at]mit[dot]edu, Linkedin
Emaan is a graduate student in the Technology and Policy program. Her current research involves human-centred evaluations for LLM safety post finetuning. Prior to MIT, Emaan completed her undergraduate degree in Computer Science from LUMS, Pakistan where she researched online safety for vulnerable communities on the internet, like children and the visually impaired.

Visiting Researchers

Aidan Kierans Aidan Kierans, kierans[at]mit[dot]edu, Website
Aidan is a visiting student at MIT and a PhD student at the University of Connecticut. His work focuses on normative competence in AI systems, including the institutions that shape model specifications and methods for evaluating and representing moral reasoning. His past research has studied AI alignment from technical and sociotechnical perspectives, including reasoning norms, multi-agent misalignment, moral evaluation, and AI governance. In his free time, he enjoys hiking, travel, and bass guitar.

Henry Castillo Henry Castillo, henryac[at]mit[dot]edu
Henry is interested in robust and reliable ways to understand and edit models, including pretraining interventions, developmental interpretability, and inherent/purposeful design. He double majored in computer science and mathematics at UT Austin and worked as a software engineer at Stripe before AAG. His interests are the top-k subset of west coast tech employee hobbies (climbing, hiking, backpacking, weightlifting, cooking, etc.)

Alumni

Adriano Hernandez Adriano Hernandez, adrianoh[at]mit[dot]edu, LinkedIn
As a large language model :), Adriano is passionate about interpretability and the science of deep learning. His work in the group focused on using insights from activation engineering to improve the robustness of language models such as himself to data poisoning attacks. Before MIT, he was first engineer at Meru, where he built retrieval augmented generation systems for document-based question-answering.

Andreas Haupt Andreas Haupt, haupt[at]csail[dot]mit[dot]edu, Website
Andy completed his Ph.D. in Engineering-Economic Systems in 2024, and was the first Ph.D. student of the Algorithmic Alignment Group.
Now: Human-Centered AI Fellow at Stanford University, and in 2026–27 an AI Institute Fellow-in-Residence at Schmidt Sciences, Digital Fellow at the Stanford Digital Economy Lab, and Technology and Human Rights Fellow at the Harvard Kennedy School.

Ariba Khan Ariba Khan, akhan02[at]mit[dot]edu
Ariba was a Master of Engineering (MEng) student who researched cultural bias and cultural alignment of Large Language Models (LLMs). Her academic interests extend to AI fairness, model debiasing, and interpretability. In her free time, Ariba enjoys listening to music, visiting art museums, and reading historical fiction novels.
Now: a member of technical staff at Crosby.

Deepika Raman Deepika Raman, deepikar[at]mit[dot]edu, LinkedIn
Deepika was a graduate student in MIT’s Technology and Policy Program. Her research interests lie in the responsible and equitable design of emerging technologies and public digital goods. In the group, she worked on participatory approaches to AI across the development lifecycle and challenges with their governance. Before MIT, Deepika worked with the University of Chicago Trust, India - leading teams tackling data governance, digital transformation, and AI policy at the Ministry of Electronics and IT, Government of India.
Now: an AI Standards Development Researcher with the AI Security Initiative, and a Non-Resident Research Fellow at the Center for Long-Term Cybersecurity, UC Berkeley.

Dana Choi Dana Choi, choie[at]mit[dot]edu, Website
Dana is broadly interested in understanding how the human normative system works and what enables cooperation. Through reverse-engineering the mechanisms of human collective intelligence, she hopes to contribute to efforts in designing and facilitating desirable interactions in our society. She draws insights from economics, anthropology, cognitive science, reinforcement learning, and social computing.
Now: an AI policy researcher at the OECD.

Hendrik Mayer Hendrik Mayer, hmayer[at]mit[dot]edu
Hendrik was a Master of Engineering (MEng) student who worked on theoretical alignment problems in reinforcement learning systems. Before joining the Algorithmic Alignment Group, he researched topics in algebraic complexity theory with Prof. Markus Bläser and worked in the Emergent Quantum Matter Group at MIT. His hobbies include playing soccer and hiking.

Jovana Kondic Jovana Kondic, jkondic[at]mit[dot]edu, Website, LinkedIn
Jovana's interests lie broadly at the intersection of probabilistic inference, social cognition, and human-robot interaction. Her research focused on building interactive AI agents that 1) effectively learn from human input, and 2) understand and act in accordance with human preferences, intentions, and values.
Now: a Ph.D. candidate at MIT CSAIL, advised by Aude Oliva.

Julian Manyika Julian Manyika, jmanyika[at]mit[dot]edu, LinkedIn
Julian was a Masters in Engineering (MEng) student who worked on alignment for Large Language Models. Julian’s interests include reward-modeling on reasoning, evaluating language model understanding of human conversation and intent, and the possibility for AI research to help inform the way we think about morality.
Now: a DPhil student in computer science at the University of Oxford, supervised by Jiarui Gan.

Magdalena Price Magdalena Price, maprice[at]mit[dot]edu
Lena was an M.Eng student whose research focused on human-computer interfaces and machine learning. Her project focused on making effective data preparation scalable and inexpensive, in the hopes of mitigating biased model results commonly seen in big data predictions.
Now: a software engineer at Google.

Max Langenkamp Max Langenkamp, maxnz[at]mit[dot]edu, Website
Max researched AI governance. His focus was on open source machine learning software and how it shapes AI research. He is especially inspired by the work of economist Elinor Ostrom and draws from fields ranging from the philosophy of science to the economics of innovation. Before that, he researched computational cognitive science, worked at the White House Office of Science and Technology Policy, and published on AI policy at the Center for Security and Emerging Technology. He has an M.Eng from MIT in electrical engineering and computer science and loves bossa nova, Tibetan mythology and the history of technology.
Now: leading the policy and hardware initiative at SecureDNA, and a research affiliate at MIT.

Mehul Damani Mehul Damani, mehul42[at]mit[dot]edu, Website
Mehul’s research aimed to improve multi-agent reinforcement learning systems using techniques and ideas from model-based RL, intrinsic motivation, curriculum learning, and reward design. Prior to joining MIT, he worked on developing general-purpose curriculum learning methods for reinforcement learning agents and on applying reinforcement learning to domains such as multi-agent pathfinding and multi-agent traffic signal control. His hobbies include reading and playing soccer.
Now: a Ph.D. student at MIT, advised by Jacob Andreas.

Olivia Siegel Olivia Siegel, osiegel[at]mit[dot]edu, LinkedIn
Olivia was a masters student interested in AI and robotics who got most excited seeing algorithms come to life on physical robots. Prior to joining the Algorithmic Alignment group, she did her undergraduate at MIT in EECS and worked on soft robots in the Distributed Robotics Lab. Like any good New Englander, her hobbies include shellfishing, cycling, and maple syrup making.
Now: at Keystone AI in San Francisco.

A. Pinar Ozisik A. Pinar Ozisik, pinaro[at]mit[dot]edu, Website
Pinar was a research scientist in the group, broadly interested in ensuring that algorithms and systems behave safely, correctly, and in line with their intended purpose in the real world. Before joining the team, Pinar was a visiting researcher in the MIT Media Lab with the Camera Culture group. She received her Ph.D. from the University of Massachusetts Amherst, where she focused on the analysis and application of concentration inequalities to preserve the desirable properties of systems.

Prajna Soni Prajna Soni, prajna[at]mit[dot]edu, Website, LinkedIn
Prajna was an S.M. student in the Technology and Policy Program and EECS. Her research interests are broadly in the evaluation of algorithmic systems, both from a technical and regulatory perspective, algorithmic fairness and tools which facilitate the development of safer and more trustworthy AI. Prior to MIT, Prajna graduated from NYU Abu Dhabi in 2020 and was awarded the Post-graduate Research Fellow at NYU Abu Dhabi where she investigated bias propagation in recommender systems.
Now: an Applied Research Scientist at Alinia AI.

Rakshit S. Trivedi Rakshit S. Trivedi, rstrivedi[at]csail[dot]mit[dot]edu, Website
Rakshit was a Postdoctoral Associate in the Computer Science and Artificial Intelligence Laboratory (CSAIL) at MIT. Prior to that, he was a Postdoctoral Fellow in EconCS at Harvard School of Engineering and Applied Sciences (SEAS) working on multi-agent reinforcement learning and imitation learning for economic design. He obtained his PhD from Georgia Institute of Technology, focusing on machine learning for networked and multi-agent systems. He is broadly interested in the development of AI capable of learning from human experiences, can quickly adapt to evolving human needs, and achieve alignment with human values. Through the lens of multi-agent reinforcement learning, he is interested in studying the effectiveness of such AI in the presence of social, economic and cultural factors.

Rui-Jie Yew Rui-Jie Yew, rjy[at]mit[dot]edu, Website
Rui-Jie was an S.M. student in Technology and Policy. Her research interests are in the human-centered and legal aspects of computation, both in the design of regulation for emerging technologies as well as in the operationalization of legal values for technical systems. In 2021, she received a joint B.A. in computer science and mathematics from Scripps College as an off-campus student at Harvey Mudd College.
Now: a Ph.D. student in computer science at Brown University, advised by Suresh Venkatasubramanian.

Stephen Casper Stephen Casper, scasper[at]mit[dot]edu, Website
Cas worked on a lot of miscellaneous things, but most of his work focused on red-teaming, robustness, and evaluations/audits. Before his Ph.D, he worked with the Harvard Kreiman Lab and the Center for Human-Compatible AI. Hobbies of his include biking, growing plants, and keeping insects.
Now: Assistant Professor at the Harvard Kennedy School, Faculty Affiliate with Harvard SEAS, and Faculty Associate with the Berkman Klein Center.

Stewart Slocum Stewart Slocum, sslocum3[at]mit[dot]edu, Website
Stewart does empirical research aimed towards making future AI systems aligned and controllable. Previously, he has worked on editing model beliefs, character training, and evaluations for capabilities and hazardous knowledge. Outside of work, he enjoys playing piano, walking in nature, salsa dancing, and meditation.
Now: on the safety team at xAI, on leave from a Ph.D. at MIT CSAIL.

Taylor Curtis Taylor Lynn Curtis, tlcurtis[at]mit[dot]edu, LinkedIn
Taylor Lynn was an S.M. candidate in technology and policy. Before that degree, she obtained a bachelor of software engineering with a minor in political science from McGill University in Montréal, Canada. Her research interests include the effectiveness of governance surrounding large, generative models, quantifying the societal impact of AI, and more generally the regulatory frameworks that function in the technology space.
Now: Policy Advisor for AI and Science in the office of Yoshua Bengio at Mila.

Timothy Kostolansky Timothy Kostolansky, timkosto[at]mit[dot]edu, Website
Tim was a Master of Engineering student interested in improving the safety of AI systems, specifically through the lens of interpretability. Before his MEng, Tim double-majored in computer science and physics at MIT. Outside of work, Tim enjoys playing sports, reading, and meditating.
Now: on the applied research team at Prime Intellect.

Timothy Qian Timothy Qian, tcqian[at]mit[dot]edu, Website, LinkedIn
Timothy was a Master of Engineering (MEng) student focused on red-teaming, robustness, and preference learning for language models. Before joining the Algorithmic Alignment Group, he conducted research in quantum information theory and proximal operators.