Research

Current Research

A non exhaustive list of directions and questions I am currently working on:

  • Interpretability
    • Science of Activation Engineering
      • What is the geometry of activations in abstract concept spaces (for example, alignment)?
      • How do we unlock better steering methods for LLMs?
      • Can we measure goals of AI systems using interp methods?
      • Representation Geometry: What are the patterns model's activations use?
  • Reinforcement Learning
    • Can implicit decision-aware auditing prevent goal misgeneralization?
    • How do we build systems that continually adapt to humans?
  • Human Simulations/Cognitive Digital Twins (CDTs)
    • How do we build systems (CDTs) that capture and simulate human/user behavior?

Publications