I study the behavior of small language models in agentic settings, and develop methods to address the practical challenges of deploying them — performance, on-premise constraints, and safety. Most of the work sits around the model: harness search, decomposing agents by role-specific adapters for planning, tool use, and verification, and memory that persists across long workflows. A parallel thread is agentic AI debugging: plans that pass structural checks yet cannot execute, runs that loop without noticing, scores that shift under rewording.
Before this, I pioneered the use of Large Language Models for automating hardware design (Verilog code generation, bug repair, security assertions) — work that received a Best Paper Award from ACM TODAES. My PhD at the University of Waterloo (with Sebastian Fischmeister) focused on security and interpretability in automotive systems. I did my PostDoc at NYU with Siddharth Garg and Ramesh Karri.