Shailja Thakur

Research Scientist IBM Research

I study the behavior of small language models in agentic settings, and develop methods to address the practical challenges of deploying them — performance, on-premise constraints, and safety. Most of the work sits around the model: harness search, decomposing agents by role-specific adapters for planning, tool use, and verification, and memory that persists across long workflows. A parallel thread is agentic AI debugging: plans that pass structural checks yet cannot execute, runs that loop without noticing, scores that shift under rewording.

Before this, I pioneered the use of Large Language Models for automating hardware design (Verilog code generation, bug repair, security assertions) — work that received a Best Paper Award from ACM TODAES. My PhD at the University of Waterloo (with Sebastian Fischmeister) focused on security and interpretability in automotive systems. I did my PostDoc at NYU with Siddharth Garg and Ramesh Karri.

Recent

2026
Aug 2026
PressBenchDrift covered in Crypto Briefing, daily.dev, and edgeX
Jun 2026
TalkConfidently Wrong: When AI Cannot Catch Its Own BugsOpen Source Summit India, Mumbai
2026
ACLTwo papers accepted at ACL 2026 — Think Like You Execute and STaD Findings
All publications