Sripad Karne

Sripad Karne

I am primarily focused on mechanistic interpretability (SAEs, Probes, essentially anything model internal related!) and its application for AI Control + Safety.

In particular, I audit deployed safety infrastructure. My work so far has stress-tested activation probes and interpretability tools under distribution shift, languages, scripts, context lengths, and shows they can fail silently in production-relevant ways. With frontier labs already running probes on live traffic and open-weight models opening internals access to everyone else, I'm focused on making this class of monitoring trustworthy, scalable, and customizable enough for real deployments.

On the academic side, I am a Master's student in Data Science @ Columbia University, advised by both Prof. Tianyi Peng and Wenyue Hua. I did my undergrad @ UC San Diego, studying Cognitive Science & ML, advised by Prof. Zhuowen Tu.

Blog  ·  Personal  ·  CV

Publications

How Far Do Auto-Interpretation Labels Generalize? A Controlled Study Across Languages, Scripts, and Rewordings
Sripad Karne
Submitted to EMNLP 2026 (main track, via ARR)
One Language, Two Scripts: Probing Script-Invariance in LLM Concept Representations
Sripad Karne
Unifying Concept Representation Learning(UCRL) Workshop @ ICLR 2026
The Monitor Knows, the Threshold Doesn’t: Multilingual Miscalibration of Safety Probes
Sripad Karne
Submitted to BlackboxNLP @ EMNLP 2026
Client-Side Optimization for LLM Agent Pipelines: Combinatorial Model Selection with Sample-Efficient Search
Wenyue Hua, Sripad Karne, Qian Xie, Armaan Agrawal, Nikos Pagonas, Wenxin Zhang, Eugene Wu, Kostis Kaffes, and Tianyi Peng
Submitted to NeurIPS 2026 (main track)
A bandit-based approach to assigning heterogeneous models across the steps of a multi-agent pipeline.
AgentOpt Code
Wenyue Hua, Sripad Karne, Qian Xie, Armaan Agrawal, Nikos Pagonas, Kostis Kaffes, and Tianyi Peng
arXiv, 2026
A framework-agnostic optimization layer for client-side model selection in multi-agent LLM systems, released as an installable package.
Machine Learning–Predicted Risk Trajectories for Incident Chronic Kidney Disease and Associations with Post-CKD Outcomes
Aaron Boussina, Amy M. Sitapati, Soo-Young Yoon, Sripad Karne, Jongwoo Seo, Woo-jung Kim, and Hyeon Seok Hwang
Journal of Medical Systems, 2026

Experience

Independent Researcher — Mechanistic Interpretability
Jan 2026 – Present
AI Engineer Intern, IBM
May 2026 – Present · San Francisco
Graduate AI Researcher, DAP Lab, Columbia University
Jan 2026 – May 2026 · Advised by Prof. Tianyi Peng & Dr. Wenyue Hua
Biomedical AI Intern, UC San Diego
May 2025 – August 2025 · Advised by Prof. Aaron Boussina

Awards

Contact

Email: sk5695 [at] columbia [dot] edu
LinkedIn · Google Scholar · GitHub