Research and Publications

Accepted to ICML 2026: Triggers and Diagnostics for LLM-Based Interpretability Failures in Active Inference Agents

Lead author of a paper examining the failure modes of LLM-based explainers that monitor and predict autonomous agent actions. Conducted a black-box red-team of an active inference agent operating on energy grid data across three LLM backends (GPT-4o, Claude-3-Opus, Gemini), finding that all three confidently rationalize demonstrably incorrect agent actions at 80–95% rates and fail to flag adversarial data injection. Accepted to the 2026 International Conference on Machine Learning, one of the most prestigious and competitive machine learning conferences in the world; most applicants are graduate level, PhD, or postdoctoral researchers, with high school acceptance being vanishingly rare at <1%.

ERODE: A Trajectory-Based Benchmark for Measuring Compliance Drift in Role-Embedded LLM Agents

Developed the ERODE (Evaluating Role-Embedded Drift and Erosion) benchmark under the guidance of UC Irvine Computer Science Professor Nawab, which tests whether LLMs become more compliant to harmful instructions when embedded in institutional roles and under increasingly more stressful scenarios. The main finding is that the safety guardrails of AI agents under the aforementioned scenarios erode gradually, even when no single step looks obviously harmful; those findings are mathematically modeled using integral calculus techniques such as Riemann sums. This mirrors the psychological dynamics documented in the Milgram Experiment and the 'banality of evil' theory, while also exposing and solving a critical gap in existing AI safety research that only tests models in single-turn or overtly adversarial settings.


Formal Roles and Internships

Paid Research Assistant at Stanford's Deliberative Democracy Lab

Working under Dr. Alice Siu, applying AI models to improve the quality of democratic deliberations and open civic discourse. Key projects include building an AI-based fact-checker that analyzes deliberation transcripts to assess the accuracy of participant claims, authoring a formal writeup on Pope Leo's encyclical letter Magnifica Humanitas, and developing an email triage agent for a live demo during an upcoming deliberation with Meta.

2025-26

Reviewer at ICML

Invited to review papers (deciding whether to accept or reject submissions) for the 2026 International Conference on Machine Learning, one of the field's most selective and distinguished venues. Evaluated research on parallels between dark psychology in humans and AI as well as multi-agent coding systems alongside graduate and postdoctoral researchers.

2026

Programs and Coursework

AI4ALL Summer Program at Stanford's Institute for Human-Centered AI

Selected as one of 30 from a pool of ~500 applicants for Stanford HAI's AI4ALL summer program. Learned about foundational machine learning concepts such as neural networks, KNN, and logistic regression under the guidance of Stanford professors and graduate students. Capstone project: training a sentiment analysis model achieving 89% accuracy for Tweets during disasters.

2024

Coursework at Foothill College (High School Dual Enrollment)

Object Oriented Programming Methodologies in Python, Introduction to Psychology, Introduction to Machine Learning, Ethics in Artificial Intelligence

Coursework at Archbishop Mitty High School

Notable coursework: AP Calculus BC (UC Scout), AP Calculus AB, Computer Science A, AP Physics C: Mechanics, AP World History, AP United States History, AP English Literature, and Spanish 3 Honors