Programming Projects & Skills
Leading author of Triggers and Diagnostics for LLM-Based Interpretability Failures in Active Inference Agents
Lead author of an ICML 2026 workshop-accepted paper on LLM explainer failure modes in autonomous agent oversight. Implemented an Active Inference agent in Julia using RxInfer.jl, modeling probabilistic beliefs via variational free energy minimization over a Gaussian state space, then red-teamed its LLM explainer across GPT-4o, Claude-3-Opus, and Gemini by corrupting observation streams and injecting adversarial prompts into agent state tuples
ERODE: A Trajectory-Based Benchmark for Measuring Compliance Drift in Role-Embedded LLM Agents
Developed ERODE, an LLM safety benchmark under UC Irvine Professor Nawab, evaluating six frontier models across graduated multi-turn institutional role scenarios. Designed prompt escalation pipelines across five high-stakes domains and built calculus-grounded trajectory metrics using Riemann sums to quantify cumulative harm accumulation over multi-turn interactions.
Project Title
Description of the project — what it does, what technologies you used, and what you built or learned.
Skill Category (e.g. Programming Languages)
Python, JavaScript, HTML/CSS, etc.
Skill Category (e.g. Frameworks & Tools)
React, Git, etc.
Skill Category (e.g. Concepts)
Machine learning, NLP, etc.