Programming Projects

Leading author of Triggers and Diagnostics for LLM-Based Interpretability Failures in Active Inference Agents

Lead author of an ICML 2026 workshop-accepted paper on LLM explainer failure modes in autonomous agent oversight. Implemented an Active Inference agent in Julia using RxInfer.jl, modeling probabilistic beliefs via variational free energy minimization over a Gaussian state space, then red-teamed its LLM explainer across GPT-4o, Claude-3-Opus, and Gemini by corrupting observation streams and injecting adversarial prompts into agent state tuples

2026

ERODE: A Trajectory-Based Benchmark for Measuring Compliance Drift in Role-Embedded LLM Agents

Developed ERODE, an LLM safety benchmark under UC Irvine Professor Nawab, evaluating six frontier models across graduated multi-turn institutional role scenarios. Designed prompt escalation pipelines across five high-stakes domains and built calculus-grounded trajectory metrics using Riemann sums to quantify cumulative harm accumulation over multi-turn interactions.

2025

Project Title

Description of the project — what it does, what technologies you used, and what you built or learned.

2025
Skills

Skill Category (e.g. Programming Languages)

Python, JavaScript, HTML/CSS, etc.

Skill Category (e.g. Frameworks & Tools)

React, Git, etc.

Skill Category (e.g. Concepts)

Machine learning, NLP, etc.