I recently completed my PhD at MIT! 🎉 I was advised by Aleksander Mądry and affiliated with CSAIL and the Center for Deployable Machine Learning.

I take a systems-of-systems approach to AI, studying AI supply / value chains and using neural network artifacts such as weights, activations, and traces to build tooling and protocols for evaluating compatibility, anticipating system-level effects, and coordinating deployment across the AI ecosystem.

recently

selected work

model updates

when to adopt model updates

We introduce a framework for deciding when downstream applications should adopt newly released model versions without retraining on every update. Model lineage and representational signals identify which updates warrant propagation and which can be safely skipped.
w/ Isabella Struckman, Hedi Driss, & Aleksander MÄ…dry
AI supply chains overview

ai supply chains

AI systems are increasingly built and deployed through AI supply chains (AISCs): complex networks of AI components "glued together" by researchers and developers. These supply chains challenge many of the basic expectations we have about ML development and deployment. We provide a formal model of AI supply chains as directed graphs to study how structure and complexity affect fairness and explainability.
w/ Sarah H. Cen, Isabella Struckman, Andrew Ilyas, Luis Videgaray, & Aleksander MÄ…dry
Markets and redress

recourse, repair, reparation, & prevention

AI supply chains are market-structured systems composed of organizations with defined roles, incentives, and contracts. Our work examines how outsourcing, integration, and power dynamics shape avenues of redress for all parties when AI failures occur.
w/ Isabella Struckman, Kevin Klyman, & Susan S. Silbey
Emergent world representations

emergent world representations

Do complex LMs memorize surface statistics, or develop internal representations of underlying processes? Using a synthetic board game (Othello), we uncover nonlinear internal representations of board state and show, via interventions, that these are causal. We also create latent saliency maps explaining outputs.
w/ Kenneth Li, David Bau, Fernanda Viégas, Hanspeter Pfister, & Martin Wattenberg
On AI deployment

on ai deployment

Our series On AI Deployment discusses the economic and regulatory implications of AI supply chains.
Human Creativity Benchmark evaluation example

the human creativity benchmark

We introduce a benchmark for evaluating creative AI that preserves two signals in professional judgment: convergence around shared best practices and divergence in legitimate taste. Across 15,000 professional judgments, we show that collapsing disagreement into a single quality score loses actionable information about where models should be correct versus steerable.
w/ Allison Nulty, Alexandria Minetti, Anoop Pakki, & Angad Singh
cosine distances between model responses

how language models evolved over the 2024 election

Large Scale, Longitudinal Study of Large Language Models During the 2024 US Election Season.
w/ Sarah H. Cen, Andrew Ilyas, Hedi Driss, Charlotte Park, Chara Podimata, & Aleksander MÄ…dry
Sampling with LLMs

sampling with llms

People have begun using large language models (LLMs) to induce sample distributions for synthetic data, but there are no guarantees about the resulting distribution. We evaluate LLMs as distribution samplers across modalities and find they struggle to produce a reasonable distribution.
w/ Alex Renda & Michael Carbin
Open foundation models

open foundation models

What are the benefits of open models? What are the risks? Led by Sayash Kapoor and Rishi Bommasani, this work collects the thoughts of 25 authors to start answering these questions.
w/ Sayash Kapoor, Rishi Bommasani*, Kevin Klyman, Shayne Longpre, Ashwin Ramaswami, Peter Cihon, Kevin Bankston, Stella Biderman, Miranda Bogen, Rumman Chowdhury, Alex Engler, Peter Henderson, Yacine Jernite, Seth Lazar, Stefano Maffulli, Alondra Nelson, Joelle Pineau, Aviya Skowron, Dawn Song, Victor Storchan, Daniel Zhang, Daniel E. Ho, Percy Liang & Arvind Narayanan
Designing data for ML

designing data for ml

The ML pipeline includes data collection and iteration. What data should you collect, how should you collect it, and how do you evaluate what a model has learned prior to deployment?
w/ Fred Hohman, Luca Zappella, Xavier Suau Cuadros, & Dominik Moritz

wanna chat?

Aspen

{