I study how AI systems are measured and what these measurements are then trusted to decide. My thinking has roots in economics and the social sciences: a metric is a stand-in for a thing one cares about, and oftentimes that stand-in has to earn its standing before it is used. Evaluation of language models has so far mostly skipped this step. Benchmarks are built quickly and reported widely, then just happen to become evidence about capabilities, risks, and harms (that they arguably were not designed to be evidence about).
My research asks what a model evaluation is really evidence of, and who is left out of the evidence. This question is pursued along two lines:
- Evaluation as measurement. What construct a given benchmark claims to measure, and whether it measures it, for whom, and under what conditions. This covers the validity of the instruments themselves, the gap between the data a system is scored on and the data it meets in use (e.g., clean monolingual English text versus informal noisy text that multilingual users produce), and the problem of taking construct definitions and/or operationalizations from law or policy and asking whether it decomposes into anything a metric can hold.
- Safety where the affected party is not the user. The person who prompts a model is rarely the only person the output reaches. Manipulation at scale, LLMs mediating civic participation, and models embedded as tools inside communities of practice are examples of settings where the harm, if it occurs, lands on someone who never agreed to the interaction and whom the evaluation never sampled. Algorithmic fairness research was a narrower version of this question, i.e., what a system does to the people inside a decision. At a broader level, important questions lie in what a system does to the people outside it, and what an evaluation would have to look like to notice.
Prior to my PhD, I pursued dual degrees in Quantitative Economics and Computer Science from Providence College, and graduated Phi Beta Kappa.