publications
publications by categories in reversed chronological order. generated by jekyll-scholar.
2026
- ATTRIBAlignment Inertia: Auditing the Durability of Training Data Influence Through Policy Override Resistance2026Under review, NeurIPS Workshop on Attributing Model Behavior at Scale (ATTRIB)
- TAIObjection Without Action: Conversational Dark Patterns in AI-Mediated Commerce2026Under review, NeurIPS Workshop on Trustworthy AI
- Palgrave
2024
- JOTSAlgorithmic Impact Assessments at Scale: Practitioners’ Challenges and NeedsJournal of Online Trust and Safety, 2024
2022
- LRECThe Measuring Hate Speech Corpus: Leveraging Rasch Measurement Theory for Data PerspectivismIn Proceedings of the 1st Workshop on Perspectivist Approaches to NLP @ LREC2022, 2022
- FAccTAssessing annotator identity sensitivity via item response theory: A case study in a hate speech corpusIn Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, 2022
- WOAHTargeted Identity Group Prediction in Hate Speech CorporaIn Proceedings of the Sixth Workshop on Online Abuse and Harms (WOAH), 2022
2021
- arXivFairness on the ground: Applying algorithmic fairness approaches to production systemsarXiv preprint arXiv:2103.06172, 2021
2018
- FAT*Translation tutorial: a shared lexicon for research and practice in human-centered software systemsIn 1st Conference on Fairness, Accountability, and Transparency, 2018