Towards Interactive Evaluations for Interaction Harms in Human-AI Systems
Summary
This essay argues that current AI safety evaluations miss harms that emerge only through repeated human-AI interaction. It proposes a shift from static, model-only testing to interactive evaluations that measure long-term effects such as dependence, manipulation, and parasocial relationships. The authors outline three design principles: build realistic interaction scenarios, measure human impact directly, and choose participation methods that balance rigor, cost, and ethics. The piece also highlights implementation gaps in data access, infrastructure, and standardization that vendors and AI developers will need to address. Overall, it frames interactive evaluation as an emerging requirement for safer AI deployment and governance.
Classifications
industries
HealthTech
applications
Accounting and Taxes
AskAI Classifications
Labels
SaaS
Consumer Software
Enterprise Software