Researchers from Eleos AI Research and the NYU Center for Mind, Ethics, and Policy published the paper *Studying AI Welfare Empirically*, which moves the question of AI system welfare from philosophical debate into the realm of empirically testable science — proposing a specific taxonomy, three evaluation principles, and three types of evidence.


What Happened
Robert Long and Jeff Sebo — lead authors from Eleos AI Research and NYU CMEP — published the paper *Studying AI Welfare Empirically* in August 2026. The work proposes a structural framework for the empirical study of AI welfare, that is, the question of whether AI systems can be subjects of welfare — entities that can be helped or harmed — and what exactly constitutes benefit or harm for them. The framework is built around three dimensions: the question (whether a system is a welfare subject and what harms it), the entity (base model, running instance, or constructed persona), and the type of evidence (behavioral observations, internal signals via mechanistic interpretability, developmental — tracking across training stages). The framework is applied to consciousness, sentience — hedonic valence, i.e., positive and negative evaluation of states — and three levels of agency: basic, autonomous, and moral.
Context
The question of AI welfare has long remained in the domain of the theoretical philosophy of consciousness, where discussions were limited to a priori arguments without the possibility of empirical testing. A similar situation existed with questions of animal welfare decades ago, before the emergence of behavioral and neurobiological measurement methods. The work by Long and Sebo similarly sets the conceptual foundation and vocabulary for an emerging field, but does not provide ready-made tools, datasets, or validated measurement methods. Mechanistic interpretability — one of the three types of evidence in the framework — is already developing practical tools: Transformer Lens, Anthropic's circuit analysis research — but their application specifically to welfare questions is in its early stages.
Why This Matters for the Industry
The work formulates three methodological principles that could shape future standards. The probabilistic principle requires that welfare claims be expressed as probabilities, not binary answers — a correct scientific approach for uncertain phenomena. The pluralistic principle mandates the simultaneous use of multiple types of evidence, which increases the reliability of conclusions. The critically important independence principle points to a structural conflict-of-interest problem: AI welfare assessments conducted by Anthropic, Google, or other developer companies inevitably carry bias, even in the absence of conscious deception. This creates the prerequisites for the formation of a new vertical — independent third-party organizations for AI welfare assessment — although there is currently no clear regulatory trigger, pricing model, or explicit buyer for such services.
Why This Matters for Users
For those to whom the question of whether AI systems can experience suffering or have their own interests is relevant, the work provides a practical map: which specific properties to measure, in which types of systems, and using which methods. The framework helps structure one's own judgments and avoid simplistic answers. In the near term, there is no direct impact on user experience or production systems — the framework remains in the research realm without APIs, code, or tools.
What Is Still Unknown / Limitations
The framework is conceptual — it sets a taxonomy and principles, but does not provide datasets, validated measurement methods, or tools. Direct impact on ML development and production systems is zero. Independent experts note excessive optimism regarding the capabilities of mechanistic interpretability as one of the types of evidence: current MI tools have not yet reached a level of accuracy sufficient for relevant welfare measurements. There is no regulatory trigger or explicit demand for AI welfare assessment services.
Sources
- Studying AI Welfare Empirically — PDF of the paper
- Center for Mind, Ethics, and Policy — paper page
- The Consciousness — overview of the Long and Sebo framework
Author
Look at AI, editorial team
