AI Exposes Institutional Failures Before It Creates New Ones
The standard framing of AI ethics treats deployment as a clean slate: an AI system enters a well-functioning environment, and the question is whether that system behaves correctly. Nitesh Chawla and Paolo Benanti argue this framing is backwards. People turn to AI because institutions have already stopped being responsive, because education systems are overburdened, because workplaces have become transactional, because civic institutions struggle to sustain connection. AI did not create these failures. It scales them and makes them visible. Once deployed, it can repair, compound, substitute for, or conceal the failures it encounters. Responsible AI, the authors contend, must evaluate both the system and the institutional rupture into which it is introduced.
The paper, published at ACM AI Summit 2026, draws on Pope Leo XIV's 2026 encyclical Magnifica Humanitas as a moral and anthropological frame, then develops a practical architecture called RISE AI for making bounded, evidence-based claims about Responsibility, Inclusivity, Safety, and Empowerment. The central contribution is not another set of ethical principles. It is a framework for determining what evidence actually warrants, and what favorable evidence cannot override.
The Move From Principles to Protocols Is Already Happening
The last decade produced a shared vocabulary for ethical AI. That vocabulary has already been translated, unevenly, into technical standards, educational curricula, organizational management systems, assurance practices, and binding legal requirements. The translation chain looks like this: principle becomes design objective, which becomes system requirement, which becomes implementation mechanism, which becomes evaluation protocol, which becomes bounded evidence claim.
The EU AI Act is the most comprehensive example. For high-risk systems, Articles 9 through 15 establish requirements for risk management, data governance, documentation, logging, human oversight, accuracy, robustness, and cybersecurity. Article 27 requires fundamental-rights impact assessments. Articles 72 and 73 extend governance into post-market monitoring and serious-incident reporting.
The mechanisms reveal where translation remains incomplete. Oversight requirements do not establish that a reviewer has the time, authority, understanding, and institutional protection needed for meaningful control. Logs do not establish which claim they support or whether an appeal produces correction. Compliance does not establish empowerment, institutional repair, or a just distribution of technological power.
The same gap appears in standards like NIST AI RMF and ISO/IEC 42001, and in assurance practices that structure claims, arguments, and evidence. A standard supplies requirements. A pattern supplies a candidate mechanism. An operational system supplies an artifact. What none of them supply is a disciplined account of what the artifact actually proves, and where proof stops.
AI as Intervention and Revelation
The paper's most distinctive claim is that AI should be treated as both intervention and revelation. It does not enter a social vacuum. It exposes failures of responsiveness, belonging, care, and accountability, and then acquires the capacity to repair, compound, substitute for, or conceal them.
This framing changes the evaluation question. The standard question is whether the system is accurate, safe, or compliant. The paper adds three more: what institutional weakness is the system being asked to compensate for, which human relationships or capacities may it displace, and what evidence would show that the deployment has strengthened rather than weakened them.
The authors develop a rupture test to operationalize this. At scoping, the team documents the preexisting failure, available human and non-AI alternatives, institutional capacity, and capability at risk of substitution. Evaluation then compares the AI-mediated system with that baseline and tests both intended repair and plausible displacement. Faster throughput coupled with reduced access to a caseworker, for example, may defeat an empowerment claim even when accuracy improves.
A repair claim requires evidence that the relevant human or institutional capability became more available, durable, or answerable. Evidence of substitution, burden shifting, or suppressed visibility qualifies or defeats it. This is comparative, not metaphorical.
Magnifica Humanitas and the Limits of Design
Pope Leo XIV's encyclical provides a broader moral frame than the technical standards it draws on. Chapter Five places technology within a broader culture of power. The universal destination of goods is extended to patents, algorithms, platforms, technological infrastructure, and data. Control over platforms, infrastructure, data, and compute lies with major economic actors that set conditions of access and participation. AI can amplify the power of those who already possess resources, expertise, data, and regulatory influence.
Concern is not only whether a system treats an individual fairly, but who owns the infrastructure, sets the agenda, and has the capacity to make or refuse technological futures. Technological innovations are not neutral: they can foster participation and justice or intensify inequality, control, and exclusion. Dignity is not one value to be optimized alongside others. It is an inalienable premise governing the legitimacy of institutions and the treatment of every person.
This creates two important limits for design-centered approaches. The first is political economy. An evaluation architecture may establish that a deployed system preserves meaningful control for its users, offers practicable contestation, and communicates uncertainty well, while saying nothing about who owns the compute on which it runs, who set the agenda under which it was built, whose labor and resources sustain it, or which asymmetries between institutions and populations it reproduces. Measurement of the artifact is not measurement of the order the artifact serves.
The second limit concerns dignity and construct validity. Construct validity presupposes an object whose variation can be observed and whose indicators can be more or less adequate to it. Dignity is neither a variable nor an indicator. Observable conditions may provide evidence that dignity has been respected or violated, but they do not measure a person's possession of dignity or determine its weight against competing outcomes. Dignity is what allows us to say that a system can perform well on every selected measure and still treat a person merely as a means.
Evidence-Bounded Deployment vs. Measurement-Bounded Governance
The paper distinguishes two complementary disciplines. Evidence-bounded deployment limits claims to what has actually been evaluated. Evidence gaps, conflicts, and expiry conditions are recorded explicitly and trigger qualification of the claim, additional evaluation, remediation, or withdrawal when necessary.
Measurement-bounded governance records constraints that favorable evidence cannot override: uses excluded regardless of performance, actions no benchmark result can license, populations whose exclusion cannot be offset by aggregate gains, and human relationships or capabilities that a deployment may not substitute away. These are constraints, not constructs to be measured. They should be explicit and reviewable, with a normative or legal basis, the persons or communities they protect, the actor authorized to interpret them, and the process required to revise them.
Article 5 of the EU AI Act illustrates this logic by prohibiting specified practices rather than inviting a favorable risk-benefit score to legitimate them. The distinction sharpens the paper's central claim: AI does not demand better design in place of moral and political judgment. It demands both.
The RISE AI Evidence Architecture
RISE addresses a narrower operational question: when institutions make claims about a particular AI system, what evidence supports those claims, and where should those claims stop? It organizes bounded system claims around four domains: Responsibility, Inclusivity, Safety, and Empowerment. Each begins with a simple question: Who answers for the system? Whose perspectives shape it and whose needs does it serve? Whom does it protect from harm? Who gains or loses agency through it?
Empowerment is not a measurable trait or a general claim of benefit. It asks whether people gain durable capability, practical choice, contestation, and access to human and institutional support relative to an explicit baseline. Each claim passes through a validity chain: normative domain becomes bounded system claim, which becomes indicator, which becomes evidence source, which becomes bounded inference.
The primary artifact is a versioned claim-evidence graph, not a universal governance score. A minimal record links a bounded claim to the system, model, data, and interface version; population and context; accountable owner; indicator and threshold; evidence provenance; limitations; and expiry or change trigger. Evidence can be direct, proxy-based, conflicting, missing, or stale.
Each graph is linked to a context-and-power record: ownership and control of models, data, compute, and platforms; financing and procurement relationships; labor and resource dependencies; agenda-setting authority; and the distribution of risks and benefits. These are not additional indicators from which a political-economy score is calculated. They make visible the institutional order within which a supported claim is made and may supply constraints that qualify or preclude deployment.
Putting It to Work: Public Benefits Eligibility
The paper traces a concrete example. Suppose an AI system supports public-benefits eligibility decisions in an institution already marked by slow processing, limited caseworker capacity, and difficult appeals. Faster decisions alone may conceal rather than repair that rupture.
The RISE chain for an empowerment claim about meaningful contestability starts with the claim: applicants can meaningfully contest adverse AI-mediated recommendations. Requirements include AI disclosure, understandable reasons, accessible online and offline appeal, authorized human review, non-retaliation, and correction of downstream effects. Mechanisms include decision notices, reason codes, appeal endpoints, review queues, adverse-action pauses, and correction and recovery workflows.
Evidence requires versioned decision, notice, appeal, review, incident, complaint, and recovery logs, plus accessibility and comprehension studies. Indicators include notice comprehension, appeal initiation and completion, abandonment, review latency, reversal, subgroup disparities, and recovery completeness. The boundary specifies system version, jurisdiction, channels, languages, populations, decisions, and evaluation period.
The context-and-power row records system and data owner, vendor and cloud dependencies, procurement terms, eligibility-policy authority, caseworker staffing effects, and control of appeal records. Constraints prohibit irreversible adverse action without authorized human review, online-only appeal, retaliation for contesting a decision, and substitution away of practicable access to a caseworker.
Logs support audit and recovery but do not establish contestability without user studies, case-file audits, subgroup coverage, and human reviewers with practical authority. The resulting claim must state what is supported and what remains unknown, such as phone-only applicants, additional languages, or another jurisdiction.
Cross-Domain Portability
In education, completion and answer accuracy do not establish empowerment. Evidence should address learning transfer, unaided performance, student authorship, mentor availability, and the ability to question or refuse guidance. In clinical decision support, predictive performance does not establish sociotechnical safety. Evidence must join model evaluation to responsibility allocation, usable override, clinician observation, subgroup outcomes, incidents, correction, and recovery. In both domains, claims remain bounded to the evaluated population, institution, workflow, version, staffing conditions, and period.
What RISE Cannot Do
The authors are explicit about limitations. The worked trace and cross-domain probes are proof-of-concept specifications, not empirical validation. Claim formulation and threshold setting still require normative judgment and are shaped by institutional power. Evidence may underrepresent those most affected, while logging and provenance can create privacy and surveillance risks. Political-economic analysis can become superficial if ownership and infrastructure are merely added as fields. A supported bounded claim is not equivalent to moral legitimacy, legal compliance, or respect for dignity.
RISE can make evidence, power, and limits visible. It cannot resolve them by measurement alone. The framework is designed to complement, not replace, existing governance frameworks. NIST AI RMF, ISO/IEC 42001, and the EU AI Act supply lifecycle outcomes, management requirements, and legal obligations. Responsible AI pattern catalogues supply reusable practices. Assurance cases structure claims, arguments, and evidence. Model and data documentation, provenance records, and operational logs provide candidate evidence. RISE records what claim the artifact supports and where that support stops.
The Research Roadmap
The next phase tests the architecture rather than elaborating it. Seven priorities structure that work. Participatory content validity asks whether affected communities judge bounded claims to cover what matters. Inter-evaluator reproducibility tests whether independent evaluators classify evidence and inference boundaries similarly. Consequential validity examines whether profiles reveal unsupported claims or failures missed by benchmarks and compliance review.
Change sensitivity asks whether claims are invalidated when models, data, interfaces, workflows, populations, ownership, procurement, or institutional capacity change. Decision utility tests whether evidence gaps and constraints alter procurement, deployment, remediation, withdrawal, or appeals. Political-economic sensitivity examines whether the context-and-power record exposes dependencies that system-level assessment misses. Operational burden and proportionality measures the expertise, time, infrastructure, and cost required.
These programs require field studies with regulators, providers, deployers, workers, and affected populations. Comparative studies can test whether AI Act logs, impact assessments, oversight measures, conformity procedures, and post-market records support the inferences these groups need. Decision studies should compare choices made with and without a RISE profile. A framework that improves documentation without changing consequential decisions would have limited practical value.