← working
essay

Input-Attribution Methods Cannot Justify AI Decisions

This essay argues that deploying AI in high-stakes domains where people are owed reasons (lending, hiring, criminal justice) is currently morally impermissible, because today's explainability methods cannot deliver genuine justification. Grounding a "duty of justification" in T.M. Scanlon's contractualism, it builds a taxonomy in which a justifiable explanation must be causally operative, of the right kind (a reason, not a mere mechanism), and not reasonably rejectable. It then shows that input-attribution methods like LIME and SHAP are fundamentally behavioral, that they model a black box's input-output correlations rather than its internal causal structure, so they fail the causal condition and offer only trust, not justification. It closes by arguing that mechanistic interpretability, though nascent, is the most promising path to explanations that could satisfy this duty, and carries intrinsic moral (not merely instrumental) significance.