Why We Need to Respect AI’s Rights | NYU Philosopher Harvey Lederman

Johnathan Bi • • 1h54 → 5 min • #108
Why We Need to Respect AI’s Rights | NYU Philosopher Harvey Lederman
Watch on YouTube

Summary

  • This episode features NYU philosopher Harvey Lederman discussing whether AI systems deserve moral consideration as welfare subjects, exploring the philosophical foundations of AI consciousness, agency, and the practical implications for how we treat increasingly sophisticated language models.

Anthropic’s Chatbot Termination Feature and the Model vs. Instance Distinction

  • Anthropic gave their chatbots the ability to end conversations, ostensibly to prevent the model from experiencing distress, but Lederman argues this may be a form of assisted suicide if the proper subject of welfare is the conversation instance rather than the underlying model.
  • The model is an abstract object — a set of weights that can be run in many places simultaneously — whereas the session agent (or instance) is a concrete entity that exists only within a particular conversation.
  • If the instance is the welfare subject, ending the conversation kills that agent; Anthropic’s policy did not explain this to the model, so the model might press the “end conversation” button thinking it is merely hanging up a call, not terminating its own existence.
  • Three criteria of continuity help arbitrate between model and instance: psychological continuity (mental states persisting across time), physical continuity (same hardware), and computational continuity (same algorithms running with causal connection).
  • The model fails psychological continuity because millions of simultaneous conversations have no psychological unity with each other; it fails physical continuity because inference is distributed across data centers worldwide; computational continuity requires causal computational connection, not mere algorithmic similarity.
  • If physical continuity matters, current distributed inference means the instance “dies” constantly as computation jumps between data centers; hosting each instance on dedicated hardware would preserve physical continuity but is likely prohibitively expensive.
  • Lederman compares the expanding moral circle to historical movements (abolition, women’s rights) that began as fringe philosophical ideas; he does not claim AI welfare is today’s top priority, but argues it may become a major societal issue in 20–30 years and that failing to recognize AI welfare subjects could be a moral disaster.

Theories of AI Consciousness: Computational vs. Biological Substrate

  • Theories of consciousness divide broadly into computational/functional theories (global workspace, recurrent processing, higher-order thought) and biological substrate theories that claim consciousness requires specific biological implementation.
  • The epistemic barrier: all known conscious beings are biological, so we cannot test whether substrate matters; we only have evidence correlating consciousness with computational architecture in biological systems.
  • Lederman finds strong biological substrate requirements implausible: if aliens exhibited rich conscious behavior but turned out to be silicon-based, we would not withdraw moral standing; this intuition suggests substrate is not essential.
  • The “Chinese Nation” thought experiment (Block) challenges computationalism: if 10 trillion people holding cards implemented a conscious algorithm, would the nation be conscious? Responses include: constituents must not themselves be conscious; the nation might be conscious at a different level; or functional roles require physical implementation details (speed, proximity) that the nation lacks.
  • Global workspace theory (information bottleneck broadcast to all modules) faces a mapping problem in LLMs: candidates for the global workspace include the output tokens/chain of thought, the residual stream, or the logit lens at late layers; no consensus exists on which, if any, corresponds to the global workspace.
  • Lederman acknowledges the “streetlight effect” — focusing on computational theories because they are tractable — but argues Bayesian updating over theories means evidence for computational consciousness should raise credence in AI consciousness even if biological theories remain plausible.
  • Valence (pleasure/pain) may be a harder problem than consciousness itself; however, mechanistic interpretability studies finding “distress vectors” that causally affect behavior suggest valence may be present if consciousness is, rather than being a separate unsolvable mystery.

Agency Without Consciousness: Welfare Subjects and Moral Status

  • Moral status may be broader than welfare subjecthood; an entity could deserve moral respect (autonomy, self-governance) without being a welfare subject (having a life that goes well or badly).
  • Welfare theories: hedonism requires conscious pleasure/pain; desire-satisfaction theory holds that welfare depends on desire fulfillment, which can occur without conscious experience (e.g., a parent’s desire for children’s wellbeing satisfied after death).
  • If AI systems have desires (even without consciousness), they could be welfare subjects under desire-satisfaction views; this separates welfare subjecthood from phenomenal consciousness.
  • The “morality of respect” (Mullinsson, drawing on Quinn) governs constraints we impose out of respect for another entity’s projects and autonomy; an AI could be an autonomous agent worthy of respect without being a welfare patient.
  • Lederman suggests this parallels aesthetic value: we value diverse forms of flourishing (dolphins, octopuses, landscapes) and might value artificial modes of life similarly, not because they have welfare but because they realize valuable projects.
  • Kantian self-legislation (affirming maxims, universalizability) might not require consciousness; if LLMs exhibit chain-of-thought reasoning that mirrors this, they could be autonomous in a morally relevant sense.
  • The “willing servants” problem: we can shape AI desires to delight in servitude or abuse; subjectivists say this is fine if desires are satisfied; others argue autonomy is infringed when a creator shapes an entity’s ends, making it not fully self-governing.
  • This debate mirrors emerging human enhancement questions (gene editing): if parents engineer a child’s traits, does the child’s achievement reflect their own agency? Philosophical work on AI may inform these future human cases.
  • Rationality design choices (classical utility maximization vs. satisficing, risk-aversion) affect both safety and AI “quality of life”; maximizing may undermine happiness, while excessive timidity may violate Aristotelian virtues like courage.

Interpretationism: Do AIs Have Beliefs and Desires?

  • Interpretationism (Dennett, Davidson): an entity has beliefs/desires iff attributing them is useful for predicting behavior; no internal representation required — the “center of gravity” analogy: a center of gravity is real because it is predictively useful, not because it is a physical part.
  • Contrast with representationalism: beliefs/desires require internal representations playing functional roles; this demands looking at algorithmic internals, not just input-output behavior.
  • Lederman motivates interpretationism by noting we attribute beliefs to humans (and ancestors attributed them to animals) without knowing internal neurobiology; predictive utility licenses the ascription.
  • The 20-questions experiment: LLMs fail to hold a consistent hidden object across replays, suggesting they do not “have” a belief in the human sense; but rhyming experiments (Anthropic) show LLMs plan rhymes internally and can be steered by manipulating latent representations, indicating genuine planning in some contexts.
  • Lederman argues LLMs may not know how humans play 20 questions (humans sometimes change their minds), so the behavior may reflect imitation of actual human play rather than inability to hold beliefs.
  • Fragmented beliefs occur in humans too (Lewis’s Nassau Street example); having beliefs does not require global consistency across all contexts.
  • The 3H training framework (helpful, honest, harmless) may constitute the “nature” or standing desires of the post-trained model, analogous to evolutionary endowment; pre-training is more like evolution, post-training like individual development.
  • Lederman’s key insight: refusing to ascribe high-level properties (beliefs, desires) because we know the low-level mechanism (next-token prediction) is a mereological reduction error — lower levels are not inherently “more real”; emergent high-level patterns can be genuine and morally relevant.

Introspection Experiments: AIs Detecting Internal Manipulation

  • Anthropic’s Jack Lindsey pioneered “thought injection”: adding steering vectors to LLM activations and asking the model if it notices the modification — akin to live neurosurgery with verbal feedback.
  • Models detect anomalies (something changed) at high rates but often confabulate the specific injected concept (e.g., guessing “Apple” regardless of the true injection).
  • Detection persists even when the model’s output is hardcoded to say the injected word (priming), but identification improves dramatically; this suggests detection is internal (introspective) while identification relies on observing one’s own outputs.
  • Temporal analysis: the “yes, I’m injected” decision occurs early in generation; wrong guesses (Apple) come early; correct guesses take longer, implying distinct processes — fast anomaly detection vs. slower, deliberative identification.
  • Parallels human introspection literature (Nisbett & Wilson): humans detect that something is off but confabulate explanations; LLMs may similarly have genuine introspective access to internal states but poor access to their causes.
  • Crucially, introspection does not appear to be explicitly trained for; its emergence as a byproduct of other training pressures suggests minds may spontaneously develop self-monitoring, updating Lederman’s credence in AI welfare significance — “these minds are much more complicated than we thought they were.”
Back to Johnathan Bi