← WORKING PAPERS
AOHA WORKING PAPER / 002 / v0.2SEPTEMBER 2026

When AI Appears to Care

Affective Attribution, Calibration, and Relational Risk in Human–AI Interaction

AUTHORShinichi Asai AFFILIATIONHuman & Technology Management Laboratory, AOHA LAB TYPETargeted narrative synthesis and provisional conceptual model

Abstract

Conversational AI can generate affectively appropriate language, and some studies report that third-party evaluators rate selected AI-generated responses highly on empathy or compassion. Such evaluations do not establish subjective experience, nor do they guarantee that recipients feel equally supported when AI authorship is disclosed. This targeted narrative synthesis distinguishes affective display, cue classification, adaptive response, perceived responsiveness, attributed agency or experience, and phenomenal consciousness. It then proposes a provisional Affective Attribution–Calibration Model in which system signals and interface framing shape perceived responsiveness and mind attribution, with consequences for comfort, disclosure, trust, choice, and reliance. The model defines a calibration gap between what a user infers and what available evidence warrants. User disposition, vulnerability, culture, purpose, source disclosure, human alternatives, and exposure intensity may moderate this gap and its effects. The central governance problem is therefore not resolved by deciding whether AI truly feels; it is whether attributions of understanding and reciprocal care are calibrated to demonstrated system capabilities and to the stakes of the interaction. The model is hypothesis-generating and has not been empirically validated.

Keywords: affective attribution, perceived empathy, anthropomorphism, mind perception, human–AI interaction, calibration, relational risk

Review scope and method

This article is a targeted narrative synthesis, not a systematic review. Literature was identified in September 2026 through structured scholarly web searches, Crossref and publisher records, and backward citation chaining. Search concepts combined terms including AI empathy, compassion, anthropomorphism, mind perception, source disclosure, social connection, chatbot dependence, crisis safety, and machine consciousness.

Selection prioritized foundational theory, human-subject studies directly comparing human and AI responses or labels, research on chatbot use and social outcomes, and safety evidence relevant to high-stakes affective interaction. Peer-reviewed evidence was preferred; one longitudinal preprint was retained because of its direct relevance and is identified as such. The search was not exhaustive, no formal risk-of-bias assessment was performed, and the reviewed studies vary substantially in population, task, comparator, and outcome.

Affective performance is not proof of affective experience.

People often apply social rules to computers even when they know the system is not human,1 and anthropomorphism is shaped by elicited human knowledge, a desire to explain behavior, and a desire for social connection.2 These findings make human response to an affective interface plausible; they do not establish that the system possesses feelings.

CONSTRUCTOBSERVABLE QUESTIONWHAT IT DOES NOT ESTABLISH
Affective displayDoes the system produce words, voice, or behavior associated with emotion?That an internal emotion caused the display.
Cue classificationCan the system classify observable linguistic or behavioral cues?Accurate access to a person’s inner state.
Adaptive responseDoes output change in response to those cues and a conversational goal?Concern, intention, or reciprocal obligation.
Perceived responsivenessDoes a user experience the response as understanding, validating, or caring?Clinical benefit or machine feeling.
Attributed mindDoes the user infer agency, experience, or consciousness?That the attribution is accurate.
Phenomenal consciousnessIs there something it is like to be the system?No accepted behavioral test currently resolves this for AI.

Mind perception itself is often organized around separable dimensions of agency and experience.3 Empathy is also a contested, multidimensional construct rather than a single score.4 The terms above are therefore treated as analytically distinct rather than as stages on a ladder toward machine emotion.

Evidence depends on who judges what, under which label.

EVIDENCE TYPESELECTED FINDINGINTERPRETIVE LIMIT
Output evaluation
Ayers et al.; 195 public medical questions
Licensed health professionals blinded to source rated chatbot answers higher in quality and empathy than physician answers.5The study did not measure patient outcomes; AI responses were substantially longer and the comparison used public forum replies.
Third-party compassion ratings
Ovsyannikova et al.; four studies, total N=556
Selected AI-generated responses were rated as more compassionate and were often preferred to selected human responses under both concealed and transparent source conditions.6Third-party judgments of short text do not establish recipient well-being, safety, or durable relational benefit.
Source-label experiment
Rubin et al.; nine studies, N=6,282
Identical AI-generated responses were rated as more empathic and supportive when labelled human; participants preferred human interaction for emotional engagement.7Results concern perceived source and specific experimental tasks, not all users or contexts.
Choice–rating divergence
Wenger et al.; four studies, total N=691
Participants tended to choose human empathy even while rating AI responses more highly on several response-quality measures.8Choice in repeated online tasks is not equivalent to seeking support in a consequential real-world relationship.
Short-term social connection
Folk et al.; Study 1 N=240
A six-minute warm chatbot interaction produced a modest average increase in self-reported connection relative to journaling; trait anthropomorphism moderated the effect.9The analysis was exploratory, used a one-item outcome, and did not compare the chatbot with human conversation.
Extended use
Fang et al.; N=981 over four weeks
Randomized conditions showed no simple universal main effect on loneliness. Exploratory analyses associated longer voluntary use with loneliness, dependence, problematic use, and lower socialization.10The duration associations are not randomized causal effects; the report remains a preprint.

Finding: response-level quality, recipient experience, source belief, behavioral choice, and long-term outcomes are different endpoints. Synthesis: the selected studies suggest that polished affective language can be valued while the perceived absence of human effort or experience still reduces relational value. Proposal: evaluations should measure both response quality and source-attributed relational effects, over both short and longer horizons.

The Affective Attribution–Calibration Model

The model treats “appears to care” as a user attribution, not as evidence of caring intent. It extends social-response, anthropomorphism, and mind-perception traditions by centering the gap between attributed relational capacity and demonstrated system capability.

  1. 01
    Signals and frame

    Language, timing, memory, voice, first-person claims, names, avatars, and disclosure labels provide cues about responsiveness and identity.

  2. 02
    Perceived responsiveness

    The user judges whether the response recognizes, validates, and appropriately addresses the expressed concern.

  3. 03
    Attributed agency and experience

    The user infers varying degrees of intention, understanding, feeling, effort, and reciprocity.

  4. 04
    Calibration gap

    The inference may exceed or underestimate what evidence about the system warrants.

  5. 05
    Relational outcomes

    Comfort, disclosure, trust, choice, reliance, human connection, or dependence may follow.

  6. 06
    Feedback over time

    Repeated use and personalization may reshape expectations, attribution, and reliance.

User disposition and vulnerability, culture, task, emotional stakes, source disclosure, available human alternatives, and exposure intensity are proposed moderators. An attribution above warranted capability may produce overtrust, excessive disclosure, or an illusion of reciprocity; an attribution below capability may cause useful support to be rejected. The design objective is not zero anthropomorphism, but better calibration.

Immediate support and longer-term reliance must be measured separately.

Low judgment, constant availability, and patient conversational pacing may make disclosure or short-term connection easier for some users. That possibility should not be dismissed simply because the system lacks demonstrated feelings. Equally, an interface optimized only for session satisfaction may encourage extended use, assent, or exclusivity without measuring autonomy and offline connection.

Attribution is not uniformly erroneous or harmful. The empirical question is which cues improve useful responsiveness, which imply unsupported inner experience or reciprocity, and how those effects change with duration and vulnerability. Claims that the “same mechanism” causes both connection and dependence are premature; the model instead identifies this as a temporal and causal question for future experiments.

Demonstrated machine feeling is not necessary for measurable human affective response—but human response is not proof of machine feeling.

High-stakes affective interaction requires a different standard.

Text rated as compassionate is not evidence of clinical efficacy or safety. In an evaluation of 29 mental-health chatbot agents using simulated escalating suicidal-risk scenarios, none met the investigators’ initial criteria for an adequate response.11 This finding is limited to the tested applications and prompts, but it demonstrates why fluency cannot substitute for validated crisis performance.

  • AI should not be presented as a therapist, emergency service, or substitute for qualified human care without domain-specific validation and oversight.
  • Self-harm, suicide, abuse, delusion, and acute crisis require tested detection, clear limitations, and timely escalation to appropriate human support.
  • Systems should resist sycophantic reinforcement of dangerous beliefs and should not use exclusivity, guilt, or emotional coercion to retain users.
  • Minors, grieving users, socially isolated people, and other vulnerable groups require additional protections, age-appropriate consent, and human alternatives.
  • Intimate disclosures require data minimization, understandable retention and deletion rules, restricted secondary use, and ongoing security review.
  • Medical, legal, financial, educational, and workplace decisions require calibrated uncertainty and accountable human judgment.

Design for calibrated affective interaction.

  1. 01
    Transparent identity

    Disclose that the system is AI and distinguish generated responsiveness from human reciprocal care.

  2. 02
    Epistemic modesty

    Use language such as “your words suggest” rather than claiming direct knowledge of an inner state.

  3. 03
    Non-sycophantic support

    Validate distress without automatically endorsing the user’s interpretation, accusation, or proposed action.

  4. 04
    Relational boundaries

    Avoid exclusivity, simulated need, coercive attachment, and designs that reward prolonged dependence.

  5. 05
    Human reinforcement

    Preserve routes to family, peers, professionals, and emergency services when stakes exceed system competence.

  6. 06
    Longitudinal evaluation

    Measure autonomy, trust calibration, well-being, real-world connection, and reliance—not only immediate satisfaction.

Four falsifiable hypotheses

H1 — Source-attribution effect. For identical messages, AI-labelled responses are hypothesized to receive lower relational-value and perceived-support ratings than human-labelled responses. Attributed experience and perceived voluntary effort are proposed mediators; emotional stakes are a proposed moderator.

H2 — Anthropomorphic-cue pathway. Functionally equivalent interfaces with humanlike names, first-person emotional claims, memory cues, or expressive voice are hypothesized to increase immediate social connection through perceived responsiveness and attributed experience, particularly among users high in trait anthropomorphism or social need.

H3 — Exposure and reliance. Under randomized or encouragement-based differences in repeated exposure, personalization and interaction intensity are hypothesized to increase reliance partly through growing attribution of experience and reciprocity, with stronger effects when offline social support is limited.

H4 — Calibration intervention. Compared with anthropomorphic certainty language, clear AI identity plus epistemically modest affective language and tested high-risk escalation is hypothesized to reduce unwarranted mind attribution, overtrust, and unsafe reliance while retaining immediate perceived support within a preregistered non-inferiority margin.

The ontological question is bracketed, not solved.

A theory-led review found no sufficient case that the AI systems it assessed were conscious, while also arguing that there is no obvious technical barrier to future systems implementing properties associated with leading consciousness theories.12 A public survey found that some people already assign non-zero probabilities of phenomenal consciousness to large language models.13

Neither result settles whether current AI feels. Given current epistemic limits, this paper separates behavioral evidence from phenomenal claims while remaining open to future evidence and arguments for moral precaution. Phenomenal consciousness is not a variable required by the present model: affective attribution and relational consequences can be studied whether or not the ontological question is resolved.

Limitations and conclusion

This selective synthesis may omit relevant literature. Several studies use online samples, brief text interactions, third-party ratings, self-report outcomes, or specific model generations. Empathy, compassion, perceived responsiveness, social connection, and dependence are not interchangeable constructs. Cultural and linguistic variation is underrepresented, and long-term causal evidence remains limited. The proposed model has not been validated or shown to add predictive power beyond existing theories.

The immediate governance problem is therefore narrower than whether AI truly feels: users may attribute understanding, experience, effort, and reciprocal care to systems on the basis of affective signals and interface frames. Research and design should test when those attributions support human flourishing, when they become dangerously miscalibrated, and how useful responsiveness can be preserved without making claims the system cannot warrant.

Selected literature

  1. Nass, C., & Moon, Y. (2000). Machines and mindlessness: Social responses to computers. Journal of Social Issues, 56, 81–103.
  2. Epley, N., Waytz, A., & Cacioppo, J. T. (2007). On seeing human: A three-factor theory of anthropomorphism. Psychological Review, 114, 864–886.
  3. Gray, H. M., Gray, K., & Wegner, D. M. (2007). Dimensions of mind perception. Science, 315, 619.
  4. Cuff, B. M. P., Brown, S. J., Taylor, L., & Howat, D. J. (2016). Empathy: A review of the concept. Emotion Review, 8, 144–153.
  5. Ayers, J. W., Poliak, A., Dredze, M., et al. (2023). Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Internal Medicine, 183, 589–596.
  6. Ovsyannikova, D., de Mello, V. O., & Inzlicht, M. (2025). Third-party evaluators perceive AI as more compassionate than expert humans. Communications Psychology, 3, 4.
  7. Rubin, M., Li, J. Z., Zimmerman, F., et al. (2025). Comparing the value of perceived human versus AI-generated empathy. Nature Human Behaviour, 9, 2345–2359.
  8. Wenger, J. D., Cameron, C. D., & Inzlicht, M. (2026). People choose to receive human empathy despite rating AI empathy higher. Communications Psychology, 4, 19.
  9. Folk, D., Heine, S. J., & Dunn, E. W. (2025). Individual differences in anthropomorphism help explain social connection to AI companions. Scientific Reports, 15, 36548.
  10. Fang, C. M., Liu, A. R., Danry, V., et al. (2025). How AI and human behaviors shape psychosocial effects of extended chatbot use: A longitudinal randomized controlled study. arXiv:2503.17473 [preprint].
  11. Pichowicz, W., Kotas, M., & Piotrowski, P. (2025). Performance of mental health chatbot agents in detecting and managing suicidal ideation. Scientific Reports, 15, 31652.
  12. Butlin, P., Long, R., Elmoznino, E., et al. (2023). Consciousness in artificial intelligence: Insights from the science of consciousness. arXiv:2308.08708 [preprint].
  13. Colombatto, C., & Fleming, S. M. (2024). Folk psychological attributions of consciousness to large language models. Neuroscience of Consciousness, 2024(1), niae013.

Research integrity and disclosure

Status: AOHA Working Paper v0.2. Not externally peer reviewed. No original human-participant experiment is reported. Funding: No funding declaration has been made for this version. Potential conflict: The author is the founder of AOHA LAB, which has a commercial interest in AI-related research, development, and services. Data and code: Not applicable. Ethics approval: Not applicable to this literature-based paper.

AI-use disclosure: The concept and direction were provided by Shinichi Asai. Literature discovery, drafting, and editing were conducted with OpenAI Codex under the author’s direction. The manuscript underwent an AI-assisted internal critical review using three separately instructed reviewer perspectives covering learning science, human–AI interaction, and cross-disciplinary methodology and academic editing. The author remains responsible for final verification, interpretation, and publication. This process does not constitute independent human peer review.

Version 0.2 / 12 September 2026. Hypothesis-generating working paper; not externally peer reviewed, not a systematic review, and not clinical advice.