When AI Appears to Care
Affective Attribution, Calibration, and Relational Risk in Human–AI Interaction
Abstract
Conversational AI can generate affectively appropriate language, and some studies report that third-party evaluators rate selected AI-generated responses highly on empathy or compassion. Such evaluations do not establish subjective experience, nor do they guarantee that recipients feel equally supported when AI authorship is disclosed. This targeted narrative synthesis distinguishes affective display, cue classification, adaptive response, perceived responsiveness, attributed agency or experience, and phenomenal consciousness. It then proposes a provisional Affective Attribution–Calibration Model in which system signals and interface framing shape perceived responsiveness and mind attribution, with consequences for comfort, disclosure, trust, choice, and reliance. The model defines a calibration gap between what a user infers and what available evidence warrants. User disposition, vulnerability, culture, purpose, source disclosure, human alternatives, and exposure intensity may moderate this gap and its effects. The central governance problem is therefore not resolved by deciding whether AI truly feels; it is whether attributions of understanding and reciprocal care are calibrated to demonstrated system capabilities and to the stakes of the interaction. The model is hypothesis-generating and has not been empirically validated.
Keywords: affective attribution, perceived empathy, anthropomorphism, mind perception, human–AI interaction, calibration, relational risk
01 / Scope
Review scope and method
This article is a targeted narrative synthesis, not a systematic review. Literature was identified in September 2026 through structured scholarly web searches, Crossref and publisher records, and backward citation chaining. Search concepts combined terms including AI empathy, compassion, anthropomorphism, mind perception, source disclosure, social connection, chatbot dependence, crisis safety, and machine consciousness.
Selection prioritized foundational theory, human-subject studies directly comparing human and AI responses or labels, research on chatbot use and social outcomes, and safety evidence relevant to high-stakes affective interaction. Peer-reviewed evidence was preferred; one longitudinal preprint was retained because of its direct relevance and is identified as such. The search was not exhaustive, no formal risk-of-bias assessment was performed, and the reviewed studies vary substantially in population, task, comparator, and outcome.
02 / Constructs
Affective performance is not proof of affective experience.
People often apply social rules to computers even when they know the system is not human,1 and anthropomorphism is shaped by elicited human knowledge, a desire to explain behavior, and a desire for social connection.2 These findings make human response to an affective interface plausible; they do not establish that the system possesses feelings.
| CONSTRUCT | OBSERVABLE QUESTION | WHAT IT DOES NOT ESTABLISH |
|---|---|---|
| Affective display | Does the system produce words, voice, or behavior associated with emotion? | That an internal emotion caused the display. |
| Cue classification | Can the system classify observable linguistic or behavioral cues? | Accurate access to a person’s inner state. |
| Adaptive response | Does output change in response to those cues and a conversational goal? | Concern, intention, or reciprocal obligation. |
| Perceived responsiveness | Does a user experience the response as understanding, validating, or caring? | Clinical benefit or machine feeling. |
| Attributed mind | Does the user infer agency, experience, or consciousness? | That the attribution is accurate. |
| Phenomenal consciousness | Is there something it is like to be the system? | No accepted behavioral test currently resolves this for AI. |
Mind perception itself is often organized around separable dimensions of agency and experience.3 Empathy is also a contested, multidimensional construct rather than a single score.4 The terms above are therefore treated as analytically distinct rather than as stages on a ladder toward machine emotion.
03 / Evidence
Evidence depends on who judges what, under which label.
| EVIDENCE TYPE | SELECTED FINDING | INTERPRETIVE LIMIT |
|---|---|---|
| Output evaluation Ayers et al.; 195 public medical questions | Licensed health professionals blinded to source rated chatbot answers higher in quality and empathy than physician answers.5 | The study did not measure patient outcomes; AI responses were substantially longer and the comparison used public forum replies. |
| Third-party compassion ratings Ovsyannikova et al.; four studies, total N=556 | Selected AI-generated responses were rated as more compassionate and were often preferred to selected human responses under both concealed and transparent source conditions.6 | Third-party judgments of short text do not establish recipient well-being, safety, or durable relational benefit. |
| Source-label experiment Rubin et al.; nine studies, N=6,282 | Identical AI-generated responses were rated as more empathic and supportive when labelled human; participants preferred human interaction for emotional engagement.7 | Results concern perceived source and specific experimental tasks, not all users or contexts. |
| Choice–rating divergence Wenger et al.; four studies, total N=691 | Participants tended to choose human empathy even while rating AI responses more highly on several response-quality measures.8 | Choice in repeated online tasks is not equivalent to seeking support in a consequential real-world relationship. |
| Short-term social connection Folk et al.; Study 1 N=240 | A six-minute warm chatbot interaction produced a modest average increase in self-reported connection relative to journaling; trait anthropomorphism moderated the effect.9 | The analysis was exploratory, used a one-item outcome, and did not compare the chatbot with human conversation. |
| Extended use Fang et al.; N=981 over four weeks | Randomized conditions showed no simple universal main effect on loneliness. Exploratory analyses associated longer voluntary use with loneliness, dependence, problematic use, and lower socialization.10 | The duration associations are not randomized causal effects; the report remains a preprint. |
Finding: response-level quality, recipient experience, source belief, behavioral choice, and long-term outcomes are different endpoints. Synthesis: the selected studies suggest that polished affective language can be valued while the perceived absence of human effort or experience still reduces relational value. Proposal: evaluations should measure both response quality and source-attributed relational effects, over both short and longer horizons.
04 / Model
The Affective Attribution–Calibration Model
The model treats “appears to care” as a user attribution, not as evidence of caring intent. It extends social-response, anthropomorphism, and mind-perception traditions by centering the gap between attributed relational capacity and demonstrated system capability.
- 01Signals and frame
Language, timing, memory, voice, first-person claims, names, avatars, and disclosure labels provide cues about responsiveness and identity.
- 02Perceived responsiveness
The user judges whether the response recognizes, validates, and appropriately addresses the expressed concern.
- 03Attributed agency and experience
The user infers varying degrees of intention, understanding, feeling, effort, and reciprocity.
- 04Calibration gap
The inference may exceed or underestimate what evidence about the system warrants.
- 05Relational outcomes
Comfort, disclosure, trust, choice, reliance, human connection, or dependence may follow.
- 06Feedback over time
Repeated use and personalization may reshape expectations, attribution, and reliance.
User disposition and vulnerability, culture, task, emotional stakes, source disclosure, available human alternatives, and exposure intensity are proposed moderators. An attribution above warranted capability may produce overtrust, excessive disclosure, or an illusion of reciprocity; an attribution below capability may cause useful support to be rejected. The design objective is not zero anthropomorphism, but better calibration.
05 / Benefit & Risk
Immediate support and longer-term reliance must be measured separately.
Low judgment, constant availability, and patient conversational pacing may make disclosure or short-term connection easier for some users. That possibility should not be dismissed simply because the system lacks demonstrated feelings. Equally, an interface optimized only for session satisfaction may encourage extended use, assent, or exclusivity without measuring autonomy and offline connection.
Attribution is not uniformly erroneous or harmful. The empirical question is which cues improve useful responsiveness, which imply unsupported inner experience or reciprocity, and how those effects change with duration and vulnerability. Claims that the “same mechanism” causes both connection and dependence are premature; the model instead identifies this as a temporal and causal question for future experiments.
06 / Boundaries
High-stakes affective interaction requires a different standard.
Text rated as compassionate is not evidence of clinical efficacy or safety. In an evaluation of 29 mental-health chatbot agents using simulated escalating suicidal-risk scenarios, none met the investigators’ initial criteria for an adequate response.11 This finding is limited to the tested applications and prompts, but it demonstrates why fluency cannot substitute for validated crisis performance.
- AI should not be presented as a therapist, emergency service, or substitute for qualified human care without domain-specific validation and oversight.
- Self-harm, suicide, abuse, delusion, and acute crisis require tested detection, clear limitations, and timely escalation to appropriate human support.
- Systems should resist sycophantic reinforcement of dangerous beliefs and should not use exclusivity, guilt, or emotional coercion to retain users.
- Minors, grieving users, socially isolated people, and other vulnerable groups require additional protections, age-appropriate consent, and human alternatives.
- Intimate disclosures require data minimization, understandable retention and deletion rules, restricted secondary use, and ongoing security review.
- Medical, legal, financial, educational, and workplace decisions require calibrated uncertainty and accountable human judgment.
07 / Design
Design for calibrated affective interaction.
- 01Transparent identity
Disclose that the system is AI and distinguish generated responsiveness from human reciprocal care.
- 02Epistemic modesty
Use language such as “your words suggest” rather than claiming direct knowledge of an inner state.
- 03Non-sycophantic support
Validate distress without automatically endorsing the user’s interpretation, accusation, or proposed action.
- 04Relational boundaries
Avoid exclusivity, simulated need, coercive attachment, and designs that reward prolonged dependence.
- 05Human reinforcement
Preserve routes to family, peers, professionals, and emergency services when stakes exceed system competence.
- 06Longitudinal evaluation
Measure autonomy, trust calibration, well-being, real-world connection, and reliance—not only immediate satisfaction.
08 / Hypotheses
Four falsifiable hypotheses
H1 — Source-attribution effect. For identical messages, AI-labelled responses are hypothesized to receive lower relational-value and perceived-support ratings than human-labelled responses. Attributed experience and perceived voluntary effort are proposed mediators; emotional stakes are a proposed moderator.
H2 — Anthropomorphic-cue pathway. Functionally equivalent interfaces with humanlike names, first-person emotional claims, memory cues, or expressive voice are hypothesized to increase immediate social connection through perceived responsiveness and attributed experience, particularly among users high in trait anthropomorphism or social need.
H3 — Exposure and reliance. Under randomized or encouragement-based differences in repeated exposure, personalization and interaction intensity are hypothesized to increase reliance partly through growing attribution of experience and reciprocity, with stronger effects when offline social support is limited.
H4 — Calibration intervention. Compared with anthropomorphic certainty language, clear AI identity plus epistemically modest affective language and tested high-risk escalation is hypothesized to reduce unwarranted mind attribution, overtrust, and unsafe reliance while retaining immediate perceived support within a preregistered non-inferiority margin.
09 / Consciousness
The ontological question is bracketed, not solved.
A theory-led review found no sufficient case that the AI systems it assessed were conscious, while also arguing that there is no obvious technical barrier to future systems implementing properties associated with leading consciousness theories.12 A public survey found that some people already assign non-zero probabilities of phenomenal consciousness to large language models.13
Neither result settles whether current AI feels. Given current epistemic limits, this paper separates behavioral evidence from phenomenal claims while remaining open to future evidence and arguments for moral precaution. Phenomenal consciousness is not a variable required by the present model: affective attribution and relational consequences can be studied whether or not the ontological question is resolved.
10 / Limits
Limitations and conclusion
This selective synthesis may omit relevant literature. Several studies use online samples, brief text interactions, third-party ratings, self-report outcomes, or specific model generations. Empathy, compassion, perceived responsiveness, social connection, and dependence are not interchangeable constructs. Cultural and linguistic variation is underrepresented, and long-term causal evidence remains limited. The proposed model has not been validated or shown to add predictive power beyond existing theories.
The immediate governance problem is therefore narrower than whether AI truly feels: users may attribute understanding, experience, effort, and reciprocal care to systems on the basis of affective signals and interface frames. Research and design should test when those attributions support human flourishing, when they become dangerously miscalibrated, and how useful responsiveness can be preserved without making claims the system cannot warrant.
References
Selected literature
- Nass, C., & Moon, Y. (2000). Machines and mindlessness: Social responses to computers. Journal of Social Issues, 56, 81–103.
- Epley, N., Waytz, A., & Cacioppo, J. T. (2007). On seeing human: A three-factor theory of anthropomorphism. Psychological Review, 114, 864–886.
- Gray, H. M., Gray, K., & Wegner, D. M. (2007). Dimensions of mind perception. Science, 315, 619.
- Cuff, B. M. P., Brown, S. J., Taylor, L., & Howat, D. J. (2016). Empathy: A review of the concept. Emotion Review, 8, 144–153.
- Ayers, J. W., Poliak, A., Dredze, M., et al. (2023). Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Internal Medicine, 183, 589–596.
- Ovsyannikova, D., de Mello, V. O., & Inzlicht, M. (2025). Third-party evaluators perceive AI as more compassionate than expert humans. Communications Psychology, 3, 4.
- Rubin, M., Li, J. Z., Zimmerman, F., et al. (2025). Comparing the value of perceived human versus AI-generated empathy. Nature Human Behaviour, 9, 2345–2359.
- Wenger, J. D., Cameron, C. D., & Inzlicht, M. (2026). People choose to receive human empathy despite rating AI empathy higher. Communications Psychology, 4, 19.
- Folk, D., Heine, S. J., & Dunn, E. W. (2025). Individual differences in anthropomorphism help explain social connection to AI companions. Scientific Reports, 15, 36548.
- Fang, C. M., Liu, A. R., Danry, V., et al. (2025). How AI and human behaviors shape psychosocial effects of extended chatbot use: A longitudinal randomized controlled study. arXiv:2503.17473 [preprint].
- Pichowicz, W., Kotas, M., & Piotrowski, P. (2025). Performance of mental health chatbot agents in detecting and managing suicidal ideation. Scientific Reports, 15, 31652.
- Butlin, P., Long, R., Elmoznino, E., et al. (2023). Consciousness in artificial intelligence: Insights from the science of consciousness. arXiv:2308.08708 [preprint].
- Colombatto, C., & Fleming, S. M. (2024). Folk psychological attributions of consciousness to large language models. Neuroscience of Consciousness, 2024(1), niae013.
Integrity
Research integrity and disclosure
Status: AOHA Working Paper v0.2. Not externally peer reviewed. No original human-participant experiment is reported. Funding: No funding declaration has been made for this version. Potential conflict: The author is the founder of AOHA LAB, which has a commercial interest in AI-related research, development, and services. Data and code: Not applicable. Ethics approval: Not applicable to this literature-based paper.
AI-use disclosure: The concept and direction were provided by Shinichi Asai. Literature discovery, drafting, and editing were conducted with OpenAI Codex under the author’s direction. The manuscript underwent an AI-assisted internal critical review using three separately instructed reviewer perspectives covering learning science, human–AI interaction, and cross-disciplinary methodology and academic editing. The author remains responsible for final verification, interpretation, and publication. This process does not constitute independent human peer review.
Version 0.2 / 12 September 2026. Hypothesis-generating working paper; not externally peer reviewed, not a systematic review, and not clinical advice.