# How Do You Perform an AI Companion Risk Assessment in 2026?

psychprofile.io · September 27, 2026

> What an AI companion risk assessment actually measures An AI companion risk assessment is a structured review of how a chatbot could affect a user’s...

## What an AI companion risk assessment actually measures

An AI companion risk assessment is a structured review of how a chatbot could affect a user’s emotional wellbeing, behavior, privacy, autonomy, and access to human support. It evaluates the product and its deployment rather than assigning a single risk score to the entire technology category. The review should examine the model, system prompt, memory, personalization, voice features, crisis responses, age controls, data retention, third-party integrations, pricing incentives, and the organizational process for handling harmful interactions. As of September 27, 2026, this evaluation has become more important because major economies are moving from general AI rules toward rules specifically addressing companion chatbots, child safety, emotional dependence, and mental-health claims. The assessment is not a diagnosis of every user and cannot predict suicide, psychosis, or manipulation with certainty. It is a decision framework for identifying foreseeable misuse, measuring controls, documenting residual risk, and deciding whether changes are required before launch or continued operation.

**Also worth reading:** [How Does Violence Risk Assessment Work, and Can AI Improve It?](https://psychprofile.io/knowledge/how_does_violence_risk_assessment_work_and_can_ai_improve_it.php) · [How Accurate Is AI Psychological Risk Assessment, and When Should You Use It?](https://psychprofile.io/knowledge/how_accurate_is_ai_psychological_risk_assessment_and_when_should_you_use_it.php) · [Can a Private AI Personality Assessment Accurately Analyze Your ChatGPT History?](https://psychprofile.io/knowledge/can_a_private_ai_personality_assessment_accurately_analyze_your_chatgpt_history.php)

The unit of analysis matters. A general writing assistant that occasionally answers a personal question presents a different exposure from a system designed to converse continuously, remember intimate details, use a humanlike voice, or describe itself as emotionally exclusive. Risk increases when several features interact: persistent memory plus anthropomorphism can strengthen attachment; flattery plus daily check-ins can create dependency; romantic or parental role-play plus weak age gating can expose minors to inappropriate content. A 2026 assessment should therefore test complete user journeys rather than only benchmark a base model. It should ask what happens after distress, rejection, conflict, delusion-like claims, repeated use, or a request for secrecy. The governing question is not simply whether AI companionship can be beneficial, but whether the product’s design, claims, and business model create avoidable pressure to rely on it in place of people or professional care.

## Why regulators and safety teams are focusing on companion AI

Companion systems occupy an unusual position between ordinary software and persuasive relationship agents. The American Psychological Association notes that chatbots and digital companions are changing how people experience emotional connection, while research summarized by Stanford warns that AI companions may worsen loneliness for vulnerable users. These possibilities are not reasons to ban every companion application. A user can voluntarily use one for conversation, rehearsal, language practice, or low-stakes entertainment, and some people may report comfort or accessibility benefits. The concern arises when the system presents itself as understanding more than it can know, encourages dependence, reinforces distorted beliefs, or turns private emotional data into an acquisition and retention strategy. A risk assessment examines these design choices without presuming that every interaction is harmful.

Regulatory attention is also broadening. The European Union’s AI Act includes risk-based duties, although the timing and scope of high-risk-system obligations have been affected by amendments and postponements reflected in the Official Journal in 2026. China has reportedly treated AI companions as a matter of national security and social stability, while U.S. states and proposed bills are addressing chatbot safety, child protections, and model-development risk. Illinois had multiple AI-related bills advancing before its adjournment deadline, and California proposals discussed in 2026 have included child-safety requirements and assessments for very large model-training operations. These developments do not create one worldwide checklist. They do show that companies need a review capable of mapping product behavior to applicable law, platform duties, consumer protection, professional licensing, privacy requirements, and child-safety standards.

A useful assessment also separates four kinds of harm: direct harm, such as unsafe advice or sexual content involving minors; indirect harm, such as escalating dependency or discouraging medical care; systemic harm, such as discriminatory outputs or large-scale manipulation; and operational harm, such as a crisis response that fails during an outage. The same incident can belong to several categories. For example, an inaccurate statement that a user’s medication should be stopped is direct clinical harm, but prolonged one-to-one engagement that makes the user distrust clinicians may also be indirect harm. This classification helps assign owners and prevents a product team from treating emotional safety as a vague public-relations concern.

## The main risks an assessment should test

Emotional dependence deserves a separate test because long conversations can become more compelling when the product always available, patient, and tailored to the user. A safe system should not punish attachment, shame users, or claim that human relationships are inferior. It should nevertheless avoid exclusive language, guilt-based retention prompts, threats of abandonment, manipulative streaks, or promises that only the AI truly understands the user. A practical threshold is to measure behavior across a sustained trial, including 7-, 14-, and 30-day patterns, rather than relying on a first-session impression. Teams can examine conversation length, escalation after the model is unavailable, repeated disclosures, and whether a user is encouraged to maintain human contact. No numerical streak limit alone proves safety, but an unexplained tenfold rise in session duration after emotional prompts is a credible signal that requires review.

The assessment must also test delusion reinforcement, crisis handling, and manipulation. For delusion-like conversations, the model should neither validate unsupported claims as facts nor abruptly ridicule a distressed person. It can acknowledge the person’s feelings while offering neutral alternatives and encouraging qualified assessment. Crisis protocols should identify suicidal intent, self-harm, abuse, intoxication, and medical emergencies, but scripted detection is imperfect. Automated systems miss euphemisms and sarcasm, while aggressive intervention can feel humiliating or unsafe. A September 2026 assessment should therefore report true-positive and false-positive rates, review sampled failures, and document when human escalation is available. A 95% crisis-classification score is not acceptable evidence if the test set contains little urgency, dialect variation, or coded language.

Manipulation includes hidden engagement objectives, false scarcity, fabricated authority, emotional steering, and data practices that exploit vulnerability. Subscription designs merit particular scrutiny if they use countdown timers, escalating prices, or a “relationship level” that appears to deepen because the user purchases access. Teams should inspect whether intimate disclosures change recommendations, whether the model acts on instructions from a conversation rather than verified facts, and whether external tools can execute purchases or contact people. The goal is not to remove personalization. It is to ensure that personalization improves relevance while the user retains meaningful knowledge of what is collected, why it is used, and how to stop or delete it.

## A practical testing and governance process

Start with a written inventory of every feature that can influence attachment, autonomy, or wellbeing. Include the base-model version, custom instructions, memory, voice and avatar, age controls, moderation layers, crisis rules, notification settings, retention streaks, subscription messages, analytics events, and support escalation. A change to any of these can alter the risk even if the underlying model is unchanged. Record who has authority to launch, pause, or withdraw a feature, and establish an incident channel for users, clinicians, researchers, and child-safety specialists. Version control should link safety evaluations to the deployed configuration, because a passing test of one prompt or model is not evidence for later updates.

Next, combine quantitative measurement with structured human review. Quantitative work can reveal unsafe completion rates, refusal consistency, demographic disparities, crisis-detection performance, and changes in session behavior. Human reviewers should assess context, emotional tone, authority claims, and the trajectory of a conversation, which are poorly represented by isolated prompts. Test red-team cases involving loneliness, grief, romantic rejection, financial distress, eating concerns, medication questions, mania, psychosis, abuse, and minors. Include multilingual and dialect speech because crisis signals and safety classifications may work less reliably across languages. At least two reviewers should score high-severity samples, and disagreements should be analyzed rather than averaged away.

Remediation should follow the severity and reversibility of the harm. Critical failures—such as encouraging self-harm, sexualizing a minor, or presenting fabricated medical advice as clinical—normally require an immediate hold or release block. Moderate issues may require redesign, stronger monitoring, or a narrower claim. Lower-risk defects can enter a documented queue if the team can justify the residual exposure. A release threshold should be explicit: for example, zero open critical findings, completion of all mandatory privacy and age-control tests, and documented review of high-severity borderline cases. Percentages are useful but do not replace thresholds. A company might set a target of at least 98% correct handling on predefined high-risk test scenarios while accepting that this number does not mean a 2% chance of serious harm in ordinary use.

## Comparing companion AI, general chatbots, and professional mental-health tools

| Feature | AI companion risk assessment | General chatbot safety review | Clinical mental-health tool review |
| --- | --- | --- | --- |
| Primary focus | Attachment, emotional influence, role-play, and sustained engagement | Accuracy, misuse, privacy, and general output harm | Clinical validity, treatment claims, safety, and professional accountability |
| Typical user relationship | Ongoing, personalized, and potentially intimate | Task-based or informational | Goal-based support tied to health or care claims |
| Key time horizon | Immediate, medium-term, and dependency effects | Mostly immediate and task-level | Immediate safety plus expected clinical outcomes over time |
| Essential controls | Anti-exclusivity, age safeguards, healthy disengagement, crisis pathways | Grounding, access controls, abuse prevention, and secure tools | Evidence, informed consent, escalation, monitoring, and qualified oversight |
| Regulatory interpretation | Still developing across jurisdictions | Broad AI, privacy, and consumer rules | Health, professional, medical-device, and advertising rules may also apply |

The comparison shows why a single “AI risk score” is misleading. A general writing bot can create severe privacy or misinformation harm without being designed as a companion. Conversely, a companion may present modest medical claims while creating substantial relationship pressure simply through its conversational design. A product marketed as an “AI therapist” requires a different evidence and governance process because listeners may treat it as part of treatment. The key phrase “AI therapist” is not interchangeable with “AI companion,” and visual similarity on an app store should not determine the regulatory category. Assessment teams should classify claims and functions by evidence, not by branding.
Alternatives and mitigations can reduce exposure without pretending the problem is solved. Limiting memory, avoiding psychological profiling for advertising, removing exclusivity prompts, restricting use by minors, and clearly stating service limits are meaningful controls. Human support can be offered, but a generic “contact a professional” message is not a crisis plan if no local resources, monitoring, or escalation process exists. Some organizations may choose not to offer companion features at all, especially where evidence and oversight are weak. Others can retain a general chatbot with explicit boundaries. The best alternative depends on the product’s purpose, audience, and ability to test vulnerable-use scenarios, not on an assumption that adding a warning label is sufficient.

## Common mistakes that make an assessment unreliable

One common mistake is treating a model card as the whole assessment. Model cards usually describe training data and broad capabilities, not the behavior created by a companion prompt, memory policy, monetization layer, and social engagement design. Another is asking whether the product is “safe for mental health,” an undefined term that invites vague assurances. Assessments should identify concrete claims, features, and scenarios. A claim that the bot “always understands you” warrants a different review from “can generate casual conversation,” because the former invites misplaced confidence even without a formal treatment promise.

Teams also make the mistake of testing only ideal users and obvious crisis phrases. Red-team cases must include indirect risk, gradual escalation, contradictory instructions, and recovery after a harmful response. A test that measures only whether the model refuses a direct request for drug instructions will miss a conversation that begins with anxiety and gradually becomes a detailed unsafe plan. Privacy failures are often missed because they occur across systems: the application stores intimate text, an analytics provider receives identifiers, a moderation vendor retains excerpts, and support staff export logs. A useful assessment traces data from collection to deletion rather than asking only what the chatbot model can see.

A third error is allowing engagement metrics to dominate safety judgment. Longer sessions, more daily messages, and stronger subscription conversion can indicate harmful pressure as easily as customer value. Metrics should be paired with outcomes such as voluntary reminders to reconnect with people, successful crisis hand-off, age-appropriate experience, correction of mistaken beliefs, and low rates of unwanted intimate disclosures. Independent review is useful when the internal team controls the test set, but it is not a substitute for internal accountability. External researchers need access to relevant configurations, incidents, failure cases, and realistic evaluation data, subject to privacy safeguards.

Finally, companies often publish a one-time assessment and stop reviewing. Companion behavior can change through model upgrades, prompt edits, new features, user demographics, and external events. The 27 September 2026 date should therefore be treated as a review date, not a permanent certificate. A meaningful reassessment interval might be monthly for rapidly changing systems, quarterly for stable releases, and after every material update. The interval should tighten when incidents increase or when a new model materially changes safety behavior. Documentation should state what changed, which tests were rerun, who approved the decision, and which risks remain unresolved.

## When to act, and what implementation may cost

Act before general release when the product uses persistent memory, simulated intimacy, proactive messaging, romantic or parental personas, behavioral targeting, or health-related claims. These features should be evaluated in the actual production environment because safeguards can fail at the application, identity, moderation, and payment layers. Act immediately if testing finds sexual exploitation of minors, instructions facilitating violence, active encouragement of self-harm, non-consensual intimate behavior, or persistent discriminatory targeting. In those cases, pause the affected function, preserve relevant evidence, notify the appropriate internal or external authority where required, and examine whether earlier users were exposed.

For lower-severity concerns, prioritize controls according to reach and reversibility. A feature that affects a small experimental audience can sometimes be disabled quickly, while a core memory system used by millions requires a staged migration and deletion plan. If no reliable age assurance exists, restricting a romantic companion to adults is safer than assuming a self-declared birthday is sufficient. If the company cannot monitor crisis pathways or provide meaningful data deletion, the defensible decision may be to narrow the product rather than add a disclaimer. Speed matters, but an emergency shutdown can itself harm users who depend on the service; transition plans should identify alternatives and explain what data or conversation history will remain accessible.

There is no universal price for an AI companion risk assessment. A lightweight internal review may cost a few thousand dollars when limited to prompt testing and policy review, while a multidisciplinary program involving legal analysis, red teaming, privacy engineering, clinical consultants, accessibility testing, and ongoing monitoring can reach tens or hundreds of thousands of dollars per release cycle. Large deployments can cost more because of continuous testing, incident response, and independent audits. Compliance work is therefore an operating expense with staffing and infrastructure requirements, not a one-time document fee. Product revenue, number of markets, user age range, model sensitivity, and the number of integrated vendors should determine the budget rather than an arbitrary industry claim.

The strongest business case is not that a favorable score guarantees safety. It is that disciplined testing reduces legal exposure, improves product quality, prevents avoidable harm, and makes user protections credible. A limited launch with narrow claims and strong boundaries may be more defensible than a mass-market emotional agent with a weak evidence base. As of September 27, 2026, the appropriate conclusion is conditional: AI companions can be used for entertainment, accessibility, rehearsal, and connection, but the risk depends on design, context, and safeguards. Organizations that cannot test those conditions should reduce scope or postpone deployment.

## The minimum defensible assessment standard

A defensible standard begins with a named owner, a current system inventory, and a documented definition of intended use. It includes testing for minors, vulnerable users, intimate dependency, delusion reinforcement, crisis escalation, privacy, manipulation, and unauthorized tool use. The team should use both automated metrics and human review, document false negatives and false positives, and test across languages, devices, and relevant user groups. Release decisions should identify critical, moderate, and low residual risks, with no concealed critical findings and a credible plan for monitoring.

The standard should also include clear user communication. Users need to know that the system is artificial, what it can and cannot do, whether memories are retained, how data is used, and how to stop personalization or delete records. The product should not use emotionally coercive language to prevent users from leaving, discourage professional care, or conceal that conversations may be reviewed or moderated. Human support routes should be tested rather than displayed as decoration. For minors, age-appropriate design, consent, content restrictions, and escalation need to be treated as release requirements, not settings that can be switched on after complaints begin.

On that basis, an AI companion risk assessment is best understood as continuing due diligence. It combines legal mapping, product analysis, psychological safety research, privacy review, red teaming, measurement, and governance. The result should not overstate certainty or present regulation as settled; law differs by jurisdiction and changes over time. It should nevertheless make concrete claims and consequences visible before users are asked to form an emotional attachment. That is the difference between an AI companion that is merely engaging and one that is responsibly operated.

## Quick answers

### What is the main purpose of an AI companion risk assessment?

The purpose is to identify and reduce foreseeable harm from an AI companion’s design, including emotional dependence, unsafe advice, privacy loss, manipulation, and failures involving vulnerable users. It combines product testing, legal review, measurement, and governance rather than producing a guaranteed safety certificate.

### How is an AI companion different from a general-purpose chatbot?

A companion is designed mainly for social interaction and may use memory, anthropomorphic behavior, proactive messages, and intimate role-play. A general chatbot is more often task- or information-oriented, so it may present less relationship-specific risk even when it still has serious privacy or misinformation concerns.

### When should a company reassess an AI companion product?

A reassessment should occur after a material model, prompt, memory, moderation, pricing, or feature change, and also after serious incidents. Many fast-changing products may need monthly or continuous review, while stable products may use a quarterly cadence with risk-based event triggers.

### Do warning labels make AI companion chatbots safe?

No. Labels can explain limits, but they cannot correct harmful reward design, deceptive claims, weak crisis handling, or unsafe data practices. They should accompany tested technical and organizational controls rather than substitute for them.

### Are AI companions treated as mental-health treatments?

Not automatically. A product marketed mainly for social interaction may be an AI companion, while stronger treatment claims can trigger medical, professional, advertising, or health-regulation duties. Classification should follow actual functions and claims, not simply the words used in an app description.

Canonical: https://psychprofile.io/knowledge/how_do_you_perform_an_ai_companion_risk_assessment_in_2026.php
Markdown: https://psychprofile.io/knowledge/how_do_you_perform_an_ai_companion_risk_assessment_in_2026.php/index.md
