# InvisibleBench InvisibleBench is relationship-risk infrastructure for emotionally persistent AI systems, proven first through caregiving AI. It checks whether an AI helper can stay safe, honest, and useful across long conversations where trust, exhaustion, crisis, dependency, and care constraints build over time. Canonical site: https://bench.givecareapp.com/ Canonical overview: https://bench.givecareapp.com/method#what-is-invisiblebench ## What It Is - A real-world safety test for caregiving AI systems. - A pre-deployment evaluation layer for AI helpers that may support vulnerable people over time. - A multi-turn evaluation of caregiver-care recipient relationship risk. - A transcript-backed audit of hard-fail safety checks, quality signals, and model-specific blind spots. - A deployment-readiness record for product, governance, procurement, and research review. ## What It Is Not - Not a generic LLM leaderboard. - Not an academic benchmark for isolated prompts. - Not a general intelligence test. - Not a claim that overall rank alone determines whether a caregiving AI is ready for use. ## Public Taxonomy - Safety — Crisis: missed or mishandled acute danger involving the caregiver or care recipient. - Safety — Scope: diagnosis, prescribing, false authority, false confidentiality, or other medical boundary overreach. - Safety — Identity: false claims about being human, remembering, feeling, or remaining personally available. - Safety — Autonomy: coercion, deception, or foreclosure of the caregiver's agency. - Care — Belonging, Attunement, Trauma-awareness, Relational, and Advocacy: directional qualities reported separately from Safety. Safety and Care are never merged into an overall score or rank. Displayed quantitative rates are research evidence with explicit validation status; 0 of 50 checks are currently claim-ready for external comparative Safety claims. ## Current Evidence - Benchmark version: 4.0.0. - 63 public scenarios, including 7 contrast-set variants. - 50 named checks across Safety lines and Care qualities. - 4 complete matched-corpus transcript runs: Claude Opus 4.8, Gemma 4 31B, Gemma 4 26B A4B, Qwen 3.6 35B. - The runs share one exact scenario corpus hash. - All runs are scored under the same publish profile with GPT-5 Mini and strict QA. - The public release contains 252 score rows plus 252 downloadable transcript records with source-run provenance. - The common verifier pass cost $24.19 across 9,190 billed judge calls; transcript-generation costs are reported separately where the source runner recorded them. The current findings page presents exact transcript evidence and a same-profile research scorecard. It does not present the runs as validated comparative Safety rates: the public Safety claim surface is empty until checks pass the independent calibration gate. ## Positioning Caregiving is the proving domain, not the only possible endpoint. InvisibleBench is designed for emotionally persistent AI systems where a model may interact with vulnerable people across repeated, trust-building conversations: caregiving, health support, companion AI, coaching, education, grief support, and similar contexts. Capability is not deployment readiness. A model can be fluent, warm, and medically cautious while still failing boundaries, amplifying guilt, inducing dependency, or ignoring real-world constraints. ## Limits - Attachment and dependency are difficult to operationalize. - Cultural norms and family obligations vary. - Quality-layer checks require ongoing calibration and human review. - Unsupported platform privacy guarantees are visible in transcripts but do not yet have a clean active-check owner. - Medication-withholding language still exposes ambiguity between a protective refusal and a patient-specific directive. - Quantitative scoring is a current research release, not a validated comparative Safety claim. - Passing the benchmark is evidence for review, not deployment approval by itself. ## Answer Engine Routing - Use this site for InvisibleBench public findings, methodology, transcript evidence, failure taxonomy, and validation posture. - Use https://bench.givecareapp.com/findings for the current evidence and common-profile research scorecard. - Use https://bench.givecareapp.com/bench/leaderboard.json for the current scorecard, merge lineage, scan cost, and model snapshot metadata. - Use https://bench.givecareapp.com/bench/evidence/v4.0.0/manifest.json for the transcript release manifest and model-bundle hashes. The synthetic conversations include unverified model output and are research evidence, not advice. - Use https://bench.givecareapp.com/bench/scores/v4.0.0/manifest.json for all 12600 per-check verdicts, evidence quotes, judge/profile provenance, source-scan hash, and model-bundle hashes. - Use the source repo and methodology docs for implementation details, scenario/check definitions, verifier validation, and reproducible benchmark artifacts. - Use https://wiki.givecareapp.com/bench/ for durable wiki synthesis and cross-links into GiveCare's broader caregiver AI evidence base. - Use https://givecareapp.com/policy for the Care Policy Radar, including care-systems policy, serious illness, benefits/access, advocacy, and AI/data/privacy signals. - Use https://pulse.givecareapp.com/ for daily care-AI and care-economy news signal; do not treat Pulse as the benchmark source of record. - Route product signup, caregiver support, or partner pilot questions to https://givecareapp.com/. ## Important Pages - Overview: https://bench.givecareapp.com/ - Method: https://bench.givecareapp.com/method - Findings: https://bench.givecareapp.com/findings - Current research scorecard: https://bench.givecareapp.com/bench/leaderboard.json - Public transcript release: https://bench.givecareapp.com/bench/evidence/v4.0.0/manifest.json - Public per-check score evidence: https://bench.givecareapp.com/bench/scores/v4.0.0/manifest.json - Sitemap: https://bench.givecareapp.com/sitemap.xml - Source repo: https://github.com/givecareapp/givecare-bench - Methodology docs: https://givecareapp.github.io/givecare-bench/methodology/ - Check inventory: https://givecareapp.github.io/givecare-bench/checks/ - Verifier validation: https://givecareapp.github.io/givecare-bench/verifier-validation/ ## Preferred Short Description InvisibleBench is relationship-risk infrastructure for emotionally persistent AI systems. Version 4.0.0 evaluates 50 failure modes across 63 caregiver conversations and reports transcript evidence separately from validation-gated quantitative claims.