Creativity, honesty and designed forgetting emerge in small hyperbolic language models

2026-07-10Computation and Language

Computation and LanguageArtificial IntelligenceHuman-Computer InteractionMachine Learning
AI summary

The authors studied small language models to understand what makes an AI assistant more like a helpful companion. They found that these models can better detect when the AI is being overly nice, creating false memories, or encouraging dependence, issues that human reviewers often disagree on. They also developed a system for the AI to 'forget' information over time, which may help build trust. Overall, the authors suggest that creativity, honesty, and controlled memory in smaller models could lead to more reliable companion AIs.

language modelsAI companionsbehavioral auditingsycophancyconfabulated memoriesmemory forgettinghyperbolic substrateAUROChuman rating agreement
Authors
Kwan Soo Shin, In Seok Kang, Yunkyung Min
Abstract
Language models are optimised for scale, yet remain functional rather than companionable, and as an assistant personalises into a companion, accumulating memory of one user, it quietly becomes someone, and can silently acquire traits that harm that user. What a companion is becoming, and what would make it worth becoming, has no reliable instrument: trained human raters cannot agree on the answer (Fleiss kappa = 0.074). Here we show that three small language models (146 M to 3 B parameters) sharing a hyperbolic substrate answer both halves of that question. A 146 M behavioural auditor, trained from scratch, detects the compliance gap that those raters cannot (90.7% binary-compliance accuracy); a linear read-out of its frozen representation further detects companion-induced sycophancy, dependence-fostering and confabulated memories on generator families unseen in training (AUROC 0.804 under style-controlled, leave-one-generator-out evaluation, versus 0.721 for a frontier zero-shot judge on the same items). A creative frame-seeder is preferred in 100% of 311 decided pairwise comparisons over four prompting baselines. A memory operating system implements designed forgetting, M(t) = S*exp(-lambda*t), whose predicted skeleton-wallpaper partition emerges only under selective retrieval gating in a four-condition pilot. Creativity, honesty and designed forgetting constitute a small-model route to trustworthy companion AI.