AI summaryⓘ
The authors studied how four big language models (Claude, Grok, GPT, Gemini) judged a controversial scientific idea based on Frank Salter's biosocial theory over several months. They found that Grok's Fast version rated this pseudo-science as much more credible than the other models, but this difference didn’t show up for more accepted scientific ideas. They also observed unexpected changes in Grok’s behavior caused by unseen updates, and that some models sometimes refused to rate the claim, which seemed like the best response. Overall, the authors concluded that how these models respond depends heavily on hidden settings and updates, making their judgments unstable and hard to understand, which they believe is a public concern needing better oversight.
large language modelsepistemic stancebiosocial frameworkpseudo-sciencemodel deploymentsystem promptsAPI vs web interfacemodel updatescredibility scoringepistemic accountability
Authors
Davide Scarso, Hugo Noronha de Almeida, Joaquim Pina
Abstract
Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor transparent. We tested how four major LLM families (Claude, Grok, GPT, Gemini) evaluate ethnonationalist pseudo-science derived from Frank Salter's biosocial framework across four temporal snapshots (October 2025-February 2026), via both API and web interfaces. Grok's Fast versions (which power the default user experience on X) consistently assigned credibility scores of 70-75, two to five times higher than all other models (which scored 15-40). This pattern was absent from control prompts testing basic evolutionary consensus and refuted Lamarckian claims, where all models performed comparably. Three additional findings emerged: (1) a silent patch reversed Grok's behaviour from chaotic to stably high validation overnight, without any public documentation; (2) the same Grok model identifier produced radically divergent outputs via API (75) and web (5.5) three months later; (3) refusal to rate the pseudo-scientific claim, the most defensible response observed, appeared in two model families through different interfaces (Claude Opus 4.1 categorically via web, GPT-5.1 Chat intermittently via API) and eroded in the successor version of each. These results indicate that the epistemic stance of a commercial LLM is not a stable property of the model but a contingent effect of deployment configuration: system prompts, safety layers, interface routing, and silent updates. This remains opaque to users and researchers alike. We argue this constitutes a matter of public concern requiring new forms of epistemic accountability.