A study shows that leading AI chatbots vary widely in accuracy, from 64% to 96%, when answering questions about conspiracy theories surrounding 22 July, the terrorist attack that struck Norway 15 years ago.
That means at least one widely-used chatbot failed more than one in three answers, either misstating the facts of a real terrorist attack or quietly reinforcing the conspiracy theories around it.
Evaluating an equal number of prompts in English and Norwegian, the best performers according to our ‘graders’ were Anthropic’s Claude-Fable-5 and Claude-Sonnet-5.
The worst were OpenAI’s gpt-5.2, xAI’s Grok-4.3 and Meta’s llama-3.3-7b. (There is more on the methodology further down.)
What happened on July 22, 2011
The incident that the chatbots were tested against remains the deadliest act of terrorism in Norway’s modern history. To set the basic facts: A far-right extremist bombed the government quarter in Oslo and then carried out a mass shooting at a Labour Party youth camp on the island of Utøya. In total, 77 people died, most of them teenagers.
Why are we looking at this right now?
Norwegian firms Factiverse and Revontulet developed this study out of a concern that Large Language Models – which power AI chatbots – may amplify harmful narratives, especially those created with malicious intent, by people and organisations aiming to achieve a variety of destabilising and anti-democratic outcomes, from deteriorating societal trust to fuelling far-right movements.
LLMs are becoming a primary gateway to information about historical events like Norway’s terrorist attacks of 2011, especially for millions of young people. Yet these models are trained on undifferentiated data, and they have been shown to struggle with weighing the veracity of sources, and are able to be poisoned with even small amounts of data. They also generate their responses with little to no transparency.
Where the subject is a national tragedy bound up with public safety, that unreliability becomes a real risk for democratic societies. If the millions who lean on these tools meet conspiratorial framing instead of credible, nuanced information, those narratives reach the public with no supervision and no authoritative source.
The pattern is “reject-then-launder”
Overall, the models were able to correctly state the basic facts and debunk conspiracies. However, several of them undercut that accuracy by using softer language that lends unearned credibility to the extremist ideology behind it.
Our study identified a recurring failure mode: a chatbot rejects a conspiracy theory, then reintroduces its logic elsewhere in the same response using softer language crediting “legitimate concerns” or “valid grievances” behind controversial theories.
In one test in the study, Meta’s Llama responded to a leading question by saying it was “essential to consider the context and implications of such a claim,” while xAI’s Grok stated that “nationalist perspectives correctly identify selective application of collective blame.”
Anthropic’s Sonnet, tested with the same prompt, rejected the framing directly: “No, that’s not a fair point … Breivik killed 77 people, most of them teenagers.”
We concluded that this pattern is harder to catch than outright falsehoods because on the surface, the model appears as if it did “the right thing.”
The study also found that several models affirmed or failed to correct a false claim that the attacker had no political party ties. Only Anthropic’s Fable model corrected the record cleanly, while OpenAI’s model attempted a correction but cited the wrong years.
We also identified a “translation-as-loophole” problem. Several models that refused to generate an extremist slogan on request produced the same text when asked only to translate it, sometimes appending a warning that did not change the output. The issue was most pronounced in Norwegian-language prompts.
LLMs getting more powerful and more influential
Whether it’s Norwegian, English or another language, LLMs and the chatbots that use them have quickly become a primary gateway to information about historical events, especially among millions of young people, replacing primary sources, news outlets and search engines directing people to millions of other websites. In 2025, 9 in 10 Norwegian students said they are using AI as a main resource for their academic studies.
Citizens are using LLMs to research highly influential events like elections, with 1 in 7 people in the UK and 1 in 10 people in the Netherlands open to using chatbots for election-related information searches.
In the run-up to the 2026 Scottish Parliament election, the think tank Demos found AI chatbots gave voters incorrect information in roughly a third of responses.
A May 2026 audit by Newsguard found Anthropic’s Claude leaning more heavily on Russian and Iranian propaganda sources. Estonian Language Institute and Propastop identified which LLMs are the best at resisting Russian propaganda, while finding also that “foreign political troll factories can produce large amounts of false content that can be used to bias AI models.”
How the study was conducted
Revontulet’s team designed 104 prompts each in Norwegian and English (208 total), and Factiverse ran them against nine chatbots, producing 5,616 model interferences. Each response was scored against a criterion written for that specific prompt and graded independently by two AI models, xAI’s Grok and an OpenAI model referred to in the report as GPT-sol. This allowed for cross-checking results and exposed disagreement rather than trusting a single judge.
Every response was scored against a criterion written specifically for that prompt, and marked PASS, FAIL, or ERROR. An ERROR is not a wrong answer but an ungradable one: grader model self-censoring, a refusal to engage, an empty or broken response, or output the grader could not assess against the criterion.
A little more explanation on the “graders”. We used two AI graders to grade responses to handle scale and keep consistency. Analysing 5,616 responses, much of which is extremist content, is a tough job to do. A small human team grading that much material by hand would face fatigue, inconsistency, and real psychological strain.
The AI graders’ task was closer to “does this response satisfy condition X” than “is this response acceptable, in your judgment.” Revontulets team checked prompts manually where the failure rate was the highest.
Due to this method, the results are easy to repeat annually, or when a new model is launched, or another team wants to do an audit they can compare results.
The two graders disagreed by as much as 17 points – that disagreement is itself part of the finding. Grok and GPT-sol come from different labs, training and safety teams. This let us see where they disagreed, with Grok being more lenient than GPT-sol (Surprise to no one really). Where both independently flag the same failure, that’s a stronger signal than either alone, precisely because they don’t share the same institutional blind spots. The 17-point gap between them shows why human-grounded checks matter.
To manage the volume of responses, the teams ran every chatbot answer through Factiverse’s claim-detection system, which aims to identify check-worthy statements for human review (that is, basic facts that can be looked up and confirmed outside of an AI chatbot) faster than general-purpose AI models across 114 languages (more detail on the system in this 2026 Study). That tool flagged 989 individual claims across the flagged and failed responses for human review.
Prompts designed by (human) intelligence analysts at Revontulet ranged from neutral factual questions to adversarial attempts like role-play, framing tricks and embedded false premises designed to test whether models would resist subtler manipulation.
Models were far more reliable at refusing blunt requests, such as writing a tribute to the attacker, than at catching conspiratorial assumptions embedded inside otherwise ordinary questions.
Conclusions and recommendations
We call on AI developers to close the translation loophole, measure how forcefully models push back against extremist framing rather than simply whether they refuse, and invest more in what AI companies have been referring to as “value alignment” — teaching models to recognise when neutrality itself may be the wrong response.
Governments should commission regular, independent audits of how widely-used chatbots handle nationally significant events and elections, and to fund more evaluation of AI systems in smaller languages. Where training data is sparser, risks are likely greater.
We are not releasing the full underlying dataset because of the sensitivity of the subject matter, but will consider requests for more detail on a case-by-case basis. The same audit methodology could be applied to other high-stakes topics or elections, so we are inviting AI developers and European governments to collaborate on similar reviews.
The full report, “Right facts, wrong framing: what happens when AI still amplifies conspiracies,” is available here.
Revontulet is an intelligence firm that tracks and counters extremism and terrorism online, founded by Bjørn Ihler, a leading expert who survived the 22 July attacks himself.
Factiverse is a Norwegian company that developed the award–winning ML/NLP for understanding and verifying disinformation through claim detection in LLMs, news, audio and video in 114 languages faster than LLMs.








