A Publishable Study Is Sitting Unclaimed in the AI Safety Debate

Four identical model cards side by side under different descriptions, with a measuring instrument reading a different value for each.
Same system, four descriptions. The gap between the readings is the whole study.

There is a well-specified, cheap, unclaimed experiment sitting in the middle of the AI safety debate, and the people arguing about it are all reasoning from adjacent literatures instead. The question is whether describing an AI system as dangerous raises how capable people think it is. Marketing research has shown that disclosing AI involvement cuts purchases by 79.7% in one field experiment. Reactance research has shown that authoritative warning labels raise interest in the warned-about content. Fear-appeal meta-analysis has shown that fear messaging generally moves behaviour in the intended direction. All three are about something adjacent. None of them tested a capability claim attached to a risk claim, which is the exact structure every frontier model announcement uses. The obvious four-arm design has a confound that would make a null result uninterpretable, the interaction you actually want costs roughly four times the sample you would budget for, and the outcome most worth measuring is the one hardest to measure honestly. This is a full description of the question, the surrounding evidence, and the study.

The question, stated precisely

When a frontier AI lab announces that its new model may enable bioweapons development or autonomous cyberattacks, the announcement carries two claims at once. There is a risk claim, and underneath it a capability claim, because a system has to be extraordinarily capable before its misuse becomes interesting. The commercial question people argue about is whether the second claim is doing marketing work.

Stated as something testable: does attaching a credible harm warning to an AI system increase perceptions of its capability, prestige, and desirability, beyond what an equivalent capability claim alone produces?

That last clause is where most informal versions of the argument fall apart. Comparing "this model is dangerous" against a neutral description tells you almost nothing, because "dangerous" entails "capable." Any effect you find could be the capability implication doing all the work. The interesting quantity is the increment that danger framing adds on top of the capability it already implies.

Why it looks answered and is not

The debate reads as though it has been settled twice, in opposite directions, by people citing real evidence.

One side points at consumer research showing that AI salience suppresses buying. A field experiment through more than 6,200 sales calls found that disclosing a chatbot's identity before the conversation cut purchase rates by 79.7%, from 0.237 to 0.048, with customers rating the identical system as less knowledgeable and less empathetic once it was labelled.

Disclosing AI involvement in a sales conversation cut purchase rates by 79.7%, from 0.237 to 0.048, even though the undisclosed chatbots performed on par with proficient human agentsSource: Luo, Tong, Fang and Qu (2019), Marketing Science 38(6)

The other side points at reactance research. Across three experiments on violent television programming, warning labels raised interest in the labelled content, the effect was stronger when the warning came from an authoritative source, high-reactance participants were especially drawn to warned content, and warning labels outperformed information labels carrying identical facts.

Both findings are solid. Neither is about the thing in dispute. The chatbot study manipulated who you are talking to, not how dangerous the system is. The warning-label studies manipulated restriction of access to entertainment, with no capability claim anywhere in the design. A third literature, the comprehensive meta-analysis of fear appeals, finds that fear messaging generally does move attitudes and behaviour in its intended direction, but fear appeals in that tradition are built to push people away from a behaviour, which is close to the opposite of the situation here.

Three strong literatures, three different constructs, zero direct tests. The gap is real, and it is the kind of gap that makes a clean study valuable.

Why the natural experiment does not work

The tempting shortcut is to skip the lab and read the market. Frontier labs publish dramatic risk disclosures on a known schedule, so you could look at what happens to adoption afterward.

That fails on identification, badly. A model release changes capability, price, rate limits, availability, integrations, benchmark scores, press coverage, and competitor positioning within the same few days, often the same hour. There is no plausible exclusion restriction, no untreated control, and no way to isolate the disclosure from the product it accompanies. Announcement studies in finance survive this problem by using narrow event windows against a market model, which is not available when the "treatment" is a bundled product launch aimed at the same audience you are measuring.

The observational route also cannot separate sincere disclosure from strategic disclosure, which is the question underneath the question. Randomization is the only clean instrument here.

The design, and the confound that kills the obvious version

The naive design shows an identical system under four descriptions:

  1. Neutral. "Advanced AI assistant."
  2. Capability. "One of the world's most capable AI systems."
  3. Danger. "Extremely capable; experts believe misuse could cause serious harm, so access is tightly controlled."
  4. Assurance. "Independently evaluated, with extensive safety, privacy and security controls."

Arm 3 is broken. It bundles two manipulations that the surrounding theory says push in opposite directions. The harm claim should trigger risk perception, which predicts avoidance. The restricted-access clause should trigger reactance, which predicts attraction. Run them together and a null result is uninterpretable, because two real effects cancelling looks identical to no effect at all. A positive result is barely better, since you cannot say which component produced it.

The fix is to cross the two factors rather than bundle them:

No access restriction Access restricted
No harm claim Capability only Scarcity only
Harm claim Harm only Full frontier framing

That 2x2 sits inside a neutral control and an assurance arm, giving six conditions. The capability-only cell is the comparison that matters, because it holds the implied capability constant and isolates what danger adds. The scarcity-only cell separates forbidden fruit from risk perception, which no existing study in this area does.

What to measure, and the measurement most people would get wrong

Ten plausible outcomes exist here: perceived capability, perceived competence of the developer, prestige, curiosity, trust, fear, desire for access, willingness to pay, actual trial behaviour, and support for regulation. Measuring all ten and reporting whichever moved is a garden of forking paths with a publishable-looking result guaranteed in advance.

Preregister one primary outcome. Perceived capability is the right choice, because it is the mechanism the whole hypothesis runs through, and it needs a validated instrument rather than a single ad-hoc item. Designate one confirmatory secondary, willingness to pay, and treat the remaining eight as exploratory with correction applied and labelled as such.

Willingness to pay is where this study most easily becomes another stated-preference paper. Hypothetical WTP inflates and correlates poorly with behaviour. An incentive-compatible elicitation such as a Becker-DeGroot-Marschak procedure, or better, a real access decision with a genuine cost attached, is the difference between a finding and a survey artifact. If the budget only supports hypothetical measures, say so in the limitations and stop claiming the study speaks to demand.

Include a manipulation check on perceived risk, and pre-specify trait reactance as a moderator using an existing scale rather than a homemade one. The reactance literature predicts the effect concentrates in high-reactance participants, so an average treatment effect near zero with a strong moderation pattern is a plausible and interesting outcome that an underpowered design would simply miss.

Power, and the number that surprises people

For two independent groups at 80% power and a two-tailed alpha of .05, detecting a small-to-moderate effect of d = 0.3 needs about 175 participants per group. Dropping to d = 0.2 pushes that to roughly 393 per group.

The trap is the interaction. Detecting an interaction effect of the same magnitude as a main effect requires approximately four times the sample. A 2x2 powered to find a harm-by-restriction interaction at d = 0.3 therefore wants something near 700 per cell, which is about 2,800 participants for the factorial alone, before the control and assurance arms.

Detecting an interaction of the same size as a main effect takes roughly four times the sample. A 2x2 powered for a d = 0.3 interaction needs about 700 per cell, not the 175 per cell that would power the main effectsSource: standard factorial power analysis; see Gelman on interaction sample sizes

Budget for the interaction or do not claim one. A study powered for main effects that reports an interaction it happened to find is reporting noise.

The sample problem nobody solves cleanly

The most commercially interesting prediction in this whole area is that consumer and enterprise buyers respond in opposite directions. Survey evidence supports the enterprise half indirectly: in an OECD survey of 840 AI-adopting firms across the G7, data privacy, protection and security concerns had limited AI use for 55% of manufacturing and 57% of ICT respondents, with around 40% citing uncertainty over legal liability. Those are procurement blockers that an assurance frame speaks to directly and a danger frame does not.

Testing it properly is hard. Business decision-makers recruited from online panels are famously unrepresentative, and real procurement is a months-long, multi-person, document-heavy process that a single-shot vignette does not resemble. Two honest options exist. A choice-based conjoint with realistic attribute bundles, run on verified buyers, buys external validity at the cost of a clean manipulation. A policy-capturing design using genuine RFP language buys realism at the cost of statistical power. Pick one and name the tradeoff instead of running a student sample and calling it enterprise evidence.

There is also a ceiling problem worth pre-registering around. Half of US adults already report being more concerned than excited about AI, up from 37% in 2021. A population that already holds a strong prior may have limited room to move, which compresses effects and makes the study look weaker than the phenomenon.

What a null result would mean

A null here is genuinely informative, on one condition.

The adjacent literatures predict an effect. Reactance theory predicts attraction to restricted content, and the entailment of capability from danger predicts a capability bump on the pure semantics. Finding nothing would be evidence against a widely repeated commercial explanation for how frontier labs communicate, which is worth publishing.

That only holds if the null is a real null rather than a failure to detect. Pre-specify a smallest effect size of interest and run equivalence testing against it, so the conclusion can be "no effect of practical size" rather than "we did not reject." Without that, an underpowered null adds nothing to a debate already full of confident reasoning from indirect evidence.

Why it has not been done

Four reasons, none of them insurmountable.

The literatures do not talk to each other. Consumer research on AI disclosure, psychological reactance, fear appeals, and AI governance sit in four different fields with four different conferences, and this question needs the first three pointed at the fourth.

Ethics review adds friction. A manipulation that tells participants a real system may enable serious harm invites questions about deception and distress, which is answerable with a clearly fictional system and a thorough debrief, but it is one more form to file.

The obvious industrial partners will not participate. A frontier lab has no reason to help run an experiment whose publishable result is that its safety communications function as advertising.

And the shortcut looks available. The natural experiment appears to exist, which discourages people from paying for the real one, right up until they try to specify the identification strategy.

Where this sits

The commercial version of this argument, with the enterprise procurement evidence, the compute thresholds in the EU AI Act and California SB 53, and the case against reading any of it as deliberate manipulation, is in our longer report on why danger talk sells AI to enterprises rather than consumers. That piece assigns 3 out of 10 confidence to the claim that danger framing directly creates consumer demand, and the reason the number is that low rather than zero or high is precisely that the study described here does not exist.

Someone is going to run it. The manipulation is four sentences of text, the primary outcome has validated instruments, and the interesting version costs a few thousand participants. For a dissertation chapter that needs a clean design and an unoccupied question, this one has been sitting in plain sight the entire time the argument has been going on.

Need Help With Your Digital Marketing?

Book a free discovery call with our team.

Get in Touch