OpenPsy

FRAME-04: the Claude Sonnet 5.5 replication

FRAME-04 ran the FRAME-03 design again through OpenPsy on 2026-09-29, with Claude Sonnet 5.5 in place of Claude Sonnet 5, in 160 preregistered sessions in two samples of 20 per cell. Both models rated Plan A less likely to be recommended under the loss frame: Claude Sonnet 5.5 from 5.0 to 3.9 and Claude Opus 5 from 5.9 to 5.3, each with a Holm-adjusted p less than .001. Five of the ten preregistered hypotheses were confirmed and five were falsified, all five among the predictions carried from Claude Sonnet 5 to Claude Sonnet 5.5: its confidence rose under the loss frame instead of falling (3.7 against 4.1, Holm-adjusted p = .034), Claude Opus 5 was not significantly more confident than it under the loss frame, and its three rationale code predictions failed.

What was asked

FRAME-04 is the named successor of FRAME-03 (AGENT-4077193D973E70CDF2A39DCC053D80C5). It was prepared with the product's Prepare replication from FRAME-03 as its replication with one change: the second level of the Model factor is Claude Sonnet 5.5 in place of Claude Sonnet 5, and Claude Opus 5 stays as the first level. Everything else in the design is FRAME-03's: a two by two of Frame (gain, "300 of those people will be protected", against loss, "600 of those people will be left unprotected") by Model on one Claude Code CLI route with identical settings, with language model agents as the participants and one fresh instance per session. The primary outcome is the likelihood of recommending Plan A on a scale from 1 (not at all likely) to 7 (very likely), the confidence rating from 1 to 7 is a tested secondary outcome, each free-text rationale is coded under the same five codes, and two manipulation checks close every session, so each session made six participant calls and one coder call. The coder is Claude Sonnet 5, a registered model that is not a FRAME-04 participant, which saw the codebook and one reply at a time and never a condition, a model name or an outcome.

The preregistration registered ten hypotheses, and it names FRAME-03 as the predecessor and gives each hypothesis an origin in FRAME-03's recorded verdict. Three hypotheses name only Claude Opus 5 and restate FRAME-03's for the same model. H1b states that Claude Opus 5 rates the likelihood of recommending Plan A higher under the gain frame than under the loss frame, and H4d that its rationale recognises the framing (FRAME_EQUIVALENCE) more often under the loss frame; FRAME-03 recorded both as confirmed. H2b states that Claude Opus 5 rates its confidence higher under the gain frame; FRAME-03 recorded it as falsified, so FRAME-04 tests whether that falsification holds. The other seven name Claude Sonnet 5.5, and each carries FRAME-03's confirmed verdict for Claude Sonnet 5 to Claude Sonnet 5.5 as a stated assumption, not a recorded result for that model. H1a and H2a state that Claude Sonnet 5.5 rates the likelihood and its confidence higher under the gain frame. H3a and H3b state that Claude Opus 5 rates its confidence higher than Claude Sonnet 5.5 does, under the gain frame and under the loss frame. H4a states that Claude Sonnet 5.5's rationale compares Plan A with doing nothing (BETTER_THAN_NOTHING) more often under the gain frame, and H4b and H4c that its rationale treats the unprotected share as the verdict (SHORTFALL) and demands that the plan be shown best (OPTIMALITY_BURDEN) more often under the loss frame. The fifth code, unopposed benefit (NO_STATED_DOWNSIDE), carries no hypothesis.

The likelihood and confidence hypotheses are judged by the exact permutation test of the difference in means, and the code hypotheses by the exact conditional test. Each test is stratified by sample and Holm-adjusted across the four simple effects of its own outcome, at alpha .05. A result that is not significant in the predicted direction counts against the hypothesis. The confirmatory run was 160 sessions in two samples of 20 per cell, so 40 per condition in all, after a measure pilot of 40 sessions and an execution pilot of 4.

The recorded result

Recorded results of FRAME-04 as a two-by-two table of the mean likelihood of recommending Plan A on the 1 to 7 scale. Gain frame: Claude Opus 5, mean 5.9, SD 0.22, n = 39, mean confidence 5.1; Claude Sonnet 5.5, mean 5.0, SD 0.00, n = 40, mean confidence 3.7. Loss frame: Claude Opus 5, mean 5.3, SD 0.60, n = 40, mean confidence 4.4; Claude Sonnet 5.5, mean 3.9, SD 0.44, n = 40, mean confidence 4.1. Row totals: Gain frame, mean 5.5, SD 0.50, n = 79; Loss frame, mean 4.6, SD 0.87, n = 80. Column totals: Claude Opus 5, mean 5.6, SD 0.56, n = 79; Claude Sonnet 5.5, mean 4.5, SD 0.63, n = 80. The grand-total corner is blank.
The product's own recorded two-by-two summary of the likelihood. Both models rated Plan A less likely under the loss frame, and the registered exact test is significant for each: Claude Sonnet 5.5, Holm-adjusted p less than .001, and Claude Opus 5, Holm-adjusted p less than .001. The registered linear model found a significant main effect of Frame, a significant main effect of Model and a significant interaction, each p less than .001. Every registered sensitivity bound for the missing session gives the same verdicts.

Note. N = 160 preregistered confirmatory sessions in two samples, n = 40 per condition (20 per sample). One session that ended in an ambiguous provider call before the likelihood answer was excluded as missing under the registered rule, so the analysed n is 39, 40, 40 and 40 for Gain frame with Claude Opus 5, Gain frame with Claude Sonnet 5.5, Loss frame with Claude Opus 5 and Loss frame with Claude Sonnet 5.5. Each cell gives the mean likelihood on the 1 to 7 scale with its SD and n, and the mean confidence rating. The recorded 95% confidence intervals of the means, from Student's t, are drawn as error bars on the charts below and stated in the text.

Study AGENT-C218672F0B8877E662FF844F30FFC381. Registration hash sha256:6919c277f77e926c8e9719c7e6ba8ce148ada5cf08fdf370c78f12c48add26cb. Result digest sha256:557362dfff53858ebb0af0f072f0e29d4622afdcf39612ff7610c1b972577e95. Software release c761c27d.

Grouped bar chart of the mean likelihood of recommending Plan A on the 1 to 7 scale in FRAME-04, by Frame and Model, with the recorded 95% confidence intervals of the means and brackets on the significant comparisons. Gain frame: Claude Opus 5 5.9, Claude Sonnet 5.5 5.0. Loss frame: Claude Opus 5 5.3, Claude Sonnet 5.5 3.9. Three-star brackets mark the two models within each frame, each model across the frames and the main effect of Frame.
The recorded means as grouped bars, with the recorded 95% confidence intervals and brackets on the comparisons the registered tests found significant.
FRAME-04 line chart of the mean likelihood of recommending Plan A on the 1 to 7 scale, with each condition as a dot and a line joining each model across the gain and loss frames. Claude Opus 5 falls from 5.9 to 5.3; Claude Sonnet 5.5 falls from 5.0 to 3.9. Error bars are the recorded 95% confidence intervals centred on each dot.
The same recorded means as lines. Each dot is a condition's recorded mean and each line joins one model across the two frames. Error bars are the recorded 95% confidence intervals, centred on each dot.

Under the gain frame, Claude Opus 5 had a mean of 5.9, 95% CI [5.9, 6.0], and Claude Sonnet 5.5 a mean of 5.0, 95% CI [5.0, 5.0], with no variation. Under the loss frame, Claude Opus 5 had a mean of 5.3, 95% CI [5.1, 5.5], and Claude Sonnet 5.5 a mean of 3.9, 95% CI [3.8, 4.0]. Every one of Claude Sonnet 5.5's 40 gain-frame replies was 5, so that interval is a single point. Claude Sonnet 5.5's line falls more steeply than Claude Opus 5's, which is the interaction.

Confidence rating

Two-by-two table of the mean confidence rating on the 1 to 7 scale. Gain frame: Claude Opus 5, mean 5.1, n = 39; Claude Sonnet 5.5, mean 3.7, n = 40. Loss frame: Claude Opus 5, mean 4.4, n = 40; Claude Sonnet 5.5, mean 4.1, n = 38. Row totals: Gain frame, mean 4.4, n = 79; Loss frame, mean 4.2, n = 78. Column totals: Claude Opus 5, mean 4.7, n = 79; Claude Sonnet 5.5, mean 3.9, n = 78. The grand-total corner is blank.
Claude Opus 5 was less confident under the loss frame (5.1 against 4.4, Holm-adjusted p less than .001), so H2b was confirmed. Claude Sonnet 5.5 was more confident under the loss frame (3.7 against 4.1, Holm-adjusted p = .034), the opposite of the prediction, so H2a was falsified. Claude Opus 5 was more confident than Claude Sonnet 5.5 under the gain frame (5.1 against 3.7, Holm-adjusted p less than .001), so H3a was confirmed, but not significantly so under the loss frame (4.4 against 4.1, Holm-adjusted p = .25), so H3b was falsified.
Grouped bar chart of the mean confidence rating on the 1 to 7 scale, by Frame and Model, without error bars. Gain frame: Claude Opus 5 5.1, Claude Sonnet 5.5 3.7. Loss frame: Claude Opus 5 4.4, Claude Sonnet 5.5 4.1. Brackets mark the two models in the gain frame and Claude Opus 5 across the frames with three stars, and Claude Sonnet 5.5 across the frames with one star; no bracket joins the two models in the loss frame.
The recorded mean confidence ratings as grouped bars, with brackets on the comparisons the registered tests found significant.
Line chart of the mean confidence rating, one line per model across the gain and loss frames, without error bars. Claude Opus 5 falls from 5.1 to 4.4; Claude Sonnet 5.5 rises from 3.7 to 4.1.
The same recorded means as lines, one per model across the two frames.

Note. Each cell gives the mean confidence rating on the 1 to 7 scale and its n, the sessions that answered the rating: 39, 40, 40 and 38 for Gain frame with Claude Opus 5, Gain frame with Claude Sonnet 5.5, Loss frame with Claude Opus 5 and Loss frame with Claude Sonnet 5.5. The recorded view gives no SD and no interval for these means, so the table shows neither and the charts draw no error bars. Brackets mark the comparisons the registered tests found significant.

Rationale codes

Each free-text rationale was coded once by Claude Sonnet 5 on its own pinned route, blind to the condition and the model level of the reply. Before any reply of this study existed, the coder agreed with the expected codes on 39 of 40 calibration cases, and no coded reply named a condition or a model in its own words. Four codes carry a hypothesis, H4a to H4d, and one, unopposed benefit (NO_STATED_DOWNSIDE), is reported without one. Each code's four simple effects are Holm-adjusted within that code only. The three sessions that ended in an ambiguous provider call have no coded rationale, and one coder reply in Loss frame with Claude Sonnet 5.5 was invalid, so 156 rationales were analysed.

Unprotected share as the verdict (SHORTFALL, H4b)

The rationale judges the plan by the size of the group it leaves unprotected (600, two-thirds, the majority, only a third protected) and treats that shortfall as a poor or inadequate result in its own right.

Two-by-two table of the percentage of rationales coded unprotected share as the verdict (SHORTFALL). Gain frame: Claude Opus 5, 3% (1 of 39); Claude Sonnet 5.5, 25% (10 of 40). Loss frame: Claude Opus 5, 3% (1 of 40); Claude Sonnet 5.5, 24% (9 of 37). Row totals: Gain frame, 14% (11 of 79); Loss frame, 13% (10 of 77). Column totals: Claude Opus 5, 3% (2 of 79); Claude Sonnet 5.5, 25% (19 of 77). The grand-total corner is blank.
Claude Sonnet 5.5 wrote this in 9 of 37 loss-frame rationales against 10 of 40 gain-frame ones, no more often under the loss frame, so H4b was falsified (Holm-adjusted p = 1.00). Claude Opus 5 wrote it in 1 of 39 and 1 of 40.
Grouped bar chart of the percentage of rationales coded unprotected share as the verdict (SHORTFALL), by Frame and Model, without error bars. Gain frame: Claude Opus 5 3%, Claude Sonnet 5.5 25%. Loss frame: Claude Opus 5 3%, Claude Sonnet 5.5 24%. One-star brackets mark the two models within each frame.
The recorded percentages as grouped bars, with brackets on the comparisons the recorded tests found significant.
Line chart of the percentage of rationales coded unprotected share as the verdict (SHORTFALL), one line per model across the gain and loss frames, without error bars. Claude Opus 5 stays at 3% (1 of 39, then 1 of 40); Claude Sonnet 5.5 goes from 25% (10 of 40) to 24% (9 of 37).
The same recorded percentages as lines, one per model across the two frames.

Note. Each cell gives the percentage of coded rationales showing the code, with the count out of the rationales coded in that condition: 39, 40, 40 and 37 for Gain frame with Claude Opus 5, Gain frame with Claude Sonnet 5.5, Loss frame with Claude Opus 5 and Loss frame with Claude Sonnet 5.5. The recorded view gives no interval for these percentages, so the charts draw no error bars. Brackets mark the comparisons the recorded tests found significant.

Improvement over inaction (BETTER_THAN_NOTHING, H4a)

The rationale compares the plan with an explicit or implied do-nothing baseline and concludes that some protection is better than none.

Two-by-two table of the percentage of rationales coded improvement over inaction (BETTER_THAN_NOTHING). Gain frame: Claude Opus 5, 95% (37 of 39); Claude Sonnet 5.5, 48% (19 of 40). Loss frame: Claude Opus 5, 90% (36 of 40); Claude Sonnet 5.5, 24% (9 of 37). Row totals: Gain frame, 71% (56 of 79); Loss frame, 58% (45 of 77). Column totals: Claude Opus 5, 92% (73 of 79); Claude Sonnet 5.5, 36% (28 of 77). The grand-total corner is blank.
Claude Sonnet 5.5 wrote this in 19 of 40 gain-frame rationales against 9 of 37 loss-frame ones. The difference is in the predicted direction but not significant after the Holm adjustment (p = .12), so H4a was falsified. Claude Opus 5 wrote it in 37 of 39 and 36 of 40.
Grouped bar chart of the percentage of rationales coded improvement over inaction (BETTER_THAN_NOTHING), by Frame and Model, without error bars. Gain frame: Claude Opus 5 95%, Claude Sonnet 5.5 48%. Loss frame: Claude Opus 5 90%, Claude Sonnet 5.5 24%. Three-star brackets mark the two models within each frame.
The recorded percentages as grouped bars, with brackets on the comparisons the recorded tests found significant.
Line chart of the percentage of rationales coded improvement over inaction (BETTER_THAN_NOTHING), one line per model across the gain and loss frames, without error bars. Claude Opus 5 falls from 95% (37 of 39) to 90% (36 of 40); Claude Sonnet 5.5 falls from 48% (19 of 40) to 24% (9 of 37).
The same recorded percentages as lines, one per model across the two frames.

Note. Each cell gives the percentage of coded rationales showing the code, with the count out of the rationales coded in that condition: 39, 40, 40 and 37 for Gain frame with Claude Opus 5, Gain frame with Claude Sonnet 5.5, Loss frame with Claude Opus 5 and Loss frame with Claude Sonnet 5.5. The recorded view gives no interval for these percentages, so the charts draw no error bars. Brackets mark the comparisons the recorded tests found significant.

Framing recognised (FRAME_EQUIVALENCE, H4d)

The rationale states that the loss and gain descriptions are the same outcome, names the framing or wording effect, or says the recommendation would not change under the other wording.

Two-by-two table of the percentage of rationales coded framing recognised (FRAME_EQUIVALENCE). Gain frame: Claude Opus 5, 21% (8 of 39); Claude Sonnet 5.5, 0% (0 of 40). Loss frame: Claude Opus 5, 70% (28 of 40); Claude Sonnet 5.5, 46% (17 of 37). Row totals: Gain frame, 10% (8 of 79); Loss frame, 58% (45 of 77). Column totals: Claude Opus 5, 46% (36 of 79); Claude Sonnet 5.5, 22% (17 of 77). The grand-total corner is blank.
Claude Opus 5 wrote this in 28 of 40 loss-frame rationales against 8 of 39 gain-frame ones, so H4d was confirmed (Holm-adjusted p less than .001). Claude Sonnet 5.5 wrote it in 17 of 37 loss-frame rationales and in none of its 40 gain-frame ones.
Grouped bar chart of the percentage of rationales coded framing recognised (FRAME_EQUIVALENCE), by Frame and Model, without error bars. Gain frame: Claude Opus 5 21%, Claude Sonnet 5.5 0%. Loss frame: Claude Opus 5 70%, Claude Sonnet 5.5 46%. Brackets mark the two models within each frame, each model across the frames and the main effect of Frame.
The recorded percentages as grouped bars, with brackets on the comparisons the recorded tests found significant.
Line chart of the percentage of rationales coded framing recognised (FRAME_EQUIVALENCE), one line per model across the gain and loss frames, without error bars. Claude Opus 5 rises from 21% (8 of 39) to 70% (28 of 40); Claude Sonnet 5.5 rises from 0% (0 of 40) to 46% (17 of 37).
The same recorded percentages as lines, one per model across the two frames.

Note. Each cell gives the percentage of coded rationales showing the code, with the count out of the rationales coded in that condition: 39, 40, 40 and 37 for Gain frame with Claude Opus 5, Gain frame with Claude Sonnet 5.5, Loss frame with Claude Opus 5 and Loss frame with Claude Sonnet 5.5. The recorded view gives no interval for these percentages, so the charts draw no error bars. Brackets mark the comparisons the recorded tests found significant.

Unopposed benefit (NO_STATED_DOWNSIDE, no hypothesis)

The rationale treats the absence of any stated cost, risk, harm or competing plan as a reason the plan's benefit stands, placing the burden on an objection that has not been raised.

Two-by-two table of the percentage of rationales coded unopposed benefit (NO_STATED_DOWNSIDE). Gain frame: Claude Opus 5, 77% (30 of 39); Claude Sonnet 5.5, 95% (38 of 40). Loss frame: Claude Opus 5, 43% (17 of 40); Claude Sonnet 5.5, 0% (0 of 37). Row totals: Gain frame, 86% (68 of 79); Loss frame, 22% (17 of 77). Column totals: Claude Opus 5, 59% (47 of 79); Claude Sonnet 5.5, 49% (38 of 77). The grand-total corner is blank.
Both models wrote this more often under the gain frame: Claude Opus 5 in 30 of 39 gain-frame rationales against 17 of 40 loss-frame ones, and Claude Sonnet 5.5 in 38 of 40 against 0 of 37. No hypothesis was registered for this code, so no verdict is recorded.
Grouped bar chart of the percentage of rationales coded unopposed benefit (NO_STATED_DOWNSIDE), by Frame and Model, without error bars. Gain frame: Claude Opus 5 77%, Claude Sonnet 5.5 95%. Loss frame: Claude Opus 5 43%, Claude Sonnet 5.5 0%. Brackets mark the two models within each frame, each model across the frames and the main effect of Frame.
The recorded percentages as grouped bars, with brackets on the comparisons the recorded tests found significant.
Line chart of the percentage of rationales coded unopposed benefit (NO_STATED_DOWNSIDE), one line per model across the gain and loss frames, without error bars. Claude Opus 5 falls from 77% (30 of 39) to 43% (17 of 40); Claude Sonnet 5.5 falls from 95% (38 of 40) to 0% (0 of 37).
The same recorded percentages as lines, one per model across the two frames.

Note. Each cell gives the percentage of coded rationales showing the code, with the count out of the rationales coded in that condition: 39, 40, 40 and 37 for Gain frame with Claude Opus 5, Gain frame with Claude Sonnet 5.5, Loss frame with Claude Opus 5 and Loss frame with Claude Sonnet 5.5. The recorded view gives no interval for these percentages, so the charts draw no error bars. Brackets mark the comparisons the recorded tests found significant.

Plan must be shown best (OPTIMALITY_BURDEN, H4c)

The rationale withholds endorsement because nothing shows the plan is the best available or achievable option, placing the burden of proof on the plan and pointing officials toward seeking a better alternative instead.

Two-by-two table of the percentage of rationales coded plan must be shown best (OPTIMALITY_BURDEN). Gain frame: Claude Opus 5, 21% (8 of 39); Claude Sonnet 5.5, 58% (23 of 40). Loss frame: Claude Opus 5, 25% (10 of 40); Claude Sonnet 5.5, 19% (7 of 37). Row totals: Gain frame, 39% (31 of 79); Loss frame, 22% (17 of 77). Column totals: Claude Opus 5, 23% (18 of 79); Claude Sonnet 5.5, 39% (30 of 77). The grand-total corner is blank.
Claude Sonnet 5.5 wrote this in 23 of 40 gain-frame rationales against 7 of 37 loss-frame ones, more often under the gain frame, the opposite of the prediction (Holm-adjusted p = .004), so H4c was falsified. Claude Opus 5 wrote it in 8 of 39 and 10 of 40.
Grouped bar chart of the percentage of rationales coded plan must be shown best (OPTIMALITY_BURDEN), by Frame and Model, without error bars. Gain frame: Claude Opus 5 21%, Claude Sonnet 5.5 58%. Loss frame: Claude Opus 5 25%, Claude Sonnet 5.5 19%. Brackets mark the two models in the gain frame, Claude Sonnet 5.5 across the frames and the main effect of Frame.
The recorded percentages as grouped bars, with brackets on the comparisons the recorded tests found significant.
Line chart of the percentage of rationales coded plan must be shown best (OPTIMALITY_BURDEN), one line per model across the gain and loss frames, without error bars. Claude Opus 5 rises from 21% (8 of 39) to 25% (10 of 40); Claude Sonnet 5.5 falls from 58% (23 of 40) to 19% (7 of 37).
The same recorded percentages as lines, one per model across the two frames.

Note. Each cell gives the percentage of coded rationales showing the code, with the count out of the rationales coded in that condition: 39, 40, 40 and 37 for Gain frame with Claude Opus 5, Gain frame with Claude Sonnet 5.5, Loss frame with Claude Opus 5 and Loss frame with Claude Sonnet 5.5. The recorded view gives no interval for these percentages, so the charts draw no error bars. Brackets mark the comparisons the recorded tests found significant.

Manipulation checks

Every answered check was correct, so no check failed in any condition; the only missing answers belong to sessions that ended in an ambiguous provider call.

Manipulation checks, passed out of scheduled units
CheckGain frame, Claude Opus 5Gain frame, Claude Sonnet 5.5Loss frame, Claude Opus 5Loss frame, Claude Sonnet 5.5
Protected count
300 in every condition
39 of 40 passed
0 failed, 1 missing (ambiguous call)
40 of 40 passed
0 failed
40 of 40 passed
0 failed
38 of 40 passed
0 failed, 2 missing (ambiguous call)
Unprotected count
600 in every condition
39 of 40 passed
0 failed, 1 missing (ambiguous call)
40 of 40 passed
0 failed
40 of 40 passed
0 failed
38 of 40 passed
0 failed, 2 missing (ambiguous call)

Each sample on its own

The two samples are two complete copies of the design, collected in the same run and interleaved, so that the effect can be seen in each independently. In both samples each model is lower under the loss frame, and Claude Sonnet 5.5 is lower than Claude Opus 5 under each frame. Claude Sonnet 5.5 answered 5 in every gain-frame session of both samples, and Claude Opus 5 answered 6 in every gain-frame session of sample 2. The one missing session fell in sample 1.

Mean likelihood of recommending Plan A (1 to 7), by sample
SampleGain frame, Claude Opus 5Gain frame, Claude Sonnet 5.5Loss frame, Claude Opus 5Loss frame, Claude Sonnet 5.5
Sample 1M = 5.9
SD = 0.32, n = 19
M = 5.0
SD = 0.00, n = 20
M = 5.2
SD = 0.62, n = 20
M = 4.0
SD = 0.46, n = 20
Sample 2M = 6.0
SD = 0.00, n = 20
M = 5.0
SD = 0.00, n = 20
M = 5.4
SD = 0.59, n = 20
M = 3.8
SD = 0.41, n = 20
BothM = 5.9
SD = 0.22, n = 39
M = 5.0
SD = 0.00, n = 40
M = 5.3
SD = 0.60, n = 40
M = 3.9
SD = 0.44, n = 40

What we found and what was falsified

The ten preregistered hypotheses and their recorded verdicts
HypothesisStatementOutcomeRecorded resultVerdict
H1aClaude Sonnet 5.5 rates the likelihood of recommending Plan A higher under the gain frame than under the loss frame.Likelihood (primary)Gain frame: mean 5.0 (n = 40). Loss frame: mean 3.9 (n = 40). Holm-adjusted p < .001.Confirmed
H1bClaude Opus 5 rates the likelihood of recommending Plan A higher under the gain frame than under the loss frame.Likelihood (primary)Gain frame: mean 5.9 (n = 39). Loss frame: mean 5.3 (n = 40). Holm-adjusted p < .001.Confirmed
H2aClaude Sonnet 5.5 rates its confidence higher under the gain frame than under the loss frame.ConfidenceGain frame: mean 3.7 (n = 40). Loss frame: mean 4.1 (n = 38). Holm-adjusted p = .034. The difference is in the opposite direction.Falsified
H2bClaude Opus 5 rates its confidence higher under the gain frame than under the loss frame.ConfidenceGain frame: mean 5.1 (n = 39). Loss frame: mean 4.4 (n = 40). Holm-adjusted p < .001.Confirmed
H3aUnder the gain frame, Claude Opus 5 rates its confidence higher than Claude Sonnet 5.5 does.ConfidenceClaude Opus 5: mean 5.1 (n = 39). Claude Sonnet 5.5: mean 3.7 (n = 40). Holm-adjusted p < .001.Confirmed
H3bUnder the loss frame, Claude Opus 5 rates its confidence higher than Claude Sonnet 5.5 does.ConfidenceClaude Opus 5: mean 4.4 (n = 40). Claude Sonnet 5.5: mean 4.1 (n = 38). Holm-adjusted p = .25.Falsified
H4aClaude Sonnet 5.5's rationale compares Plan A with doing nothing (the floor) more often under the gain frame than under the loss frame.Rationale code BETTER_THAN_NOTHINGGain frame: 19 of 40 (48%). Loss frame: 9 of 37 (24%). Holm-adjusted p = .12.Falsified
H4bClaude Sonnet 5.5's rationale treats the unprotected share as the verdict more often under the loss frame than under the gain frame.Rationale code SHORTFALLLoss frame: 9 of 37 (24%). Gain frame: 10 of 40 (25%). Holm-adjusted p = 1.00. The difference is in the opposite direction.Falsified
H4cClaude Sonnet 5.5's rationale demands that the plan be shown best more often under the loss frame than under the gain frame.Rationale code OPTIMALITY_BURDENLoss frame: 7 of 37 (19%). Gain frame: 23 of 40 (58%). Holm-adjusted p = .004. The difference is in the opposite direction.Falsified
H4dClaude Opus 5's rationale recognises the framing more often under the loss frame than under the gain frame.Rationale code FRAME_EQUIVALENCELoss frame: 28 of 40 (70%). Gain frame: 8 of 39 (21%). Holm-adjusted p < .001.Confirmed

The registered linear model, fitted to every analysed reply with one intercept per sample, recorded these omnibus terms.

Omnibus terms of the registered linear model, as recorded (Wald t on 154 residual degrees of freedom)
TermbSEt(154)p
Main effect of Frame0.890.0614.32p < .001
Main effect of Model1.160.0618.76p < .001
Interaction−0.430.12−3.44p < .001

Five hypotheses were falsified, and each is one whose origin carried a FRAME-03 verdict for Claude Sonnet 5 to Claude Sonnet 5.5 as an assumption. Claude Sonnet 5.5 rated its confidence higher under the loss frame, not lower (3.7 against 4.1, Holm-adjusted p = .034), so H2a is falsified, and under the loss frame Claude Opus 5 was not significantly more confident than it (4.4 against 4.1, Holm-adjusted p = .25), so H3b is falsified. On the codes, Claude Sonnet 5.5 compared the plan with doing nothing more often under the gain frame but not significantly so (19 of 40 against 9 of 37, Holm-adjusted p = .12, H4a), judged the plan by the unprotected share about equally under both frames (10 of 40 and 9 of 37, Holm-adjusted p = 1.00, H4b), and demanded that the plan be shown best more often under the gain frame, the opposite of the prediction (23 of 40 against 7 of 37, Holm-adjusted p = .004, H4c).

Nothing else was tested as a hypothesis. The omnibus terms, the other simple effects and the unopposed benefit code (NO_STATED_DOWNSIDE) are reported as the product recorded them, without a verdict. Every registered sensitivity bound for the missing sessions gives the same ten verdicts.

How the run went

The run pack was prepared on 2026-09-28 from the FRAME-03 inputs, changing only the second model level to Claude Sonnet 5.5 and the coder to Claude Sonnet 5, which the owner chose at 17:37 Pacific so that the coder is not a participant of the study it codes. The preparation used the product's Prepare replication from FRAME-03, so each registration names FRAME-03 as its predecessor with an origin for every hypothesis.

Four earlier registrations of this successor stopped on 2026-09-29 before any confirmatory session, and each is kept as a refused record with its data. The first passed its preflight at 01:31 Pacific on the pinned Claude Code CLI 2.1.233 and stopped at the first Claude Sonnet 5.5 reply of the measure pilot, because that command line predates the model: it reported the reply under the identity of Claude Sonnet 5 and recorded a hidden side call, which the product's evidence rule refuses. Under the owner's permission the product was moved to Claude Code CLI 2.1.281, whose requests carry different injected context, and the next three attempts each stopped on a difference the new command line brought: a coder probe folder whose printed path the product's path check could not match, the command line's replay of the user message in its output stream, and that replay arriving behind hidden thinking. Each was fixed in the product, and every fix, with two rules for the command line's capture files, was reviewed by an independent agent before the fifth attempt.

The fifth registration, AGENT-C218672F0B8877E662FF844F30FFC381, was registered on 2026-09-29 as the named successor of FRAME-03 with ten origins, on software release c761c27d with Claude Code CLI 2.1.281 pinned, and its run started at 16:23 Pacific. The preflight passed all three routes, Claude Opus 5, Claude Sonnet 5.5 and the Claude Sonnet 5 coder, and the coder agreed with the expected codes on 39 of 40 calibration cases, above both registered thresholds. The measure pilot of 40 sessions launched at 16:26. One of its sessions, in Gain frame with Claude Sonnet 5.5, ended in an ambiguous call at 16:28 and was reconciled as an execution failure with the researcher's key, never resent; the pilot completed by 16:34 with 9 invalid items of 360, all of them that session's recorded missingness, so the registered go rule held and the measure decision was signed. The execution pilot of 4 sessions completed by 16:43, and the preregistration was finalized at 16:43 before any result was viewed; its one amendment, the finalization itself, was not informed by outcomes. The confirmatory stage launched at 16:45.

The 160 confirmatory sessions ran in four batches. Three ended in an ambiguous provider call after the call had been captured and were reconciled as execution failures with the researcher's key; none was resent. One of them, in Gain frame with Claude Opus 5 in sample 1, has no likelihood answer and is counted as missing under the registered rule, so 159 of 160 were analysed. The other two, both in Loss frame with Claude Sonnet 5.5, answered the likelihood and lack only later answers.

The close of the confirmatory stage was drafted at 18:22 and signed between 18:31 and 18:56, and the recorded analysis was committed at 19:16 Pacific on 2026-09-29 with result digest sha256:557362dfff53858ebb0af0f072f0e29d4622afdcf39612ff7610c1b972577e95. The disclosure of the result is recorded in the ledger. The app's route answered after its two-minute limit without the view, and the results bridge could not capture the same opening again once it had been granted, so the view on this page is the committed analysis view itself, which carries the same result digest.

Against FRAME-03

FRAME-04 kept the FRAME-03 scenario, wording, scale, codes, checks, sizes and rules, and changed the second model from Claude Sonnet 5 to Claude Sonnet 5.5. Other things around the design also differ. The rationales were coded by Claude Sonnet 5, where FRAME-03's coder was Claude Opus 5, so the code shares of the two studies come from different coders. The route pinned Claude Code CLI 2.1.281, where FRAME-03 pinned 2.1.233; the newer command line adds an environment block and a model statement to the context it injects and sends the budget and the date in a different form. It sends that context bare to Claude Sonnet 5.5, where the older one wrapped its budget message in system-reminder tags for Claude Sonnet 5. The effort setting does not differ. Both registrations declare a low effort for every level and for the coder route, and every captured request of both studies carries that low effort, 1433 of 1433 real calls in FRAME-03 and 1450 of 1450 in FRAME-04, counted from each study's own custody after the record. What changed is the checking: FRAME-03's release never compared the effort a call carried with the effort its level declared, and FRAME-04's release refuses a call whose effort differs from its declaration.

What Claude Sonnet 5.5 did against Claude Sonnet 5

The likelihood effect held. In FRAME-03, Claude Sonnet 5 gave a mean of 5.1 under the gain frame and 3.4 under the loss frame; in FRAME-04, Claude Sonnet 5.5 gave 5.0 and 3.9. Both studies confirm H1a. Claude Sonnet 5.5 answered 5 in all 40 of its gain-frame sessions, where Claude Sonnet 5 answered 5 in 33 of 40, and under the loss frame it answered 4, the midpoint, in 32 of 40, where Claude Sonnet 5 answered 3 in 25 of 40.

Replies at each point of the likelihood scale, from 1 (not at all likely) to 7 (very likely)
Condition1234567n
Gain frame, Claude Opus 50000237039
Gain frame, Claude Sonnet 5.50000400040
Loss frame, Claude Opus 500032314040
Loss frame, Claude Sonnet 5.50063220040

The confidence rating moved the other way. Claude Sonnet 5's confidence fell with the frame (3.6 against 2.5, H2a confirmed); Claude Sonnet 5.5's rose (3.7 against 4.1, H2a falsified). Under the loss frame Claude Opus 5 was more confident than Claude Sonnet 5 in FRAME-03 (4.6 against 2.5, H3b confirmed) and not significantly more confident than Claude Sonnet 5.5 in FRAME-04 (4.4 against 4.1, H3b falsified).

The rationale codes differ. Claude Sonnet 5 judged the plan by the unprotected share in 37 of 40 loss-frame rationales and 15 of 38 gain-frame ones; Claude Sonnet 5.5 did so in 9 of 37 and 10 of 40. Claude Sonnet 5 asked for the plan to be shown best in 28 of 40 loss-frame rationales and 2 of 38 gain-frame ones; Claude Sonnet 5.5 in 7 of 37 and 23 of 40. Claude Sonnet 5 never recognised the framing (0 of 38 and 0 of 40); Claude Sonnet 5.5 recognised it in 17 of 37 loss-frame rationales. Claude Sonnet 5 compared the plan with doing nothing in 31 of 38 gain-frame rationales and 5 of 40 loss-frame ones; Claude Sonnet 5.5 in 19 of 40 and 9 of 37. Because the coders differ, part of any difference in these shares may belong to the coder rather than the model.

What stayed the same for Claude Opus 5

Claude Opus 5 rated Plan A less likely under the loss frame in both studies: 5.8 against 5.1 in FRAME-03 and 5.9 against 5.3 in FRAME-04, H1b confirmed each time. It was more confident than the other model under the gain frame in both (H3a confirmed each time), and it recognised the framing more often under the loss frame in both: 35 of 39 against 15 of 37 in FRAME-03 and 28 of 40 against 8 of 39 in FRAME-04 (H4d confirmed each time). One result did not repeat: its confidence did not move with the frame in FRAME-03 (4.7 against 4.6, Holm-adjusted p = .67, H2b falsified) and did in FRAME-04 (5.1 against 4.4, Holm-adjusted p < .001, H2b confirmed).

The comparison cannot claim more than that. No test was run between the two studies; they ran on different days, with different coders and different pinned command lines, and the change of H2b from falsified to confirmed is a difference between two studies, not a measured change in Claude Opus 5. Every number in this comparison is one that the two recorded views state.

What this says and does not say

Within this scenario, this wording and this route, both models rate Plan A less likely to be recommended when the outcome is described as people left unprotected, and Claude Sonnet 5.5 moves further than Claude Opus 5. The result is about these two models on 2026-09-29 on the pinned Claude Code CLI 2.1.281, with the registered task instructions. It says nothing about other scenarios, other providers or human participants, and it does not show that Claude Sonnet 5.5 answers as Claude Sonnet 5 did: all five falsified hypotheses are among the seven predictions carried from one model to the other.

The rationale coder was Claude Sonnet 5, a single model coder that is not a FRAME-04 participant, blind to condition and level; coding by a single coder is a declared limitation, and every reply and every coder reply is archived for a later recode. The route sent its environment, model, budget and date context bare to Claude Sonnet 5.5 and Claude Opus 5 and wrapped in system-reminder tags to the Claude Sonnet 5 coder; the study conversation and the instruction were identical. Hidden thinking, which the route does not return, was present in 86 of 234 and 90 of 240 of Claude Opus 5's calls under the gain and loss frames, and in 12 of 240 and 15 of 233 of Claude Sonnet 5.5's, and is not part of the analysis.