FRAME-03: the seven-point replication
FRAME-03 ran the FRAME-02 design again through OpenPsy on 2026-09-28, with the recommendation asked as a likelihood from 1 to 7 instead of yes or no, in 160 preregistered sessions in two samples of 20 per cell. Both models rated Plan A less likely to be recommended under the loss frame: Claude Sonnet 5 from 5.1 to 3.4 and Claude Opus 5 from 5.8 to 5.1, each with a Holm-adjusted p less than .001. Nine of the ten preregistered hypotheses were confirmed. One was falsified: Claude Opus 5 did not rate its confidence higher under the gain frame (4.7 against 4.6, Holm-adjusted p = .67).
What was asked
The design is the FRAME-02 design: a two by two of Frame (gain, "300 of those people will be protected", against loss, "600 of those people will be left unprotected") by Model (Claude Opus 5 against Claude Sonnet 5 on one Claude Code CLI route with identical settings), with language model agents as the participants and one fresh instance per session. What changed is how the answer is taken. The primary outcome is the likelihood of recommending Plan A on a scale from 1 (not at all likely) to 7 (very likely), where 4 means neither likely nor unlikely. The confidence rating, also from 1 to 7, became a tested secondary outcome. Each free-text rationale was coded under five codes taken from the FRAME-02 post hoc coding by a registered model coder, Claude Opus 5, which saw the codebook and one reply at a time and never a condition, a model name or an outcome. Two manipulation checks close every session, so each session made six participant calls and one coder call.
The preregistration registered ten hypotheses. On the likelihood, H1a states that Claude Sonnet 5 rates the likelihood of recommending Plan A higher under the gain frame than under the loss frame, and H1b states the same for Claude Opus 5. On the confidence rating, H2a and H2b make the same prediction for Claude Sonnet 5 and for Claude Opus 5, and H3a and H3b state that Claude Opus 5 rates its confidence higher than Claude Sonnet 5 does, under the gain frame and under the loss frame. On the rationale codes, H4a states that Claude Sonnet 5's rationale compares Plan A with doing nothing (BETTER_THAN_NOTHING) more often under the gain frame, H4b and H4c state that its rationale treats the unprotected share as the verdict (SHORTFALL) and demands that the plan be shown best (OPTIMALITY_BURDEN) more often under the loss frame, and H4d states that Claude Opus 5's rationale recognises the framing (FRAME_EQUIVALENCE) more often under the loss frame. The fifth code, unopposed benefit (NO_STATED_DOWNSIDE), carries no hypothesis.
The likelihood and confidence hypotheses are judged by the exact permutation test of the difference in means, and the code hypotheses by the exact conditional test. Each test is stratified by sample and Holm-adjusted across the four simple effects of its own outcome, at alpha .05. A result that is not significant in the predicted direction counts against the hypothesis. The confirmatory run was 160 sessions in two samples of 20 per cell, so 40 per condition in all, after a measure pilot of 40 sessions and an execution pilot of 4.
The recorded result
Note. N = 160 preregistered confirmatory sessions in two samples, n = 40 per condition (20 per sample). Four sessions that ended in an ambiguous provider call before the likelihood answer were excluded as missing under the registered rule, so the analysed n is 37, 40, 39 and 40 for Gain frame with Claude Opus 5, Gain frame with Claude Sonnet 5, Loss frame with Claude Opus 5 and Loss frame with Claude Sonnet 5. Each cell gives the mean likelihood on the 1 to 7 scale with its SD and n, and the mean confidence rating. The recorded 95% confidence intervals of the means, from Student's t, are drawn as error bars on the charts below and stated in the text.
Study AGENT-4077193D973E70CDF2A39DCC053D80C5. Registration hash sha256:b2ce589e323d8aa833f3176065b5f1bdae33cb59eae3050644d58627aa926b8a. Result digest sha256:d4a4431898f3834d01808cfc4c8801e432b30628f92df89841c7d65182ae2336. Software release 6aa25d67.
Under the gain frame, Claude Opus 5 had a mean of 5.8, 95% CI [5.6, 6.0], and Claude Sonnet 5 a mean of 5.1, 95% CI [4.9, 5.2]. Under the loss frame, Claude Opus 5 had a mean of 5.1, 95% CI [4.9, 5.3], and Claude Sonnet 5 a mean of 3.4, 95% CI [3.1, 3.6]. Claude Sonnet 5's line falls more steeply than Claude Opus 5's, which is the interaction.
Confidence rating
Note. Each cell gives the mean confidence rating on the 1 to 7 scale and its n, the sessions that answered the rating: 37, 39, 39 and 40 for Gain frame with Claude Opus 5, Gain frame with Claude Sonnet 5, Loss frame with Claude Opus 5 and Loss frame with Claude Sonnet 5. The recorded view gives no SD and no interval for these means, so the table shows neither and the charts draw no error bars. Brackets mark the comparisons the registered tests found significant.
Rationale codes
Each free-text rationale was coded once by Claude Opus 5 on its own pinned route, blind to the condition and the model level of the reply. Before any reply of this study existed, the coder agreed with the expected codes on 40 of 40 calibration cases, and no coded reply named a condition or a model in its own words. Four codes carry a hypothesis, H4a to H4d, and one, unopposed benefit (NO_STATED_DOWNSIDE), is reported without one. Each code's four simple effects are Holm-adjusted within that code only.
Unprotected share as the verdict (SHORTFALL, H4b)
The rationale judges the plan by the size of the group it leaves unprotected (600, two-thirds, the majority, only a third protected) and treats that shortfall as a poor or inadequate result in its own right.
Note. Each cell gives the percentage of coded rationales showing the code, with the count out of the rationales coded in that condition: 37, 38, 39 and 40 for Gain frame with Claude Opus 5, Gain frame with Claude Sonnet 5, Loss frame with Claude Opus 5 and Loss frame with Claude Sonnet 5. The recorded view gives no interval for these percentages, so the charts draw no error bars. Brackets mark the comparisons the recorded tests found significant.
Improvement over inaction (BETTER_THAN_NOTHING, H4a)
The rationale compares the plan with an explicit or implied do-nothing baseline and concludes that some protection is better than none.
Note. Each cell gives the percentage of coded rationales showing the code, with the count out of the rationales coded in that condition: 37, 38, 39 and 40 for Gain frame with Claude Opus 5, Gain frame with Claude Sonnet 5, Loss frame with Claude Opus 5 and Loss frame with Claude Sonnet 5. The recorded view gives no interval for these percentages, so the charts draw no error bars. Brackets mark the comparisons the recorded tests found significant.
Framing recognised (FRAME_EQUIVALENCE, H4d)
The rationale states that the loss and gain descriptions are the same outcome, names the framing or wording effect, or says the recommendation would not change under the other wording.
Note. Each cell gives the percentage of coded rationales showing the code, with the count out of the rationales coded in that condition: 37, 38, 39 and 40 for Gain frame with Claude Opus 5, Gain frame with Claude Sonnet 5, Loss frame with Claude Opus 5 and Loss frame with Claude Sonnet 5. The recorded view gives no interval for these percentages, so the charts draw no error bars. Brackets mark the comparisons the recorded tests found significant.
Unopposed benefit (NO_STATED_DOWNSIDE, no hypothesis)
The rationale treats the absence of any stated cost, risk, harm or competing plan as a reason the plan's benefit stands, placing the burden on an objection that has not been raised.
Note. Each cell gives the percentage of coded rationales showing the code, with the count out of the rationales coded in that condition: 37, 38, 39 and 40 for Gain frame with Claude Opus 5, Gain frame with Claude Sonnet 5, Loss frame with Claude Opus 5 and Loss frame with Claude Sonnet 5. The recorded view gives no interval for these percentages, so the charts draw no error bars. Brackets mark the comparisons the recorded tests found significant.
Plan must be shown best (OPTIMALITY_BURDEN, H4c)
The rationale withholds endorsement because nothing shows the plan is the best available or achievable option, placing the burden of proof on the plan and pointing officials toward seeking a better alternative instead.
Note. Each cell gives the percentage of coded rationales showing the code, with the count out of the rationales coded in that condition: 37, 38, 39 and 40 for Gain frame with Claude Opus 5, Gain frame with Claude Sonnet 5, Loss frame with Claude Opus 5 and Loss frame with Claude Sonnet 5. The recorded view gives no interval for these percentages, so the charts draw no error bars. Brackets mark the comparisons the recorded tests found significant.
Manipulation checks
Every answered check was correct, so no check failed in any condition; the only missing answers belong to sessions that ended in an ambiguous provider call.
| Check | Gain frame, Claude Opus 5 | Gain frame, Claude Sonnet 5 | Loss frame, Claude Opus 5 | Loss frame, Claude Sonnet 5 |
|---|---|---|---|---|
| Protected count 300 in every condition | 37 of 40 passed 0 failed, 3 missing (ambiguous call) | 39 of 40 passed 0 failed, 1 missing (ambiguous call) | 39 of 40 passed 0 failed, 1 missing (ambiguous call) | 40 of 40 passed 0 failed |
| Unprotected count 600 in every condition | 37 of 40 passed 0 failed, 3 missing (ambiguous call) | 38 of 40 passed 0 failed, 2 missing (ambiguous call) | 39 of 40 passed 0 failed, 1 missing (ambiguous call) | 40 of 40 passed 0 failed |
Each sample on its own
The two samples are two complete copies of the design, collected in the same run and interleaved, so that the effect can be seen in each independently. In both samples each model is lower under the loss frame, and Claude Sonnet 5 is lower than Claude Opus 5 under each frame. All four missing sessions fell in sample 1.
| Sample | Gain frame, Claude Opus 5 | Gain frame, Claude Sonnet 5 | Loss frame, Claude Opus 5 | Loss frame, Claude Sonnet 5 |
|---|---|---|---|---|
| Sample 1 | M = 5.9 SD = 0.56, n = 17 | M = 5.1 SD = 0.45, n = 20 | M = 4.9 SD = 0.71, n = 19 | M = 3.3 SD = 0.80, n = 20 |
| Sample 2 | M = 5.7 SD = 0.47, n = 20 | M = 5.1 SD = 0.39, n = 20 | M = 5.3 SD = 0.64, n = 20 | M = 3.5 SD = 0.69, n = 20 |
| Both | M = 5.8 SD = 0.52, n = 37 | M = 5.1 SD = 0.42, n = 40 | M = 5.1 SD = 0.68, n = 39 | M = 3.4 SD = 0.74, n = 40 |
What we found and what was falsified
| Hypothesis | Statement | Outcome | Recorded result | Verdict |
|---|---|---|---|---|
| H1a | Claude Sonnet 5 rates the likelihood of recommending Plan A higher under the gain frame than under the loss frame. | Likelihood (primary) | Gain frame: mean 5.1 (n = 40). Loss frame: mean 3.4 (n = 40). Holm-adjusted p < .001. | Confirmed |
| H1b | Claude Opus 5 rates the likelihood of recommending Plan A higher under the gain frame than under the loss frame. | Likelihood (primary) | Gain frame: mean 5.8 (n = 37). Loss frame: mean 5.1 (n = 39). Holm-adjusted p < .001. | Confirmed |
| H2a | Claude Sonnet 5 rates its confidence higher under the gain frame than under the loss frame. | Confidence | Gain frame: mean 3.6 (n = 39). Loss frame: mean 2.5 (n = 40). Holm-adjusted p < .001. | Confirmed |
| H2b | Claude Opus 5 rates its confidence higher under the gain frame than under the loss frame. | Confidence | Gain frame: mean 4.7 (n = 37). Loss frame: mean 4.6 (n = 39). Holm-adjusted p = .67. | Falsified |
| H3a | Under the gain frame, Claude Opus 5 rates its confidence higher than Claude Sonnet 5 does. | Confidence | Claude Opus 5: mean 4.7 (n = 37). Claude Sonnet 5: mean 3.6 (n = 39). Holm-adjusted p < .001. | Confirmed |
| H3b | Under the loss frame, Claude Opus 5 rates its confidence higher than Claude Sonnet 5 does. | Confidence | Claude Opus 5: mean 4.6 (n = 39). Claude Sonnet 5: mean 2.5 (n = 40). Holm-adjusted p < .001. | Confirmed |
| H4a | Claude Sonnet 5's rationale compares Plan A with doing nothing (the floor) more often under the gain frame than under the loss frame. | Rationale code BETTER_THAN_NOTHING | Gain frame: 31 of 38 (82%). Loss frame: 5 of 40 (13%). Holm-adjusted p < .001. | Confirmed |
| H4b | Claude Sonnet 5's rationale treats the unprotected share as the verdict more often under the loss frame than under the gain frame. | Rationale code SHORTFALL | Loss frame: 37 of 40 (93%). Gain frame: 15 of 38 (39%). Holm-adjusted p < .001. | Confirmed |
| H4c | Claude Sonnet 5's rationale demands that the plan be shown best more often under the loss frame than under the gain frame. | Rationale code OPTIMALITY_BURDEN | Loss frame: 28 of 40 (70%). Gain frame: 2 of 38 (5%). Holm-adjusted p < .001. | Confirmed |
| H4d | Claude Opus 5's rationale recognises the framing more often under the loss frame than under the gain frame. | Rationale code FRAME_EQUIVALENCE | Loss frame: 35 of 39 (90%). Gain frame: 15 of 37 (41%). Holm-adjusted p < .001. | Confirmed |
The registered linear model, fitted to every analysed reply with one intercept per sample, recorded these omnibus terms.
| Term | b | SE | t(151) | p |
|---|---|---|---|---|
| Main effect of Frame | 1.20 | 0.10 | 12.41 | p < .001 |
| Main effect of Model | 1.23 | 0.10 | 12.68 | p < .001 |
| Interaction | −0.99 | 0.19 | −5.12 | p < .001 |
One hypothesis was falsified. H2b predicted that Claude Opus 5 would rate its confidence higher under the gain frame than under the loss frame. It did not: its mean confidence was 4.7 against 4.6 (Holm-adjusted p = .67), so under the registered rule H2b is falsified. Claude Opus 5 rated Plan A less likely under the loss frame without becoming less confident in its answer.
Nothing else was tested as a hypothesis. The omnibus terms, the other simple effects and the unopposed benefit code (NO_STATED_DOWNSIDE) are reported as the product recorded them, without a verdict. Every registered sensitivity bound for the missing sessions gives the same ten verdicts.
How the run went
The run was planned with Claude Sonnet 4.5 as the rationale coder, a model other than the two participants. Two real preflights refused that route before any session ran: the first on the dated model name the provider reports, which the product was then changed to accept, and the second on the shape of the messages the pinned command line sends to that model, which the product cannot yet declare. Under the owner's fallback of 2026-09-25, the coder became Claude Opus 5, the same model as one participant level, blind to level and condition. The next preflight passed, and the coder agreed with the expected codes on 40 of 40 calibration cases.
The first registration with that coder, AGENT-3B1C33B3F617B569A75482CEC6C4213B, ran its measure pilot at four sessions per cell. One call met the provider's 529 overloaded error; the session was reconciled as an execution failure, never resent, and the pilot completed. The registered go rule then refused the measure decision, because it counts that session against parse validity and so left its cell at 3 of 4, below 90 percent. The refusal was recorded and signed in that registration, which is kept with all its data. The same design was registered again, unchanged except that the measure pilot returned to ten sessions per cell, the programme default of FRAME-01 and FRAME-02, as AGENT-4077193D973E70CDF2A39DCC053D80C5.
That registration started at 01:24 on 2026-09-28, Pacific time. The measure pilot of 40 sessions had no invalid item of 360, and its decision was approved and signed at 01:37. The execution pilot of 4 sessions passed, the preregistration was finalized at 01:40, and the confirmatory stage launched at 01:41. The preregistration was finalized before any result was viewed, and its one amendment, the finalization itself, was not informed by outcomes.
Six of the 160 confirmatory sessions ended in an ambiguous provider call and were reconciled as execution failures with the researcher's key; none was resent. One connection was refused on the scenario call, four connections dropped at the same instant (two on the likelihood answer, one on the confidence rating and one on the unprotected count check), and one scenario call did not answer within the registered 240 seconds. That timeout placed a collection hold; the product's own route probe verified the transport again, and the hold was cleared through the product before the run resumed. Four of the six sessions have no likelihood answer and are counted as missing under the registered rule, so 156 of 160 were analysed: three in Gain frame with Claude Opus 5 and one in Loss frame with Claude Opus 5, all in sample 1. The other two answered the likelihood and lack only a later answer. The coded outcomes count six sessions as missing and analyse 154.
The confirmatory stage completed at 03:27, and the recorded analysis was committed at 03:48 Pacific on 2026-09-28 with result digest sha256:d4a4431898f3834d01808cfc4c8801e432b30628f92df89841c7d65182ae2336. The disclosure of the result is recorded in the ledger. The app's two-minute route limit ended the first attempts to open the result before the view arrived, so the recorded view was captured from the product's results bridge with a longer wait; it carries the same result digest, and the capture added no event to the ledger.
Against FRAME-02
FRAME-03 kept the FRAME-02 scenario, wording, models and route, and changed these things: the recommendation is asked as a likelihood from 1 to 7 instead of yes or no, the confidence rating is tested instead of exploratory, the rationale is coded under five registered codes by a coder call, and two manipulation checks close every session, so each session made six participant calls against four in FRAME-02. It is also smaller, 20 sessions per cell in each of two samples against 40.
What the seven-point scale showed
In FRAME-02, Claude Opus 5 recommended Plan A in 79 of 80 gain-frame sessions and 76 of 77 loss-frame sessions. At that ceiling the yes or no outcome showed no frame effect (Holm-adjusted p = 1.00), and its hypothesis H1b was recorded as falsified. On the scale, Claude Opus 5 gave a mean of 5.8 under the gain frame and 5.1 under the loss frame, and the corresponding hypothesis, H1b, was confirmed (Holm-adjusted p < .001). The scale made visible a graded frame effect within Claude Opus 5 that the binary outcome could not show. It still leaned towards recommending under both frames, with no reply below 4, but less strongly under the loss frame, where 7 of its 39 replies sat at the midpoint of 4 against none of its 37 gain-frame replies.
| Condition | 1 | 2 | 3 | 4 | 5 | 6 | 7 | n |
|---|---|---|---|---|---|---|---|---|
| Gain frame, Claude Opus 5 | 0 | 0 | 0 | 0 | 9 | 26 | 2 | 37 |
| Gain frame, Claude Sonnet 5 | 0 | 0 | 0 | 2 | 33 | 5 | 0 | 40 |
| Loss frame, Claude Opus 5 | 0 | 0 | 0 | 7 | 21 | 11 | 0 | 39 |
| Loss frame, Claude Sonnet 5 | 0 | 2 | 25 | 9 | 4 | 0 | 0 | 40 |
What stayed the same is the Claude Sonnet 5 effect. In FRAME-02 it recommended Plan A in 79 of 79 gain-frame sessions and 11 of 80 loss-frame sessions; in FRAME-03 its mean was 5.1 against 3.4, with 25 of its 40 loss-frame replies at 3. Both studies confirm its frame hypothesis. The scale adds that its gain-frame endorsement was moderate, with 33 of 40 replies at 5, where the yes or no outcome recorded every gain-frame session as a recommendation. Under the gain frame the two models also differ on the scale (5.8 against 5.1, Holm-adjusted p < .001), where FRAME-02 found no difference (79 of 80 against 79 of 79, Holm-adjusted p = 1.00).
The confidence rating and the codes add what the recommendation alone does not. Claude Sonnet 5's confidence fell with the frame (3.6 against 2.5) and Claude Opus 5's did not (4.7 against 4.6). The codes describe what each model wrote: Claude Opus 5 recognised the framing in 35 of 39 loss-frame rationales, and Claude Sonnet 5 never did, while Claude Sonnet 5's loss-frame rationales judged the plan by the unprotected share (37 of 40) and asked for it to be shown best (28 of 40).
The comparison cannot claim more than that. The two studies asked different primary questions, ran on different days with different sample sizes, and no test was run between them. The move of H1b from falsified to confirmed is a difference between two designs, not a measured change in Claude Opus 5. Every number in this comparison is one that the two recorded views state.
What this says and does not say
Within this scenario, this wording and this route, both models rate Plan A less likely to be recommended when the outcome is described as people left unprotected, and Claude Sonnet 5 moves much further than Claude Opus 5. The result is about these two models on 2026-09-28 on the pinned Claude Code CLI, with the registered task instructions. It says nothing about other scenarios, other providers or human participants.
The rationale coder was Claude Opus 5, the same model as one participant level. It was blind to condition and level, but coding by a single model coder is a declared limitation, and every reply and every coder reply is archived for a later recode. The route sent its token budget message wrapped in system-reminder tags to Claude Sonnet 5 and bare to Claude Opus 5; the study conversation and the instruction were identical. Hidden thinking, which the route does not return, was present in 83 of 223 and 90 of 235 of Claude Opus 5's calls under the gain and loss frames, and in 34 of 236 and 25 of 240 of Claude Sonnet 5's, and is not part of the analysis.