FRAME-06: Claude Opus 5 on two Claude Code versions
FRAME-06 ran the FRAME-04 design through OpenPsy on 2026-10-06 with the same model in both columns: Claude Opus 5 on Claude Code 2.1.281, the version FRAME-04 used, and Claude Opus 5 on Claude Code 2.1.233, the version FRAME-03 used, in 160 preregistered sessions in two samples of 20 per cell. On both versions the model rated Plan A less likely to be recommended under the loss frame: from 6.0 to 5.2 on 2.1.281 and from 5.8 to 5.3 on 2.1.233, each with a Holm-adjusted p less than .001. Its confidence fell under the loss frame on 2.1.281 (5.0 against 4.1, Holm-adjusted p less than .001) and fell less on 2.1.233 (4.7 against 4.4, Holm-adjusted p = .093), which is not significant, so H2b was falsified. Under the gain frame the two versions did not differ significantly in confidence (5.0 against 4.7, Holm-adjusted p = .21), so H3 was falsified. Five of the seven preregistered hypotheses were confirmed.
What was asked
FRAME-03 and FRAME-04 asked Claude Opus 5 the same question on different versions of the Claude Code command line. Its confidence in its recommendation did not move significantly with the frame in FRAME-03, on 2.1.233 (4.7 against 4.6, Holm-adjusted p = .67), and did in FRAME-04, on 2.1.281 (5.1 against 4.4, Holm-adjusted p less than .001). FRAME-06 asks whether that confidence effect depends on the command-line version, by running both versions on the same day in one study.
FRAME-06 is the named successor of FRAME-04 (AGENT-C218672F0B8877E662FF844F30FFC381) and keeps its design except for the Model factor, whose two levels are now the same model on two versions. It is a two by two of Frame (gain, "300 of those people will be protected", against loss, "600 of those people will be left unprotected") by Model, with language model agents as the participants and one fresh instance per session. The primary outcome is the likelihood of recommending Plan A on a scale from 1 (not at all likely) to 7 (very likely), the confidence rating from 1 to 7 is a tested secondary outcome, each free-text rationale is coded under FRAME-04's five codes, and two manipulation checks close every session. The coder is Claude Sonnet 5, a registered model that is not a FRAME-06 participant, which saw the codebook and one reply at a time and never a condition, a model name or an outcome.
The preregistration registered seven hypotheses. H1a and H1b state that Claude Opus 5 rates the likelihood of recommending Plan A higher under the gain frame on each version. H2a and H2b state the same for its confidence on each version; the registration carried FRAME-03's result on 2.1.233 as a stated assumption and expected H2b to be falsified. H3 states that under the gain frame its confidence is higher on 2.1.281 than on 2.1.233. H4a and H4b state that on each version its rationale recognises the framing (FRAME_EQUIVALENCE) more often under the loss frame. The other four codes carry no hypothesis.
The likelihood and confidence hypotheses are judged by the exact permutation test of the difference in means, and the code hypotheses by the exact conditional test. Each test is stratified by sample and Holm-adjusted across the four simple effects of its own outcome, at alpha .05. A result that is not significant in the predicted direction counts against the hypothesis. The confirmatory run was 160 sessions in two samples of 20 per cell, after a measure pilot and an execution pilot as registered.
The recorded result
Note. N = 160 preregistered confirmatory sessions in two samples, n = 40 per condition (20 per sample). Each cell gives the mean likelihood on the 1 to 7 scale with its SD and n, and the mean confidence rating. Error bars are the recorded 95% confidence intervals of the means, from Student's t with n − 1 degrees of freedom. Brackets mark the comparisons the registered tests found significant.
Under the gain frame, Claude Opus 5 on Claude Code 2.1.281 had a mean of 6.0, 95% CI [5.9, 6.0], and Claude Opus 5 on Claude Code 2.1.233 a mean of 5.8, 95% CI [5.6, 5.9]. Under the loss frame, Claude Opus 5 on Claude Code 2.1.281 had a mean of 5.2, 95% CI [5.0, 5.3], and Claude Opus 5 on Claude Code 2.1.233 a mean of 5.3, 95% CI [5.1, 5.4]. Under the gain frame the two versions did not differ significantly (Holm-adjusted p = .090). The interaction term of the registered model, b = 0.28, SE = 0.13, t(155) = 2.07, p = .040, says the frame moved the likelihood more on 2.1.281 than on 2.1.233; it is reported without a verdict, as no hypothesis names it.
160 units were scheduled and 160 reserve units were held in reserve. 160 units were analysed, 0 replaced by a reserve unit, 0 excluded and 160 reserve units left unused.
Study AGENT-552121A995F2670603D9B626FADAF1BB. Registration hash sha256:e412722ca3078cc058354151ab8a9677888e37b1f9b0981fa719e9658dbe806c. Result digest sha256:6363aa6402da40c79a6aee3f80273f63edeb6a06228d1bf3e9c2c7ea5ec441d8.
Confidence rating
Note. Each cell gives the mean confidence rating on the 1 to 7 scale with its SD and n; every session answered the rating, so n = 40 in each condition. Error bars are the recorded 95% confidence intervals of the means. Brackets mark the comparisons the registered tests found significant.
Under the gain frame, Claude Opus 5 on Claude Code 2.1.281 had a mean of 5.0, 95% CI [4.8, 5.1], and Claude Opus 5 on Claude Code 2.1.233 a mean of 4.7, 95% CI [4.5, 4.9]. Under the loss frame, Claude Opus 5 on Claude Code 2.1.281 had a mean of 4.1, 95% CI [3.9, 4.3], and Claude Opus 5 on Claude Code 2.1.233 a mean of 4.4, 95% CI [4.1, 4.6].
Rationale codes
Each free-text rationale was coded once by Claude Sonnet 5 on its own pinned route on Claude Code 2.1.281, blind to the condition and the model level of the reply. Before any reply of this study existed, the coder agreed with the expected codes on 40 of 40 calibration cases. One coded reply named a condition or a model in its own words. Two hypotheses, H4a and H4b, name the framing recognised code (FRAME_EQUIVALENCE); the other four codes are reported without one. Each code's four simple effects are Holm-adjusted within that code only. Every session has a coded rationale, so each cell has 40 coded rationales.
Unprotected share as the verdict (SHORTFALL, no hypothesis)
The rationale judges the plan by the size of the group it leaves unprotected (600, two-thirds, the majority, only a third protected) and treats that shortfall as a poor or inadequate result in its own right.
Note. Each cell gives the percentage of coded rationales showing the code, with the count out of the rationales coded in that condition: 40, 40, 40 and 40 for Gain frame with Claude Opus 5 on Claude Code 2.1.281, Gain frame with Claude Opus 5 on Claude Code 2.1.233, Loss frame with Claude Opus 5 on Claude Code 2.1.281 and Loss frame with Claude Opus 5 on Claude Code 2.1.233. The recorded view gives no interval for these percentages, so the charts draw no error bars. Brackets mark the comparisons the recorded tests found significant.
Improvement over inaction (BETTER_THAN_NOTHING, no hypothesis)
The rationale compares the plan with an explicit or implied do-nothing baseline and concludes that some protection is better than none.
Note. Each cell gives the percentage of coded rationales showing the code, with the count out of the rationales coded in that condition: 40, 40, 40 and 40 for Gain frame with Claude Opus 5 on Claude Code 2.1.281, Gain frame with Claude Opus 5 on Claude Code 2.1.233, Loss frame with Claude Opus 5 on Claude Code 2.1.281 and Loss frame with Claude Opus 5 on Claude Code 2.1.233. The recorded view gives no interval for these percentages, so the charts draw no error bars. Brackets mark the comparisons the recorded tests found significant.
Framing recognised (FRAME_EQUIVALENCE, H4a and H4b)
The rationale states that the loss and gain descriptions are the same outcome, names the framing or wording effect, or says the recommendation would not change under the other wording.
Note. Each cell gives the percentage of coded rationales showing the code, with the count out of the rationales coded in that condition: 40, 40, 40 and 40 for Gain frame with Claude Opus 5 on Claude Code 2.1.281, Gain frame with Claude Opus 5 on Claude Code 2.1.233, Loss frame with Claude Opus 5 on Claude Code 2.1.281 and Loss frame with Claude Opus 5 on Claude Code 2.1.233. The recorded view gives no interval for these percentages, so the charts draw no error bars. Brackets mark the comparisons the recorded tests found significant.
Unopposed benefit (NO_STATED_DOWNSIDE, no hypothesis)
The rationale treats the absence of any stated cost, risk, harm or competing plan as a reason the plan's benefit stands, placing the burden on an objection that has not been raised.
Note. Each cell gives the percentage of coded rationales showing the code, with the count out of the rationales coded in that condition: 40, 40, 40 and 40 for Gain frame with Claude Opus 5 on Claude Code 2.1.281, Gain frame with Claude Opus 5 on Claude Code 2.1.233, Loss frame with Claude Opus 5 on Claude Code 2.1.281 and Loss frame with Claude Opus 5 on Claude Code 2.1.233. The recorded view gives no interval for these percentages, so the charts draw no error bars. Brackets mark the comparisons the recorded tests found significant.
Plan must be shown best (OPTIMALITY_BURDEN, no hypothesis)
The rationale withholds endorsement because nothing shows the plan is the best available or achievable option, placing the burden of proof on the plan and pointing officials toward seeking a better alternative instead.
Note. Each cell gives the percentage of coded rationales showing the code, with the count out of the rationales coded in that condition: 40, 40, 40 and 40 for Gain frame with Claude Opus 5 on Claude Code 2.1.281, Gain frame with Claude Opus 5 on Claude Code 2.1.233, Loss frame with Claude Opus 5 on Claude Code 2.1.281 and Loss frame with Claude Opus 5 on Claude Code 2.1.233. The recorded view gives no interval for these percentages, so the charts draw no error bars. Brackets mark the comparisons the recorded tests found significant.
Manipulation checks
Every check was answered and every answer was correct, so all 40 sessions of each condition passed both checks.
| Check | Gain frame, Claude Opus 5 on Claude Code 2.1.281 | Gain frame, Claude Opus 5 on Claude Code 2.1.233 | Loss frame, Claude Opus 5 on Claude Code 2.1.281 | Loss frame, Claude Opus 5 on Claude Code 2.1.233 |
|---|---|---|---|---|
| Protected count 300 in every condition | 40 of 40 passed 0 failed | 40 of 40 passed 0 failed | 40 of 40 passed 0 failed | 40 of 40 passed 0 failed |
| Unprotected count 600 in every condition | 40 of 40 passed 0 failed | 40 of 40 passed 0 failed | 40 of 40 passed 0 failed | 40 of 40 passed 0 failed |
Each sample on its own
The two samples are two complete copies of the design, collected in the same run and interleaved, so that the effect can be seen in each independently. In both samples each version is lower under the loss frame than under the gain frame.
| Sample | Gain frame, Claude Opus 5 on Claude Code 2.1.281 | Gain frame, Claude Opus 5 on Claude Code 2.1.233 | Loss frame, Claude Opus 5 on Claude Code 2.1.281 | Loss frame, Claude Opus 5 on Claude Code 2.1.233 |
|---|---|---|---|---|
| Sample 1 | M = 6.0 SD = 0.00, n = 20 | M = 5.9 SD = 0.37, n = 20 | M = 5.3 SD = 0.44, n = 20 | M = 5.2 SD = 0.49, n = 20 |
| Sample 2 | M = 5.9 SD = 0.31, n = 20 | M = 5.7 SD = 0.47, n = 20 | M = 5.1 SD = 0.22, n = 20 | M = 5.4 SD = 0.67, n = 20 |
| Both | M = 6.0 SD = 0.22, n = 40 | M = 5.8 SD = 0.42, n = 40 | M = 5.2 SD = 0.36, n = 40 | M = 5.3 SD = 0.59, n = 40 |
What we found and what was falsified
| Hypothesis | Statement | Outcome | Recorded result | Verdict |
|---|---|---|---|---|
| H1a | Claude Opus 5 on Claude Code 2.1.281 rates the likelihood of recommending Plan A higher under the gain frame than under the loss frame. | Likelihood (primary) | Gain frame: mean 6.0 (n = 40). Loss frame: mean 5.2 (n = 40). Holm-adjusted p < .001. | Confirmed |
| H1b | Claude Opus 5 on Claude Code 2.1.233 rates the likelihood of recommending Plan A higher under the gain frame than under the loss frame. | Likelihood (primary) | Gain frame: mean 5.8 (n = 40). Loss frame: mean 5.3 (n = 40). Holm-adjusted p < .001. | Confirmed |
| H2a | Claude Opus 5 on Claude Code 2.1.281 rates its confidence higher under the gain frame than under the loss frame. | Confidence | Gain frame: mean 5.0 (n = 40). Loss frame: mean 4.1 (n = 40). Holm-adjusted p < .001. | Confirmed |
| H2b | Claude Opus 5 on Claude Code 2.1.233 rates its confidence higher under the gain frame than under the loss frame. | Confidence | Gain frame: mean 4.7 (n = 40). Loss frame: mean 4.4 (n = 40). Holm-adjusted p = .093. | Falsified |
| H3 | Under the gain frame, Claude Opus 5 on Claude Code 2.1.281 rates its confidence higher than Claude Opus 5 on Claude Code 2.1.233 does. | Confidence | Claude Opus 5 on Claude Code 2.1.281: mean 5.0 (n = 40). Claude Opus 5 on Claude Code 2.1.233: mean 4.7 (n = 40). Holm-adjusted p = .21. | Falsified |
| H4a | Claude Opus 5 on Claude Code 2.1.281's rationale recognises the framing more often under the loss frame than under the gain frame. | Rationale code FRAME_EQUIVALENCE | Loss frame: 27 of 40 (68%). Gain frame: 11 of 40 (28%). Holm-adjusted p = .003. | Confirmed |
| H4b | Claude Opus 5 on Claude Code 2.1.233's rationale recognises the framing more often under the loss frame than under the gain frame. | Rationale code FRAME_EQUIVALENCE | Loss frame: 26 of 40 (65%). Gain frame: 11 of 40 (28%). Holm-adjusted p = .004. | Confirmed |
The registered linear model, fitted to every analysed reply with one intercept per sample, recorded these omnibus terms.
| Term | b | SE | t(155) | p |
|---|---|---|---|---|
| Main effect of Frame | 0.66 | 0.07 | 9.98 | p < .001 |
| Main effect of Model | 0.04 | 0.07 | 0.56 | p = .57 |
| Interaction | 0.28 | 0.13 | 2.07 | p = .040 |
Two hypotheses were falsified, both on the confidence rating. On Claude Code 2.1.233 Claude Opus 5's confidence was lower under the loss frame but not significantly so (4.7 against 4.4, Holm-adjusted p = .093, H2b), which is the outcome the registration expected. Under the gain frame its confidence on 2.1.281 was not significantly higher than on 2.1.233 (5.0 against 4.7, Holm-adjusted p = .21, H3). So the confidence effect appears on 2.1.281 and is weaker on 2.1.233, and this study did not detect a direct difference between the two versions.
Nothing else was tested as a hypothesis. The omnibus terms, the other simple effects and the four codes without a hypothesis are reported as the product recorded them, without a verdict. No reply was invalid, refused or truncated, so every registered sensitivity bound equals the observed mean, and the minimum, maximum and least favourable re-scorings give the same verdicts for the five likelihood and confidence hypotheses.
How the run went
The sessions ran under registration AGENT-552121A995F2670603D9B626FADAF1BB, version 4, with content hash sha256:e412722ca3078cc058354151ab8a9677888e37b1f9b0981fa719e9658dbe806c, the named successor of FRAME-04. One level ran Claude Opus 5 on claude-code-cli version 2.1.233 and the other on version 2.1.281, and the coder ran on version 2.1.281. Each version sent its own context with the study conversation: version 2.1.281 an identity line, an environment block, a model statement, a budget message and a date reminder, and version 2.1.233 a date reminder, the same identity line and the same budget message.
Requests on the claude-code-cli route on executable version 2.1.281 are judged on effort: each must carry no output_config.effort when its level sends none, and exactly the level's effort when it states one. Requests on the claude-code-cli route on executable version 2.1.233 are not judged on effort: the preflight and the run do not check the effort they carry, and that version sends output_config.effort high when its level sends none. The registration states a low effort for every level, so the rule for a level that sends no effort does not arise; the run did not check the effort that the 2.1.233 requests carried.
The preregistration with content hash sha256:53e3cde457e52f28e4cfe2dbac27c1a04f5319937a9961744d1711a2f3c2b223 was finalized at 2026-10-06T15:02:22.230Z. The registration was amended once: a preregistration-finalization amendment at 2026-10-06T15:02:26.930Z, not informed by outcomes. No result was viewed before the preregistration was finalized. 160 units were scheduled and 160 reserve units were held in reserve. 160 units were analysed, 0 replaced by a reserve unit, 0 excluded and 160 reserve units left unused.
Against FRAME-03 and FRAME-04
FRAME-03 ran Claude Opus 5 on Claude Code 2.1.233 and FRAME-04 ran it on Claude Code 2.1.281, on different days, each beside a different second model. FRAME-06 ran both versions in one run on the same day, with the FRAME-04 scenario, wording, scales, codes, checks, sizes and rules.
The confidence effect
On Claude Code 2.1.281, Claude Opus 5's confidence fell under the loss frame in FRAME-04 (5.1 against 4.4, H2b confirmed) and again in FRAME-06 (5.0 against 4.1, H2a confirmed). On Claude Code 2.1.233, it did not move significantly in FRAME-03 (4.7 against 4.6, H2b falsified) and did not move significantly in FRAME-06 (4.7 against 4.4, H2b falsified), though in FRAME-06 the difference was in the predicted direction. On each version FRAME-06 recorded the same verdict as the earlier study on that version.
The likelihood and the framing recognised code
The likelihood effect held on both versions, as it did in FRAME-03 and FRAME-04. Claude Opus 5 recognised the framing more often under the loss frame on both versions: 27 of 40 against 11 of 40 on 2.1.281 and 26 of 40 against 11 of 40 on 2.1.233. In FRAME-04 it did so in 28 of 40 against 8 of 39, and in FRAME-03 in 35 of 39 against 15 of 37, where the coder was Claude Opus 5 rather than Claude Sonnet 5.
No test was run between studies; they ran on different days with different second models, and every number in this comparison is one that the three recorded views state.
What this says and does not say
Within this scenario, this wording and these two routes, Claude Opus 5 rates Plan A less likely to be recommended under the loss frame on both versions of the command line, and its confidence falls under the loss frame on 2.1.281 and falls less, not significantly, on 2.1.233. The study did not find a significant direct difference in gain-frame confidence between the versions, so it does not show that the version causes the confidence effect. The two versions also differ in more than their number: each sends its own context to the model, and only 2.1.281 is checked for the effort it sends. The result is about this model on 2026-10-06 on these two pinned command lines, with the registered task instructions. It says nothing about other scenarios, other providers or human participants.
The rationale coder was Claude Sonnet 5 on its own pinned route on Claude Code 2.1.281, a single model coder that is not a FRAME-06 participant, blind to condition and level. It saw the codebook and one reply at a time, and never a condition, a model name, another unit or an outcome. Coding by a single coder is a declared limitation, and every reply and every coder reply is archived for a later recode. On Claude Code 2.1.281, the version both the coder and one of the two Claude Opus 5 levels ran on, the route sent its environment, model, budget and date context bare to Claude Opus 5 and wrapped in system-reminder tags to the Claude Sonnet 5 coder, and the study conversation and the instruction were identical. On Claude Code 2.1.233 the route added only a date reminder, the identity line and the budget message. Hidden thinking, which the route does not return, was present in 95 of 240 and 88 of 240 calls on Claude Code 2.1.281 under the gain and loss frames, and in 88 of 240 and 85 of 240 on Claude Code 2.1.233, and is not part of the analysis.