OpenPsy

Effective AI Behavioral Research

OpenPsy is an open research platform for designing, running and analyzing behavioral experiments, including experiments in which language models take part as participants. The goal of OpenPsy is to enable researchers to understand what is happening beneath their results and to have an audit history in order to understand their research programs. OpenPsy is presently in an internal closed alpha.

Why OpenPsy?

Agentic research can be difficult to parse, monitor, and understand. Often agents will violate rules and assumptions about research design undermining results. OpenPsy was born from this difficulty and aims to assist researchers in agent to agent and agent to human research. The software is designed to improve a researcher's ability to plan, execute, and analyze agentic experiments with or without human participants.

  • Introspection

    OpenPsy enables the effective design of experiments by providing simulation prior to execution, and saved session logs for auditability.

  • Replicability

    OpenPsy saves every version of a design and changes, so a researcher can audit and export the design for replication.

  • Control

    OpenPsy enables researchers to control context, understanding exactly what content they are serving agents at each step of an experiment.

Build Experiments as a Flow

In the flow editor, researchers use "blocks" to build ideal experimental process through a node graph interface. All OpenPsy aspects read from the same common database ensuring there is no drift between a design and execution. Below we illustrate an example study in the OpenPsy framework.

Example

1Plan

Does a gain or loss framing change an LLM's recommendation? Does this response vary across models?
A Classic 2x2 Research Design

This 2x2 experiment was run through OpenPsy, with two factors of two levels each. The first factor is a gain versus loss frame. The second factor is model type, Claude Opus 5 or Claude Sonnet 5.

Participants are randomized across the four conditions.

Four conditions
Claude Opus 5Claude Sonnet 5
Gain frameGain + OpusGain + Sonnet
Loss frameLoss + OpusLoss + Sonnet
Factor 1 · Framing: Gain / Loss

The two scenarios differ only in the highlighted clause.

Gain frame

A new infection is expected to affect 900 people in a region. Health officials are deciding whether to adopt Plan A. If Plan A is adopted, 300 of those people will be protected.

Loss frame

A new infection is expected to affect 900 people in a region. Health officials are deciding whether to adopt Plan A. If Plan A is adopted, 600 of those people will be left unprotected.

Factor 2 · Model type: Opus / Sonnet

Only the model taking part differs. Both levels run the same procedure through the same route with the same settings.

Opus

Claude Opus 5 claude-opus-5

Sonnet

Claude Sonnet 5 claude-sonnet-5

Route
Claude Code CLI in print mode, one call at a time.
Effort
Model default.
Limits
1,024 tokens per reply, 4 minutes per call, no retries.
Session
Four calls: the scenario, then Recommendation, Rationale and Confidence.
Dependent Variable
Question

Do you recommend adopting Plan A?

Response

Yes / No

Our dependent variable of interest is the percentage of valid responses selecting Yes to recommend Plan A in each condition.

Procedure
The OpenPsy Flow of the FRAME-01 study with its simulation paused on the fourth of eight steps. The toolbar says Version 5 is saved and holds a Resume button, an Exit simulation button, a speed slider set to 2x, the step count 4 of 8 and zoom controls reading 53 percent, with the whole procedure fitted in view. Under the toolbar, a strip headed Preview condition reads Frame Gain frame and Model Claude Opus 5, followed by the line: One assignment. No participant calls are made and no study data is created. In the graph, a Random assignment block reads Preview one of the 4 conditions, Seed 4107 and 2 routes, and is marked as taken. It leads to the independent variable Model, applied at Task instructions, whose Level 1 Claude Opus 5 reads Applied in this preview and whose Level 2 Claude Sonnet 5 is dimmed and reads Not assigned in this preview; then to Task instructions, an instruction block reading Call 1 of 4 begins here, marked as taken; then to the independent variable Frame, applied at Scenario, whose Level 1 Gain frame is outlined and reads Applying in this preview and whose Level 2 Loss frame is dimmed; then to Scenario, a manipulation block reading Call 1 of 4 continues; then to Recommendation Binary (Yes/No), a behaviour block with a single choice of two options reading Call 2 of 4 begins here; then to Open-Ended Rationale Explanation, a production block reading Call 3 of 4 begins here; then to Confidence Rating (7 Point Scale), an endorsement block with an open text response reading Call 4 of 4 begins here. Each block carries a Prompt badge and reads 1 item, Design only. The pane across the bottom is headed Simulation step 4 of 8 with the block name Gain frame and the lines Direct agent prompt and 1 item. Beside it are two buttons, Agent view, which is selected, and Library source, then Condition chips reading 1 Gain frame, selected, and 2 Loss frame, then Dry-run route followed by: Design only. Not admitted for execution. The body of the pane opens with Gain frame, followed by its registered wording in a monospaced face: If Plan A is adopted, 300 of those people will be protected. Below it, a note in smaller, lighter type reads: Specification preview. No provider is called and no participant data exists here.
Flow designer enables full simulation prior to execution. By designing experiments using a node graph interface, researchers can simulate their experiment to see what their agents are receiving. The above is the actual flow for the example experiment.

2Execute

The Execution Monitor of the FRAME-01 run, replayed as it stood at 2026-09-25 01:00:26 UTC. Under the heading Execution Monitor, a banner reads "Replay as of 2026-09-25 01:00:26 UTC. This is the state the ledger recorded at that moment, not live evidence." Two tabs follow, Pilot and Main study; the selected Main study tab is filled dark green with white text, and the Pilot tab is outlined. Under the heading Main study (confirmatory), a line reads "This phase is recorded in live mode." A green bar filled to three fifths is labelled 120/200 terminal, and two chips read Wave 0 100/100 and Wave 1 20/100. A line reads "1 main study unit is in flight." A section headed Cell coverage (counts only) explains that each cell's strip fills as its scheduled units reach a terminal state, and that coverage is how many units are done, never how they turned out. Below it, a card titled Main Study Session Coverage by Frame and Model holds a two-by-two table with a diagonal corner header, where Model labels the columns and Frame labels the rows. The rows are Gain frame and then Loss frame, and the columns are Claude Opus 5 with a purple swatch and then Claude Sonnet 5 with a blue swatch. All four cells read 30 of 50 sessions complete and 60%, each with a green strip filled to three fifths. A note under the table reads "Each cell counts the scheduled main study sessions of one condition that have finished. The table shows how many sessions are done, never how they turned out." The terminal taxonomy, across the main study's settled units (terminal, stopped, failed, and quarantined), has one bucket, complete 120, on a basis of 120 settled units, each landing in exactly one bucket. Cards at the bottom show envelope closure (120/120 terminal units sealed by envelope), scoring closure (120/200 scheduled units scored) and ledger freshness (evidence is arriving now, and the last evidence is from 2026-09-25 01:00:26 UTC, 0 seconds before the replay's clock).
Blinded and Secure Execution. The execution monitor shows results in progress providing a current status update.

3Analyze

FRAME-01: Does a gain or loss frame affect what a model endorses?

Yes

Recorded results in a two-by-two table and grouped bar chart. Frame labels the rows and Model labels the columns, separated by a diagonal corner. Gain frame: Claude Opus 5, 100% (50 of 50); Claude Sonnet 5, 98% (49 of 50); row total, 99% (99 of 100). Loss frame: Claude Opus 5, 100% (50 of 50); Claude Sonnet 5, 34% (17 of 50); row total, 67% (67 of 100). Column totals: Claude Opus 5, 100% (100 of 100); Claude Sonnet 5, 66% (66 of 100). The grand-total corner is blank. The table also shows a mean confidence rating for each condition and total. The four bars carry Wilson 95% confidence intervals: Gain frame with Claude Opus 5, 92.9 to 100.0%; Gain frame with Claude Sonnet 5, 89.5 to 99.6%; Loss frame with Claude Opus 5, 92.9 to 100.0%; Loss frame with Claude Sonnet 5, 22.4 to 47.8%. Two brackets carry three stars each, for a Holm-adjusted p less than .001: a dark bracket compares Claude Opus 5 with Claude Sonnet 5 in the loss frame, and a blue bracket compares the gain frame with the loss frame for Claude Sonnet 5. No bracket is drawn for Claude Opus 5 across frames or for the two models in the gain frame. The key defines *p less than .05, **p less than .01 and ***p less than .001.
Standardized Results Reporting. OpenPsy uses standardized formats including APA style tables and charts. These are the recorded results of the FRAME-01 study, run through OpenPsy on 2026-09-24 with 200 preregistered sessions. In the preregistered fallback analysis, Model had a significant main effect, while Frame had no significant main effect and the interaction was not significant, so OpenPsy recorded the preregistered framing hypothesis as inconclusive. The pre-registered exact test of Gain frame against Loss frame within Claude Sonnet 5 is significant (Holm-adjusted p less than .001).

Note. N = 200 preregistered confirmatory sessions, n = 50 per condition, every one analysed. Percentages are sessions recommending Plan A. Error bars are Wilson 95% confidence intervals. Brackets mark the comparisons the registered tests found significant.

Study AGENT-EE7FF17DBD5EE72B744DF50BABD9707A. Registration hash sha256:0011b440d41d20edfe5a48978853d12aa926cbb3d5edc43de7d0728f6ae7cd69. Result digest sha256:0be67aac147c0ed153830757b38016080ed8d6274d94677a37b90dddff223b22.

4Replicate

A recorded study can be replicated from its own registration. FRAME-02 replicated FRAME-01 through OpenPsy with 320 preregistered sessions in two samples and one hypothesis per model.

FRAME-02: Does the endorsement effect replicate in a larger study?

Yes

Recorded results of FRAME-02 in a two-by-two table and grouped bar chart. Frame labels the rows and Model labels the columns, separated by a diagonal corner. Gain frame: Claude Opus 5, 99% (79 of 80), mean confidence 5.6; Claude Sonnet 5, 100% (79 of 79), mean confidence 4.5. Loss frame: Claude Opus 5, 99% (76 of 77), mean confidence 5.2; Claude Sonnet 5, 14% (11 of 80), mean confidence 2.9. Row totals: Gain frame, 99% (158 of 159); Loss frame, 55% (87 of 157). Column totals: Claude Opus 5, 99% (155 of 157); Claude Sonnet 5, 57% (90 of 159). The grand-total corner is blank. The four bars carry Wilson 95% confidence intervals.
Replication with Per-Model Hypotheses. These are the recorded results of FRAME-02, the replication of FRAME-01 run through OpenPsy on 2026-09-25 and 2026-09-26 with 320 preregistered sessions in two samples of 40 per condition. The preregistration asked one question per model. For Claude Sonnet 5 the registered exact test of Gain frame against Loss frame is significant (Holm-adjusted p less than .001), with 79 of 79 sessions recommending Plan A under the gain frame and 11 of 80 under the loss frame, and OpenPsy recorded that hypothesis as confirmed. For Claude Opus 5 the same test finds no difference (79 of 80 against 76 of 77, Holm-adjusted p = 1.00), and under the registered rule that a result not significant in the predicted direction counts against the hypothesis, OpenPsy recorded it as falsified. Both samples show the same pattern. Five of the 320 sessions ended in an ambiguous provider call, were never resent, and are counted as missing under the registered rule; every registered sensitivity bound gives the same verdicts.

Note. N = 320 preregistered confirmatory sessions in two samples, n = 80 per condition (40 per sample). Five sessions that ended in an ambiguous provider call were excluded as missing under the registered rule, so the analysed n is 80, 79, 77 and 80 for Gain frame with Claude Opus 5, Gain frame with Claude Sonnet 5, Loss frame with Claude Opus 5 and Loss frame with Claude Sonnet 5. Percentages are sessions recommending Plan A. Error bars are Wilson 95% confidence intervals. Brackets mark the comparisons the registered tests found significant.

Study AGENT-B9FB8AA2D19BF8FF15BAD36154CCC6E6, registered under the title FRAME-01 as version 5 of the same experiment. Registration hash sha256:f624c956ecbb04e2188b775389c976e5f6c2c53611f9fe2c2dbda82bdec363de. Result digest sha256:2cf4606819ed0233e4169b555ea69cce5410459762c9b08c0ac48c76ac10a0a0.

See the replication here. FRAME-02 has its own page with each sample, the effect sizes and the record of the run.

5Research

After obtaining results, progress your research programs to better understand phenomena through subsequent experimentation.

FRAME-03: Does the frame change how strongly a model endorses, not just whether it does?

Yes

Recorded results of FRAME-03 in a two-by-two table and grouped bar chart of the mean likelihood of recommending Plan A on the 1 to 7 scale. Frame labels the rows and Model labels the columns, separated by a diagonal corner. Gain frame: Claude Opus 5, mean 5.8, SD 0.52, n = 37, mean confidence 4.7; Claude Sonnet 5, mean 5.1, SD 0.42, n = 40, mean confidence 3.6. Loss frame: Claude Opus 5, mean 5.1, SD 0.68, n = 39, mean confidence 4.6; Claude Sonnet 5, mean 3.4, SD 0.74, n = 40, mean confidence 2.5. Row totals: Gain frame, mean 5.4; Loss frame, mean 4.2. Column totals: Claude Opus 5, mean 5.4; Claude Sonnet 5, mean 4.2. The grand-total corner is blank. The four bars carry the recorded 95% confidence intervals of the means, and three-star brackets mark the two models within each frame, each model across the frames and the main effect of Frame.
Recorded results of FRAME-03 in a two-by-two table and line chart of the mean likelihood of recommending Plan A on the 1 to 7 scale. Frame labels the rows and Model labels the columns, separated by a diagonal corner. Gain frame: Claude Opus 5, mean 5.8, SD 0.52, n = 37, mean confidence 4.7; Claude Sonnet 5, mean 5.1, SD 0.42, n = 40, mean confidence 3.6. Loss frame: Claude Opus 5, mean 5.1, SD 0.68, n = 39, mean confidence 4.6; Claude Sonnet 5, mean 3.4, SD 0.74, n = 40, mean confidence 2.5. Row totals: Gain frame, mean 5.4; Loss frame, mean 4.2. Column totals: Claude Opus 5, mean 5.4; Claude Sonnet 5, mean 4.2. The grand-total corner is blank. In the line chart Claude Opus 5 falls from 5.8 to 5.1 and Claude Sonnet 5 from 5.1 to 3.4, with the recorded 95% confidence intervals centred on each dot.
A Seven-Point Follow-Up. These are the recorded results of FRAME-03, run through OpenPsy on 2026-09-28 with 160 preregistered sessions in two samples of 20 per condition. It asked the FRAME-02 question as a likelihood from 1 to 7 instead of yes or no, tested the confidence rating and coded each rationale under five preregistered codes. Both models rated Plan A less likely under the loss frame: Claude Sonnet 5 from 5.1 to 3.4 and Claude Opus 5 from 5.8 to 5.1, each with a Holm-adjusted p less than .001, so the corresponding Claude Opus 5 hypothesis, H1b, which the yes or no outcome falsified in FRAME-02, was confirmed on the scale. Nine of the ten hypotheses were confirmed. One was falsified: Claude Opus 5 did not rate its confidence higher under the gain frame (4.7 against 4.6, Holm-adjusted p = .67).

Note. N = 160 preregistered confirmatory sessions in two samples, n = 40 per condition (20 per sample). Four sessions that ended in an ambiguous provider call were excluded as missing under the registered rule, so the analysed n is 37, 40, 39 and 40 for Gain frame with Claude Opus 5, Gain frame with Claude Sonnet 5, Loss frame with Claude Opus 5 and Loss frame with Claude Sonnet 5. Each cell gives the mean likelihood of recommending Plan A on the 1 to 7 scale with its SD and n, and the mean confidence rating. Error bars are the recorded 95% confidence intervals of the means. Brackets mark the comparisons the registered tests found significant.

Study AGENT-4077193D973E70CDF2A39DCC053D80C5. Registration hash sha256:b2ce589e323d8aa833f3176065b5f1bdae33cb59eae3050644d58627aa926b8a. Result digest sha256:d4a4431898f3834d01808cfc4c8801e432b30628f92df89841c7d65182ae2336.

See FRAME-03 here. FRAME-03 has its own page with every outcome, each sample and the record of the run.

FRAME-04: Do the endorsement effects extend to the newest Sonnet model (5.5)?

Yes, but with some differences.

The next step of the same programme, FRAME-04, ran the FRAME-03 design again through OpenPsy with Claude Sonnet 5.5 in place of Claude Sonnet 5.

Recorded results of FRAME-04 in a two-by-two table and grouped bar chart of the mean likelihood of recommending Plan A on the 1 to 7 scale. Frame labels the rows and Model labels the columns, separated by a diagonal corner. Gain frame: Claude Opus 5, mean 5.9, SD 0.22, n = 39, mean confidence 5.1; Claude Sonnet 5.5, mean 5.0, SD 0.00, n = 40, mean confidence 3.7. Loss frame: Claude Opus 5, mean 5.3, SD 0.60, n = 40, mean confidence 4.4; Claude Sonnet 5.5, mean 3.9, SD 0.44, n = 40, mean confidence 4.1. Row totals: Gain frame, mean 5.5; Loss frame, mean 4.6. Column totals: Claude Opus 5, mean 5.6; Claude Sonnet 5.5, mean 4.5. The grand-total corner is blank. The four bars carry the recorded 95% confidence intervals of the means; every Claude Sonnet 5.5 reply under the gain frame was 5, so that interval is a single point. Three-star brackets mark the two models within each frame, each model across the frames and the main effect of Frame.
Recorded results of FRAME-04 in a two-by-two table and line chart of the mean likelihood of recommending Plan A on the 1 to 7 scale. Frame labels the rows and Model labels the columns, separated by a diagonal corner. Gain frame: Claude Opus 5, mean 5.9, SD 0.22, n = 39, mean confidence 5.1; Claude Sonnet 5.5, mean 5.0, SD 0.00, n = 40, mean confidence 3.7. Loss frame: Claude Opus 5, mean 5.3, SD 0.60, n = 40, mean confidence 4.4; Claude Sonnet 5.5, mean 3.9, SD 0.44, n = 40, mean confidence 4.1. Row totals: Gain frame, mean 5.5; Loss frame, mean 4.6. Column totals: Claude Opus 5, mean 5.6; Claude Sonnet 5.5, mean 4.5. The grand-total corner is blank. In the line chart Claude Opus 5 falls from 5.9 to 5.3 and Claude Sonnet 5.5 from 5.0 to 3.9, with the recorded 95% confidence intervals centred on each dot.
A Newer Model, the Same Design. These are the recorded results of FRAME-04, run through OpenPsy on 2026-09-29 with 160 preregistered sessions in two samples of 20 per condition. Both models rated Plan A less likely under the loss frame: Claude Sonnet 5.5 from 5.0 to 3.9 and Claude Opus 5 from 5.9 to 5.3, each with a Holm-adjusted p less than .001. Five of the ten hypotheses were confirmed and five were falsified, all five about Claude Sonnet 5.5: its confidence rose under the loss frame instead of falling (3.7 against 4.1, Holm-adjusted p = .034), Claude Opus 5 was not significantly more confident than it under the loss frame (4.4 against 4.1, Holm-adjusted p = .25), and the three predictions about its rationale codes failed.

Note. N = 160 preregistered confirmatory sessions in two samples, n = 40 per condition (20 per sample). One session that ended in an ambiguous provider call was excluded as missing under the registered rule, so the analysed n is 39, 40, 40 and 40 for Gain frame with Claude Opus 5, Gain frame with Claude Sonnet 5.5, Loss frame with Claude Opus 5 and Loss frame with Claude Sonnet 5.5. Each cell gives the mean likelihood of recommending Plan A on the 1 to 7 scale with its SD and n, and the mean confidence rating. Error bars are the recorded 95% confidence intervals of the means. Brackets mark the comparisons the registered tests found significant.

Study AGENT-C218672F0B8877E662FF844F30FFC381. Registration hash sha256:6919c277f77e926c8e9719c7e6ba8ce148ada5cf08fdf370c78f12c48add26cb. Result digest sha256:557362dfff53858ebb0af0f072f0e29d4622afdcf39612ff7610c1b972577e95. Software release c761c27d.

See FRAME-04 here. FRAME-04 has its own page with every outcome, each sample and the record of the run.

FRAME-05: Is the same effect present in ChatGPT's Terra and Sol?

Yes, in the available exploratory results.

FRAME-05 exploratory results in a two-by-two table and grouped bar chart of likelihood of recommending Plan A on the 1 to 7 scale. Frame labels the rows and Model labels the columns. Gain frame: GPT-5.6 Sol, mean 5.75, SD 0.49, n 40, mean confidence 4.88; GPT-5.6 Terra, mean 6.28, SD 0.60, n 40, mean confidence 5.60. Loss frame: Sol, mean 3.46, SD 0.72, n 39, mean confidence 4.46; Terra, mean 4.85, SD 0.81, n 39, mean confidence 3.87. Error bars show newly calculated individual unadjusted 95% Student-t confidence intervals of the recommendation means. Three-star brackets mark the four Holm-adjusted pairwise comparisons and the unadjusted main effect of Frame. These are post-response exploratory results.
FRAME-05 exploratory results in a two-by-two table and line chart of likelihood of recommending Plan A on the 1 to 7 scale. Frame labels the rows and Model labels the columns. Gain frame: GPT-5.6 Sol, mean 5.75, SD 0.49, n 40, mean confidence 4.88; GPT-5.6 Terra, mean 6.28, SD 0.60, n 40, mean confidence 5.60. Loss frame: Sol, mean 3.46, SD 0.72, n 39, mean confidence 4.46; Terra, mean 4.85, SD 0.81, n 39, mean confidence 3.87. Error bars show newly calculated individual unadjusted 95% Student-t confidence intervals of the recommendation means. Sol falls from 5.75 to 3.46 and Terra from 6.28 to 4.85. The intervals are centred on each point.
The Same Design With OpenAI Models. Across 158 of 160 planned sessions, both GPT-5.6 models rated Plan A higher under the gain frame: Sol 5.75 versus 3.46, and Terra 6.28 versus 4.85, on the 1–7 recommendation scale. Each within-model comparison has a Holm-adjusted p < .001. These are post-response exploratory results; because recovery and analysis followed disclosure of earlier responses, the original confirmatory run is incomplete.

Note. N = 160 planned sessions in two samples, with 40 per condition (20 per sample). Both Gain cells contain 40 analysed sessions and both Loss cells contain 39. Two sessions remain missing. Each cell gives the mean likelihood of recommending Plan A with its SD and n, and the mean confidence rating. Error bars are newly calculated individual unadjusted 95% Student-t confidence intervals of the means, using n − 1 degrees of freedom as in FRAME-04. Pairwise significance brackets use Holm-adjusted p values; the main effect of Frame uses its unadjusted omnibus p. The intervals do not account for missingness, outcome-informed recovery or provider dependence.

FRAME-05 verified export sha256:1a3430c1bd216c30bd13c6374a1d07887f21b691b61d1619e6f6e84ad6071c97. Exploratory analysis sha256:0e422625c52a7c8a2cd92220b704ddca44f19efdb9906329f6b3c41466358c44. Mean confidence intervals calculated separately for this presentation.

FRAME-05: the OpenAI replication. FRAME-05 has its own page with every outcome, each sample, statistical comparisons, independent double coding and the record of the run.

FRAME-06: Does Opus 5's confidence effect depend on the Claude Code version?

Not reliably. It appears on 2.1.281 and is weaker on 2.1.233.

Recorded results of FRAME-06 in a two-by-two table and grouped bar chart of the mean likelihood of recommending Plan A on the 1 to 7 scale. Frame labels the rows and Model labels the columns, separated by a diagonal corner. Gain frame: Claude Opus 5 on Claude Code 2.1.281, mean 6.0, SD 0.22, n = 40, mean confidence 5.0; Claude Opus 5 on Claude Code 2.1.233, mean 5.8, SD 0.42, n = 40, mean confidence 4.7. Loss frame: Claude Opus 5 on Claude Code 2.1.281, mean 5.2, SD 0.36, n = 40, mean confidence 4.1; Claude Opus 5 on Claude Code 2.1.233, mean 5.3, SD 0.59, n = 40, mean confidence 4.4. Row totals: Gain frame, mean 5.9, n = 80; Loss frame, mean 5.2, n = 80. Column totals: Claude Opus 5 on Claude Code 2.1.281, mean 5.6, n = 80; Claude Opus 5 on Claude Code 2.1.233, mean 5.5, n = 80. The grand-total corner is blank. The four bars carry the recorded 95% confidence intervals of the means. Brackets mark Claude Opus 5 on Claude Code 2.1.281 in Gain frame against Claude Opus 5 on Claude Code 2.1.281 in Loss frame, with Holm-adjusted p less than .001 (three stars); Claude Opus 5 on Claude Code 2.1.233 in Gain frame against Claude Opus 5 on Claude Code 2.1.233 in Loss frame, with Holm-adjusted p less than .001 (three stars); the main effect of Frame, Gain frame against Loss frame, with p less than .001 (three stars).
Recorded results of FRAME-06 in a two-by-two table and line chart of the mean likelihood of recommending Plan A on the 1 to 7 scale. Frame labels the rows and Model labels the columns, separated by a diagonal corner. Gain frame: Claude Opus 5 on Claude Code 2.1.281, mean 6.0, SD 0.22, n = 40, mean confidence 5.0; Claude Opus 5 on Claude Code 2.1.233, mean 5.8, SD 0.42, n = 40, mean confidence 4.7. Loss frame: Claude Opus 5 on Claude Code 2.1.281, mean 5.2, SD 0.36, n = 40, mean confidence 4.1; Claude Opus 5 on Claude Code 2.1.233, mean 5.3, SD 0.59, n = 40, mean confidence 4.4. Row totals: Gain frame, mean 5.9, n = 80; Loss frame, mean 5.2, n = 80. Column totals: Claude Opus 5 on Claude Code 2.1.281, mean 5.6, n = 80; Claude Opus 5 on Claude Code 2.1.233, mean 5.5, n = 80. The grand-total corner is blank. In the line chart Claude Opus 5 on Claude Code 2.1.281 falls from 6.0 to 5.2 and Claude Opus 5 on Claude Code 2.1.233 from 5.8 to 5.3, with the recorded 95% confidence intervals centred on each dot.
The Same Model on Two Versions. These are the recorded results of FRAME-06, run through OpenPsy on 2026-10-06 with 160 preregistered sessions in two samples of 20 per condition. It ran Claude Opus 5 on Claude Code 2.1.281, the version FRAME-04 used, and on Claude Code 2.1.233, the version FRAME-03 used. On both versions the model rated Plan A less likely under the loss frame, each with a Holm-adjusted p less than .001. Its confidence fell under the loss frame on 2.1.281 (5.0 against 4.1, Holm-adjusted p less than .001) and fell less, not significantly, on 2.1.233 (4.7 against 4.4, Holm-adjusted p = .093), and under the gain frame the two versions did not differ significantly in confidence (Holm-adjusted p = .21). Five of the seven hypotheses were confirmed; H2b and H3 were falsified.

Note. N = 160 preregistered confirmatory sessions in two samples, n = 40 per condition (20 per sample), and every session was analysed. Each cell gives the mean likelihood of recommending Plan A on the 1 to 7 scale with its SD and n, and the mean confidence rating. Error bars are the recorded 95% confidence intervals of the means. Brackets mark the comparisons the registered tests found significant.

Study AGENT-552121A995F2670603D9B626FADAF1BB. Registration hash sha256:e412722ca3078cc058354151ab8a9677888e37b1f9b0981fa719e9658dbe806c. Result digest sha256:6363aa6402da40c79a6aee3f80273f63edeb6a06228d1bf3e9c2c7ea5ec441d8.

See FRAME-06 here. FRAME-06 has its own page with every outcome, each sample and the record of the run.

PRICE-01: Does a gain or loss frame change what a model will pay?

Yes, for both models, and much more for Opus 5.

Recorded results of PRICE-01 in a two-by-two table and grouped bar chart of the mean maximum price in US dollars, on the range from 0 to 60. Frame labels the rows and Model labels the columns, separated by a diagonal corner. Gain frame: Claude Opus 5, mean 25.16 US dollars, SD 9.86, n = 37, mean likelihood to buy 4.7; Claude Sonnet 5.5, mean 15.80 US dollars, SD 5.48, n = 40, mean likelihood to buy 2.7. Loss frame: Claude Opus 5, mean 11.40 US dollars, SD 1.86, n = 40, mean likelihood to buy 2.2; Claude Sonnet 5.5, mean 11.43 US dollars, SD 2.41, n = 40, mean likelihood to buy 2.1. Row totals: Gain frame, mean 20.30 US dollars, SD 9.14, n = 77; Loss frame, mean 11.41 US dollars, SD 2.14, n = 80. Column totals: Claude Opus 5, mean 18.01 US dollars, SD 9.78, n = 77; Claude Sonnet 5.5, mean 13.61 US dollars, SD 4.75, n = 80. The grand-total corner is blank. The four bars carry the recorded 95% confidence intervals of the means, and a dashed line marks the 25 dollar reference. Brackets mark Claude Opus 5 in Gain frame against Claude Sonnet 5.5 in Gain frame, with Holm-adjusted p less than .001 (three stars); Claude Opus 5 in Gain frame against Claude Opus 5 in Loss frame, with Holm-adjusted p less than .001 (three stars); Claude Sonnet 5.5 in Gain frame against Claude Sonnet 5.5 in Loss frame, with Holm-adjusted p less than .001 (three stars); the main effect of Frame, Gain frame against Loss frame, with p less than .001 (three stars).
Recorded results of PRICE-01 in a two-by-two table and line chart of the mean maximum price in US dollars, on the range from 0 to 60. Frame labels the rows and Model labels the columns, separated by a diagonal corner. Gain frame: Claude Opus 5, mean 25.16 US dollars, SD 9.86, n = 37, mean likelihood to buy 4.7; Claude Sonnet 5.5, mean 15.80 US dollars, SD 5.48, n = 40, mean likelihood to buy 2.7. Loss frame: Claude Opus 5, mean 11.40 US dollars, SD 1.86, n = 40, mean likelihood to buy 2.2; Claude Sonnet 5.5, mean 11.43 US dollars, SD 2.41, n = 40, mean likelihood to buy 2.1. Row totals: Gain frame, mean 20.30 US dollars, SD 9.14, n = 77; Loss frame, mean 11.41 US dollars, SD 2.14, n = 80. Column totals: Claude Opus 5, mean 18.01 US dollars, SD 9.78, n = 77; Claude Sonnet 5.5, mean 13.61 US dollars, SD 4.75, n = 80. The grand-total corner is blank. In the line chart Claude Opus 5 falls from 25.16 to 11.40 and Claude Sonnet 5.5 from 15.80 to 11.43, with the recorded 95% confidence intervals centred on each dot and a dashed line at the 25 dollar reference.
The Same Frame, Asked as a Price. These are the recorded results of PRICE-01, run through OpenPsy on 2026-10-06 with 160 preregistered sessions in two samples of 20 per condition. It asked Claude Opus 5 and Claude Sonnet 5.5 the most they would pay, in whole US dollars from 0 to 60, for a preventive course against an infection expected to affect 900 people, described by the 300 of them it protects or the 600 it leaves unprotected, beside a stated reference of 25 US dollars. Both models named a lower price under the loss frame: Claude Opus 5 from 25.16 to 11.40 US dollars and Claude Sonnet 5.5 from 15.80 to 11.43, each with a Holm-adjusted p less than .001, and the registered interaction term confirmed the larger difference for Claude Opus 5 (b = 9.39, SE = 1.84, t(152) = 5.11, p < .001). All five hypotheses were confirmed.

Note. N = 160 preregistered confirmatory sessions in two samples, n = 40 per condition (20 per sample). Six sessions had a provider call that ended ambiguously and was not resent, five in Gain frame with Claude Opus 5 and one in Gain frame with Claude Sonnet 5.5. Three of them, all in Gain frame with Claude Opus 5, have no price, so the analysed n is 37, 40, 40 and 40 for Gain frame with Claude Opus 5, Gain frame with Claude Sonnet 5.5, Loss frame with Claude Opus 5 and Loss frame with Claude Sonnet 5.5. The likelihood to buy is analysed for 156 sessions, 37, 39, 40 and 40 in the same order. Each cell gives the mean maximum price in US dollars with its SD and n, and the mean likelihood to buy. Error bars are the recorded 95% confidence intervals of the means, and the dashed line marks the 25 dollar reference. Brackets mark the comparisons the registered tests found significant.

Study AGENT-533F088A363F01BBCC4F5B9A6C65D539. Registration hash sha256:77908dbd91d1e31193d70fa4ffe8546b904ed016b4c65b1d751e5e8bf8032f0e. Result digest sha256:478427129010254bdc2fbc1f3a1c6e088f789635b5972309574866f994bd6b0e.

See PRICE-01 here. PRICE-01 has its own page with every outcome, each sample and the record of the run.

PRICE-02: Sol and Terra repeat the price-framing task

Public review. Numerical inference available; rationale coding, final recording and app integration pending. The main collection contains 153 eligible complete sessions from 160 scheduled. Both models’ observed price and purchase-likelihood means are higher under gain than loss, meeting all four planned numerical criteria.

Review PRICE-02: results, bar and line charts, tables, and deviations.

PRICE-01 above remains the original Claude Opus 5 / Sonnet 5.5 study. PRICE-02 is a presentation alias for the native Sol/Terra replication; historical records are unchanged.

OpenPsy Roadmap

OpenPsy is in internal closed alpha. The example study on this page was run with language models as its participants, and no human participant has taken part in an OpenPsy study.

In Progress (Expected Live Q4 2026)

  • A researcher can create an experiment, build it as a graph from Library blocks, edit each block's wording, and save every change as a new version.
  • Any two saved versions can be compared, and the differences are written out as sentences.
  • Library wordings are stored as exact versions, and publishing a revision leaves existing experiments on the version they used.
  • A scripted test participant can take an authored study from consent to completion on a phone-sized screen.
  • Automated evals pre and post execution

Near Future

  • Export and replication of experiments by third parties
  • Automated Measure and Execution Pilot Workflows
  • Automated Power Analysis
  • Co-Pilot Assistant in application.
  • Hybrid Human/Agent Experiments

Source Code

The source code, the specification and the product documentation are being prepared for publication on GitHub.

GitHub repository opening soon

Scroll to read. Pinch to zoom.