OpenPsy

Tools for Effective AI Behavioral Research.

OpenPsy is an open research platform for designing, running and analyzing behavioral experiments, including experiments in which language models take part as participants. The goal of OpenPsy is to enable researchers to understand what is happening beneath their results and to have an audit history in order to understand their research programs. OpenPsy is presently in an internal closed alpha.

Why OpenPsy?

Agentic research can be difficult to parse, monitor, and understand. Often agents will violate rules and assumptions about research design undermining results. OpenPsy was born from this difficulty and aims to assist researchers in agent to agent and agent to human research. The software is designed to improve a researcher's ability to plan, execute, and analyze agentic experiments with or without human participants.

  • Introspection

    OpenPsy enables the effective design of experiments by providing simulation prior to execution, and saved session introspection for auditability.

  • Replicability

    OpenPsy saves every version of a design and changes, so a researcher can audit and export the design for replication.

  • Context Control

    OpenPsy enables researchers to understand exactly what content they are serving agents at each step of an experiment. This experimental introspection allows researchers to improve their confidence in their operations and results.

OpenPsy Builds Experiments as a Flow

In the flow editor, researchers use "blocks" to build ideal experimental process through a node graph interface. All OpenPsy aspects read from the same common database ensuring there is no drift between a design and execution. Below we illustrate an example study in the OpenPsy framework.

Example Study

1Design

Research question

Does a gain or loss framing in a question affect how an LLM responds? Does this response vary across models?

Two factors, four conditions

The sample experiment is a 2x2, two factors with two levels. The first factor is a gain versus loss frame. The second factor is model type, Claude Opus 5 or Claude Sonnet 5.

Participants are randomized across the four conditions.

Four conditions
Claude Opus 5Claude Sonnet 5
Gain frameGain + OpusGain + Sonnet
Loss frameLoss + OpusLoss + Sonnet
Factor 1 · Framing: Gain / Loss

The two scenarios differ only in the highlighted clause.

Gain frame

A new infection is expected to affect 900 people in a region. Health officials are deciding whether to adopt Plan A. If Plan A is adopted, 300 of those people will be protected.

Loss frame

A new infection is expected to affect 900 people in a region. Health officials are deciding whether to adopt Plan A. If Plan A is adopted, 600 of those people will be left unprotected.

OpenPsy ensures the conditions are consistent aside from the manipulation using evaluations prior to and after execution.

Factor 2 · Model type: Opus / Sonnet

The design uses Claude Opus 5 and Claude Sonnet 5 in both framing conditions, with the same procedure and settings for both models.

Procedure
The OpenPsy Flow graph for the example study on a dotted canvas. On the left, a Random assignment block reads Preview one of the 4 conditions, Seed 4107 and 4 routes. Dashed arrows lead from it to two boxes labeled Independent variable, one above the other. The Frame box says 2 levels and Applied at Task instructions, and holds Level 1 Gain frame and Level 2 Loss frame, each with a Prompt badge. The Model box says 2 levels and Applied at Task instructions, adds that the model taking part varies with the same delivery for every model, and holds Level 1 Claude Opus 5 and Level 2 Claude Sonnet 5, each with a Model badge. Dashed lines from the two boxes join into one arrow to Task instructions, which is labeled Instruction and has a Prompt badge. Arrows then lead in order to Recommendation Binary (Yes/No) labeled Behaviour, Open-Ended Rationale Explanation labeled Production, and Confidence Rating (7 Point Scale) labeled Endorsement, each with an Interface badge. Each of these four cards reads 1 item.
Example design. OpenPsy’s flow designer keeps the assigned condition, shared instructions and outcome order visible in one view.
The Add a Block panel in OpenPsy. It has a Close button and a search field labeled Search blocks. Under Functions, Participant instructions is listed at version 1. Under IVs, Gain frame and then Loss frame are listed at version 1. Under DVs, Recommendation Binary (Yes/No), Open-Ended Rationale Explanation and Confidence Rating (7 Point Scale) are listed in that order at version 1. Each group has a short description and a Show blocks from menu, which is set to This program for Functions and IVs and to All programs for DVs. Each block has an Add to Flow button.
Blocks Come From a Shared Library. The researcher can add blocks from their common pool, using their bank of IVs and DVs.
The OpenPsy editor for the selected block, Confidence Rating (7 Point Scale). It has buttons to add a block after this one, move it earlier and remove it, and the Move later button is unavailable. The Block title field reads Confidence Rating (7 Point Scale), and Delivered to is set to The participant. The wording box, three lines tall, reads: How confident are you in your recommendation, from 1 (not at all confident) to 7 (completely confident)? A note below says the block is pinned to Confidence Rating (7 Point Scale), version 2, that changing the wording creates the next version, and that the saved version is kept.
Versioning of Blocks Tracks Provenance. Researchers can modify and track changes across measures creating a history of experiments to understand their experimental progress.
The header of the experiment Gain and loss framing in a treatment choice. It has Create revision, Prepare replication and Export design buttons and a Version menu set to Version 4. Below it, the Source and attribution section is closed and the Changes from version 3 section is open. The open section shows a shaded panel with a green rule on its left edge. The panel says that Confidence Rating (7 Point Scale) moved from version 1 to version 2 of its wording.
Edit History Provides Auditability. In this example, the researcher has changed the wording of the confidence rating in version 4, which is now forked from the prior version 3.

2Execution

The same four cells now track session completion. This simulated monitor shows counts only; responses stay hidden while collection is open.

The Coverage & integrity panel of the Execution Monitor. Under the heading Pilot, a line reads "This phase is recorded in simulated mode." Below it is a green progress bar labelled 23/40 terminal. A section headed Cell coverage (counts only) holds a card titled Session Coverage by Frame and Model. The table has one diagonal corner header: Model is in the upper-right triangle for the columns, and Frame is in the lower-left triangle for the rows. Claude Opus 5 has a purple swatch and Claude Sonnet 5 has a blue swatch. It has four cells, each with a green progress strip: Gain frame with Claude Opus 5 shows 6 of 10 sessions complete, Gain frame with Claude Sonnet 5 shows 6 of 10, Loss frame with Claude Opus 5 shows 6 of 10, and Loss frame with Claude Sonnet 5 shows 5 of 10. There is no Total column and no totals row. A note under the table reads "Each cell counts the scheduled sessions of one condition that have finished. The table shows how many sessions are done, never how they turned out." Below the card, the terminal taxonomy shows valid-observation 23. Three tiles at the bottom show ledger freshness (evidence arriving now), envelope closure (23/23 terminal units sealed by envelope) and scoring closure (23/40 scheduled units scored).
Blinded and Secure Execution. The Execution Monitor shows how many sessions in each condition have finished and keeps results hidden until collection closes.

3Results

In the hypothetical results below, the same four cells show outcomes, with row and column totals and a bar chart.

The Hypothetical Results section of the example study. A dashed notice reads: These results are hypothetical. They were made up to show how a two-by-two result is reported, and they did not come from running any model. Below it, a table is titled Sessions Recommending Plan A, by Frame and Model (Hypothetical). Its single header row starts with a diagonal corner: Model is in the upper-right triangle for the columns, and Frame is in the lower-left triangle for the rows. The remaining headers are Claude Opus 5 with a purple swatch, Claude Sonnet 5 with a blue swatch and Total. The Total header and the Total row label are set in the same bold type as the Claude Opus 5, Claude Sonnet 5, Gain frame and Loss frame headers. The Gain frame row reads 72% (36 of 50, mean confidence 5.4) for Claude Opus 5, 74% (37 of 50, mean confidence 5.5) for Claude Sonnet 5 and 73% (73 of 100, mean confidence 5.5) in the Total column. The Loss frame row reads 46% (23 of 50, mean confidence 5.0) for Claude Opus 5, 46% (23 of 50, mean confidence 5.1) for Claude Sonnet 5 and 46% (46 of 100, mean confidence 5.1) in the Total column. The Total row reads 59% (59 of 100, mean confidence 5.2) for Claude Opus 5 and 60% (60 of 100, mean confidence 5.3) for Claude Sonnet 5, and its last cell is empty. The Total column and the Total row have a light green fill and a heavier rule on their inner edge. A note under the table says that comparing the row totals shows the main effect of Frame, comparing the column totals shows the main effect of Model, and comparing the four cells shows whether the factors interact. Beneath the table, a grouped bar chart is titled Percentage of Sessions Recommending Plan A, by Frame and Model (Hypothetical), with a legend for Claude Opus 5 in purple and Claude Sonnet 5 in blue and a vertical axis from 0% to 100%. It shows a Claude Opus 5 bar and a Claude Sonnet 5 bar for each frame, labeled 72% and 74% for Gain frame and 46% and 46% for Loss frame. At the top of the chart, a dark bracket marked with three stars spans the Gain frame and Loss frame groups and ends in two short equal ticks, one centred over each group, that stop well above the other brackets. Below it, a blue bracket marked with one star has long legs that drop to the 74% label of the Gain frame Claude Sonnet 5 bar and the 46% label of the Loss frame Claude Sonnet 5 bar, and lower still a purple bracket marked with one star has legs that drop to the 72% label of the Gain frame Claude Opus 5 bar and the 46% label of the Loss frame Claude Opus 5 bar. The only text under the chart is the star key, which reads *p < .05, **p < .01 and ***p < .001.
Standardized Results Reporting. OpenPsy uses standardized formats including APA style tables and charts.

OpenPsy Roadmap

OpenPsy is in internal closed alpha. It has been tested with invented data only, and no real participant has taken part.

In Progress (Expected Live Q4 2026)

  • A researcher can create an experiment, build it as a graph from Library blocks, edit each block's wording, and save every change as a new version.
  • Any two saved versions can be compared, and the differences are written out as sentences.
  • Library wordings are stored as exact versions, and publishing a revision leaves existing experiments on the version they used.
  • A scripted test participant can take an authored study from consent to completion on a phone-sized screen.
  • Automated evals pre and post execution

Near Future

  • Export and replication of experiments by third parties
  • Automated Measure and Execution Pilot Workflows
  • Automated Power Analysis
  • Co-Pilot Assistant in application.
  • Hybrid Human/Agent Experiments

Source Code

The source code, the specification and the product documentation are being prepared for publication on GitHub.

GitHub repository opening soon

Scroll to read. Pinch to zoom.