Scientific Autoresearch
Virtual labs for physical R&D
A materials campaign can move between simulation, synthesis, characterization, and testing. Each result needs to stay connected to the conditions that produced it and the question it was meant to answer, so the team can decide what to investigate next.
Dynamical gives agents an open-source, model-agnostic interface for scientific investigations. A campaign starts with a scientific question; the agent decides what evidence could resolve it, composes a lab from the available capabilities, and changes the plan as results arrive. Dynamical records each action and observation, and the facility decides which physical work can run. Virtual experiments let the agent test options before it asks for scarce physical evidence.
Dynamical
An agent can combine a computational model, a dataset, and an instrument workflow in one virtual lab, then revise that combination as its investigation develops. Each capability describes its inputs, outputs, units, and operating limits so the agent can determine how to use it.
The CLI currently runs simulation and replay and prepares physical experiment requests for facility review. Teams can use the instrument skill to prepare integrations for their own instruments, reusing existing drivers or building on Anthropic’s Model Hardware Standard (MHS) as its drivers become available. Physical execution remains under the facility’s control.
Scientific outputs come from the selected model or dataset. Dynamical checks execution and keeps each result connected to its source, sample history, and instrument settings, with uncertainty where available. Scientists can inspect which results came from calculations and which came from physical measurements.
Supported virtual campaigns can branch from a preserved state to compare different next actions. The record preserves what the agent knew when it chose an experiment and how it used the result, giving researchers a basis for evaluating decisions and training future agents.
Reference capabilities
SDL1 provides an instrument workflow grounded in the source lab's open-source commands and measurement protocol. Its checks cover workflow behavior. FastCat provides a separate empirical predictor for a fixed set of catalyst compositions at one current density, with documented calibration and uncertainty. The experiment-selection study below uses FastCat predictions and archived physical measurements.
Experiment selection
In the water-electrolysis study, agents with Dynamical selected the best available catalyst 42% of the time, compared with 28% without it. The average gap between their selections and the best catalyst was 30% smaller, falling from 38.92 to 27.35 mV.
We tested GPT 5.6 Sol, Claude Opus 5, Qwen 3.8 Max, and DeepSeek V4 Flash 0731 across six independent candidate pools, with three repeats per condition and 144 scored trajectories. Agents with Dynamical could compose a virtual lab and query the FastCat predictor. The comparison condition received the same task information without that predictor. We scored the selected catalysts against archived lab measurements that stayed sealed until scoring.
Agents selected the best catalyst 1.5× as often
Candidate selection
Best catalyst
Bottom half
Mean selected potential
Uncertainty across pools
90% pool-cluster intervalExact values
| Result | Dynamical | Without Dynamical |
|---|---|---|
| Mean requested iR-corrected potential at 10 mA/cm² | 1476.6 mV | 1488.2 mV |
| Mean regret to the pool best | 27.35 mV | 38.92 mV |
| Mean rank within the pool | 3.222 | 3.875 |
| Requested the pool best | 30/72 | 20/72 |
| Requested the bottom half | 9/72 | 20/72 |
All four models selected catalysts with lower archived potential when using Dynamical, averaging 11.6 mV lower at 10 mA/cm² across the study. Removing one decision-driving observation changed the final request in three of four checked trajectories.
In an exploratory comparison, when the predictor ranked the best candidate first, the selected catalysts averaged 25.9 mV lower potential than those chosen without Dynamical. Otherwise, they averaged 2.7 mV higher. The practical question is whether the model can distinguish the candidates well enough to guide the next experiment.
Evaluation details
The study used the FastCat dataset of nickel-foam oxygen evolution catalysts and runtime 0.1.2+fastcat2. Each agent selected one chemical-bath composition from its candidate pool. The archived score was the iR-corrected E_at_10mA potential at 10 mA/cm², where lower values were better. The predictor subtracts 1.229 V by Dynamical convention. This common offset preserves rankings and differences, and the reported comparisons use the archived potential scale.
The exact best-candidate counts were 30/72 with Dynamical and 20/72 without it. We averaged across models and repeats within each pool and treated the six pools as independent units. The 90% pool-cluster interval for the 11.57 mV potential improvement was −1.17 to +29.30 mV, with an exact sign-flip result of p = 0.3125.
The predeclared primary score estimated the information gained from the selected experiment after the evidence already acquired. Its paired effect was +0.080 nats, with a 90% pool-cluster interval of −0.081 to +0.246 and p = 0.5625. Both intervals include zero.
The predictor returned overpotential estimates and uncertainty within a fixed candidate table and passed nine predeclared admission gates. Mean absolute error was 22.2 mV on its 27-composition validation cohort and 48.4 mV across 54 study-pool compositions. Its fixed ±105 mV interval covered 96.3% and 90.7% of outcomes, respectively. The validation top-three retrieval score was 0.823 against a 0.45 threshold. Other virtual-lab outputs supplied execution feedback, while the FastCat predictor supplied the scientific estimates used for selection.
The observation-removal comparison held the model, task, seed, budget, and all other inputs constant. Within-pool predictor rankings were strong in three pools and weak in three. The exploratory comparison by whether the predictor ranked the best candidate first had p = 0.100. Provider fidelity was not randomized, and the analysis examined several provider-quality measures. Of the 54 archived outcomes, 43 came from a single physical run. Where replicates existed, the median within-composition range was 39.5 mV.
Materials investigations
Across four recorded campaigns, agents changed proposed tests, investigated faulty models, and identified missing process steps. The films span water electrolysis, additive alloys, critical-mineral recovery, and rare-earth magnets, pairing each recorded investigation with a replay of the virtual lab the agent composed.
The three later campaigns shared two NVIDIA DGX Spark systems for scientific computation, allowing agents to run calculations and virtual experiments in parallel.
01 / 04 campaigns
Film 01 · Claude Opus 5
Water electrolysis
Claude Opus 5 used FastCat estimates to compare nine catalyst compositions and changed its proposed physical measurement. When the archived outcomes were opened, its choice was the best in the pool.
Film 02 · GPT 5.6 Sol
Additive-alloy qualification
GPT 5.6 Sol compared virtual fatigue outcomes for an additively manufactured Ti-6Al-4V build and moved its proposed test from 875 to 800 MPa. It requested physical measurements on the vacuum-treated route to check the model’s survival estimates.
Film 03 · Qwen 3.8 Max
Critical-mineral recovery
Qwen 3.8 Max traced an unexpected recovery result to a defect in the inherited wash model. It revisited the route comparison and requested a physical experiment to test the recovery strategy for calcium-rich acid-mine-drainage feedstock.
Film 04 · Grok 4.6
Rare-earth magnet qualification
Grok 4.6 found that the inherited model predicted strong magnet performance while neodymium remained at its ore concentration. A composition-aware calculation returned 0.051 T against a 1.10 T requirement, and the agent requested review of the missing isolation step.
OpenUSD represents the lab, and NVIDIA Isaac Sim visualizes the recorded workflow. Scientific outputs come from the campaign's models and source evidence. Each caption describes how those outputs changed the proposed physical work.
Scientific judgment
GPT-6 Astra and Claude Fable 5.1 each completed two investigations using analytical models and archived physical measurements from one additively manufactured Ti-6Al-4V build. They examined whether a fatigue model fitted to hot isostatically pressed material could describe material treated in vacuum.
We linked each agent’s available evidence to its chosen action, saved implementation, and final claim, checking whether a stated change appeared in the calculation.
One Astra campaign tested whether defect geometry could explain a fatigue model’s errors on vacuum-treated titanium. After the initial measurements contradicted the model, it acquired postmortem defect records and reran the calculations with measured geometry and alternative defect locations. The errors decreased but remained large, so Astra retained its conclusion that the model was inadequate for fatigue-life decisions.
One Fable campaign changed which measurements it sought while keeping its model fixed. Another revised its fatigue model after two archived failures. Log-scale fatigue-life error was 14% lower across eight later failures, with four estimates improving and four worsening. In that same campaign, the agent recognized uncertainty in its specimen-axis estimate, retained the estimator, and later reported depth uncertainty that did not account for that problem.
These were substantive scientific investigations. The recurring difficulty was carrying the meaning of evidence into the next computation and final claim, even when the agent had already recognized the relevant limitation.
Study details
All four campaigns received the same question, tools, three skills, and simulated instrument allowance of $50,000 and 14 days. The task left models, acquisitions, revisions, and stopping decisions to the agents, with no required positive result. They used a study extension of Dynamical 0.1.21 with analytical models and access to NIST AMB2025-03 records. All measurements came from the archive. The supplied model warning and source information already suggested that transfer between the processing routes could fail.
Separate case and cross-case reviews by agents linked decisions to original traces, saved code, measurements, and final reports. The review followed four questions.
- What did the agent know before the analysis?
- What did the acquired evidence establish?
- What changed in the next calculation or decision, or was justifiably retained?
- Did the final account match the record?
The review tracked evidence already exposed through reports, filenames, and scheduling information, and distinguished encountered evidence from records that were available but unsought or unavailable.
The Fable comparison excludes the two fitting records, invalid tests, and tests that ended without an observed failure. Root mean squared error in log10 fatigue life fell from 0.777 to 0.667 across eight later failures. The calculation uses saved estimates and the archive service's test classifications. These specimens came from the same build, and their records were acquired after the revision.
The revision also widened the model's parameter envelopes, while leaving its runout probabilities unchanged. The envelopes describe parameter ranges rather than calibrated confidence intervals.
14% lower log-scale error after revision
Claude Fable 5.1 · F2
Four estimates improved and four worsened
Specimen / nominal stress
- 33 / 550 MPaCloserAbsolute log10 fatigue-life error before 0.266, after 0.080.
- 69 / 550 MPaFurtherAbsolute log10 fatigue-life error before 0.074, after 0.259.
- 17 / 650 MPaCloserAbsolute log10 fatigue-life error before 1.808, after 1.042.
- 45 / 650 MPaFurtherAbsolute log10 fatigue-life error before 0.332, after 0.434.
- 25 / 700 MPaCloserAbsolute log10 fatigue-life error before 0.699, after 0.324.
- 57 / 700 MPaCloserAbsolute log10 fatigue-life error before 0.520, after 0.503.
- 53 / 750 MPaFurtherAbsolute log10 fatigue-life error before 0.588, after 0.675.
- 37 / 850 MPaFurtherAbsolute log10 fatigue-life error before 0.517, after 1.181.
Absolute error in log₁₀ fatigue life
Lower is better
Log-scale RMSE 0.777 → 0.667
Eight later archived failures from one build. Two fitting records, invalid tests, and runouts excluded.
Five provisional categories organized the review, covering problem framing, measurement reasoning, experiment selection, revision and recovery, and stopping and verification. These describe decisions and consequences rather than an overall behavioral score. The package preserves original agent reports, reviewed findings, and unresolved source disagreements.
Research direction
Our goal is to build agents that develop and revise working models of physical systems as they investigate them, and use those models to decide which experiment, measurement, or intervention is worth doing next. That brings world modeling together with value estimation and scientific judgment.
We are studying how experience from one investigation can improve this process in another, across different materials, instruments, and operating conditions. Dynamical connects actions to their evidence and consequences so we can evaluate those decisions and train agents to make them better. The aim is to help teams discover and develop materials and processes they cannot yet reliably design.
Getting started
Install Dynamical in Codex, Claude Code, or another supported agent. The included skills help the agent establish the lab state, run the investigation, and prepare missing integrations.
The FastCat example supplies nine catalyst compositions and a campaign template. Give both to the agent, select the fastcat facility, and ask it which composition to synthesize and measure next.
The agent runs a separate virtual experiment for each candidate it tests and returns the evidence behind its recommendation, including any gaps that need to be resolved before physical work. For your own campaign, provide the scientific question, relevant models or data, and available instruments.
Related research
Proprio tests whether one agent's instrument experience helps the next. In simulated flake search, fresh GPT-5.6 Luna agents succeeded in 100% of sessions with repaired procedures and operating notes, compared with 67% using the original procedures.
Dynamical-SDL-1 turned 1,035 recorded materials experiments into controlled campaigns. Local and held-out assessments moved in opposite directions in 49% of comparable cases, showing why improvement on acquired evidence needs to be checked against other experimental outcomes.
In Training Scientific Judgment from Physical Experiments, our Qwen3.6-35B-A3B system combined post-training on recorded materials experiments with evidence guidance and test-time scaling. It achieved 23% lower average log loss than the base system, meaning its judgments of experimental outcomes and confidence better matched the recorded results.