Making Science Programmable.

Autonomous R&D from engineering objectives to qualified materials and processes for complex physical products.

Research

Training Scientific Judgment from Physical Experiments

We trained scientific judgment in a 35B open model on replay environments compiled from lab experiments, then scaled inference. The final system completed every campaign, kept its claims within the measured evidence boundary, improved forecasts about unseen chemistry, and reached the open-weight frontier on SDL-1.

August 2026

Dynamical-SDL-1: Measuring How Scientific Agents Use Experimental Evidence

Dynamical-SDL-1 compiles a complete physical reaction map into controlled long-horizon campaigns that measure how efficiently scientific agents convert experimental evidence into better decisions. Across six systems and six campaign assignments, local evidence assimilation, held-out transfer, and adaptive experiment selection emerge as distinct capabilities.

July 2026

Simulator-Verified Skill Acquisition for Scientific Instruments

Proprio gives an agent a persistent simulator loop to draft, execute, inspect, and repair an instrument operating skill, while an independent verifier it cannot change decides what enters the catalog. Verified feedback produced 14 non-regressive repairs from 18 paired drafts where blind retrying produced none, and one frozen protocol acquired, repaired, verified, and evolved skills across three external instrument families with zero invalid promotions. Verified in simulation. Hardware validation remains separate.

July 2026

Can a Self-Driving-Lab Agent Tell When the Evidence Is Enough?

We turn the historical record of a lab into source-located replay tasks that measure evidence-boundary judgment before a self-driving lab is trusted to run on its own. Across six frontier models and 1,872 trajectories, agents reach a valid decision on 90% of runs and the reference-equivalent path on 72%, and no model clears the benchmark.

June 2026

Scaling Test-Time Verification for Novel Materials

Crystal diffusion models encode property signals in their hidden states that they never use during sampling. Probe-gradient guidance steers unconditional generation toward target properties at test time, comparable to conditional models at over 50x the speed.

April 2026