The Future Is Thousands of Labs
A thesis on experiments that compound.
The abundance of software is creating more demand for hardware.
More intelligence now means more power, more compute, more autonomy, more aerospace capacity, more resilient supply chains, and more systems operating at physical extremes. To build that physical future, we need new material capabilities.
What is worth exploring? And once a candidate exists, what are we actually capable of building?
For most of history, the first question was the hard one. Good hypotheses were hard to find, and discovery moved at the speed of human imagination disciplined by experiment.
Models can now propose structures, mechanisms, synthesis routes, and experimental plans faster than any physical lab can validate them. The frontier is moving from a shortage of plausible ideas to a shortage of trusted contact with reality.
Before a lab is a room of instruments, a furnace schedule, a sample queue, or a set of methods, it is a disciplined way of letting the physical world answer. It turns matter into evidence, evidence into judgment, and judgment into the next experiment.
The physical world is becoming the training environment for AI, but today's labs weren't built to emit training signal.
From execution to experience
Experimental knowledge once lived only in skilled hands, notebooks, and the credibility of witnesses. Instruments widened what the lab could see; the 1915 Nobel Prize went to reading crystal structure from X-rays, and electron microscopy pushed past the optical limit. Computation made parts of the search space navigable before anything was made, and the Materials Genome Initiative made the infrastructure thesis explicit. Automation changed the tempo, closing loops between planning, synthesis, characterization, and analysis. The agentic lab extends the sequence by making experimental experience trainable.
An agentic lab connects physical results to the models, simulations, and decisions guiding the next experiment. Experiment as code separates intent from execution. The physical result still has to return as a structured object that can train agents, update verifiers, calibrate simulators, and become replayable experience.
The lab could leave much unsaid because the people around the workflow supplied the missing context. A candidate can clear every computational screen and never form in the furnace. A synthesis can fail twice and succeed the third time because someone changed a precursor and remembered why. A measurement can be trusted because of how the sample was prepared, not only because of what the instrument returned. An agent needs the context behind these records to learn from them. A failed synthesis becomes more informative when the record preserves the material state, process conditions, and reason for the next attempt. A characterization result can help calibrate a simulator when its provenance connects the signal to the specimen. Failures, deviations, and expert corrections deserve the same care as successful runs.
The fields where AI compounds fastest are the ones where checking an answer is cheap. Code compiles and runs against its tests. A proof either checks or it does not. A materials claim is checked by a furnace, a diffractometer, and months of testing, and generation has already outrun that check; GNoME reported 2.2 million new crystal predictions, including 380,000 predicted stable materials.
When candidate generation becomes abundant, verification becomes scarce. When agents can propose more plausible actions than a lab can execute, the bottleneck becomes deciding which physical actions are worth taking and returning what happened as signal.
What we build
Dynamical is building a materials neolab for extreme environments, developing material-process capabilities and the scientific systems needed to discover, make, understand, and scale them. Customer requirements and internally led research give this work its direction. A useful result changes what a product can do, how reliably it can be made, or the conditions under which it can operate.
We start with a consequential requirement and keep the scientific route open. Service temperature, lifetime, throughput, and delivered cost impose real constraints; an inherited alloy, geometry, or fabrication route may be a choice worth revisiting. Material, manufacturing process, and component design can develop together. The opportunity might be a new composition, a different interface, or a process that makes an established material useful in a new application.
Deep materials expertise and a shared scientific platform develop together. The platform connects research questions, physics and instrument models, experimental plans, observations, and decisions. Our planned reference lab will support repeated interventions, trusted measurements, and validation of material-process capabilities, alongside specialist facilities and manufacturing partners. The goal is to carry useful capability into production with the process knowledge, measured performance, operating limits, and evidence needed for adoption.
There are two ways this work should compound. A material-process capability can reach additional products, machines, suppliers, and operating conditions. At the same time, the experience of developing it can improve the models and agents conducting the next investigation. Repeated adoption supports deeper research and new capabilities; better scientific understanding should reduce how much must be rediscovered in each program. We measure physical transfer and scientific learning separately because each determines a different part of that value.
Our published research examines the decisions that make this possible. VOE-Bench turns recorded manufacturing evidence into tests of whether an agent has enough support for its conclusion. Across 1,872 runs, 58% of invalid decisions referenced evidence the agent had not inspected. SDL-1 separates understanding an experimental result from choosing an informative next experiment. In Scientific Autoresearch, agents conducted substantive investigations, but did not consistently carry the meaning of evidence into their next calculation or final claim.
These observations guide improvement. Our post-trained Qwen system, combined with evidence guidance and test-time scaling, achieved 23% lower average log loss than the base system, improving its assessments of experimental outcomes and confidence. Proprio showed how corrections to a simulated instrument procedure could help a fresh agent operate more reliably. Recorded experiments make these comparisons reproducible and expose scientific behavior that the original datasets never contained. New physical campaigns extend the work to interventions whose outcomes do not yet exist.
Information throughput
Real throughput is information throughput, how much an experiment improves a consequential decision per instrument hour, sample, dollar, and expert intervention. A measurement is valuable when it distinguishes explanations that would lead us to do different things. That may mean changing the material, revising a process, requesting a more discriminating measurement, or retaining an approach that the evidence still supports.
A scientific agent needs a working model of more than the material. It must understand what conditions the process delivered, how those conditions changed the specimen, what the instrument measured, and how that measurement relates to performance. A discrepancy might come from missing physics, a drifting sensor, or a process that never delivered the intended conditions. Each explanation calls for a different next experiment.
The resulting record connects a prediction to an intervention, a calibrated observation, a discrepancy, and the decision that follows. It preserves raw signals and instrument state alongside process history, material behavior, and functional outcomes. These connections make experimental environments useful for evaluating and training scientific agents. We can examine whether an agent chose informative evidence, diagnosed the right problem, changed the affected calculation, and preserved the correction in its final decision.
In a digital model, loss can move backward through a network. In scientific discovery, the analogous signal has to move through the choices that produced an outcome, from the candidate proposed, to the simulation that screened it, the protocol that tried to make it, the instrument that measured it, and the interpretation that shaped the next action. Experimental success and failure become feedback for both the scientific models and the policy deciding how to use them.
The highest-value signal is often human judgment. A staff scientist can see that a phase assignment is too convenient, that a substrate peak is being mistaken for the material, that a sample charged under the beam, that a result is not yet strong enough to move toward qualification. Representations of experimental evidence have to preserve what made that judgment useful and how it changed the investigation.
Beyond optimization
As generation, simulation, and lab throughput improve, we can explore more possibilities before committing to a physical experiment. The larger opportunity is to discover capabilities that existing materials and processes cannot deliver. A scientific agent must be able to question the assumed mechanism, the available measurements, and the route to the objective as evidence arrives.
Our test-time verification work shows that property information inside a crystal generator can guide its generation process. The next step is to connect that control to models of processing, measurement, and useful performance. Physical observations can calibrate those models, expose where they fail, and identify which calculations deserve more trust. This matters most at the defects, kinetics, and extreme conditions where model predictions can become unreliable.
The research ambition extends beyond retaining the last successful procedure. Can experience teach an agent how to learn an unfamiliar physical system? That means identifying which uncertainty matters, deciding what to measure, distinguishing a flawed model from a flawed observation, and revising its understanding while the investigation is still underway. Training and evaluation must test these capabilities on new scientific problems, instruments, and physical conditions.
This is also why the work matters for frontier models. Materials campaigns can produce independently assessable examples of scientific judgment with consequences grounded in experiment. Those examples can become environments and training data whose value is measured by how much they improve subsequent investigations. The physical capability creates value in its application, while the scientific experience can improve how we discover the next one.
Thousands of labs
The future is still physical. It is furnaces, films, powders, wafers, coupons, thermal cycles, corrosion tests, failed batches, and expert doubt.
Software demand is becoming materials demand, and materials demand is becoming qualification demand. Qualified materials and components are now a rate limit on strategic capacity. University shared facilities, national labs, industrial R&D labs, characterization centers, pilot lines, foundries, and qualification labs hold the expertise and equipment that turn a promising result into a usable capability. Connecting their work means preserving the scientific context as an experiment moves between models, instruments, and people.
Scientists define the problems, develop explanations, challenge measurements, and set the standard of proof. Agents expand the investigations they can conduct and the experience they can draw on. When materials research moves at the speed of product development, the material and the machine can be designed together, with hardware and software advancing against the same physical evidence. An engineer's operating requirement becomes the starting point for discovering a material-process capability, testing its limits, and producing the evidence that supports qualification.
Our mission is to make science programmable. Thousands of existing labs become programmable, verifiable, and compounding.
We are shaping a new way of approaching agentic science, and we are looking for people who want to build it with us.