Research

Scaling Test-Time Verification for Novel Materials

PublishedApril 7, 2026
UpdatedSeptember 9, 2026

We tested whether property estimates could guide crystal generation toward a desired band gap, an electronic property that distinguishes metals, semiconductors, and insulators. The study used Crystalite, a roughly 67-million-parameter diffusion transformer that generates crystal structures without a property target, and MatterGen, a diffusion model conditioned on desired material properties.

With Crystalite’s generator fixed, a small trained probe guided its generation process and increased the share of candidates in the predicted 4–6 eV band-gap range from 0% to 24%. A Crystalite checkpoint trained on a more balanced mix of metals and insulators reached 43% with guidance, while passing lattice checks for 100% of candidates and geometry checks for 99.6%.

Generating candidates is one part of materials discovery. Diffusion models like Crystalite (Hadži Veljković et al.) and MatterGen (Zeni et al.) propose novel candidates by jointly denoising atom identities, fractional coordinates (the position of each atom within the repeating unit cell), and lattice geometry from noise. Neural potentials screen for thermodynamic stability in seconds. Of over half a million candidates proposed by GNoME, fewer than one in five exhibited predicted synthesizability in subsequent analysis. The synthesis gap is where materials discovery stalls.

Multi-fidelity verification pipelines stack these layers, each adding confidence and cost. The question is where in this pipeline verification should start.

Scientific agents need to decide which evidence to trust and when to request a more expensive check. This work improves the candidate pool they evaluate by guiding generation and selecting which structures receive further computational tests.

Test-time verification pipeline: candidates flow from a generative model through three increasingly expensive verification stages (cheap constraint checks, proxy models, high-fidelity oracle/lab) with a training signal feedback loop from verified solutions back to the generator
Multi-stage verification pipeline. Candidates pass through increasingly expensive checks. Verified outcomes feed back as training signal to improve generation.

Self-correcting search (Hazra et al., Goodfire Research) exploits a gap between what diffusion models encode internally and what they express in their outputs. A small trained probe reads MatterGen's GemNet hidden states to predict what the band gap of the final structure will be at each intermediate denoising step. Self-correcting search uses this prediction as a feedback loop. At each step, the probe evaluates the proposal, and Metropolis-Adjusted Langevin Ascent accepts or rejects it based on whether it moves toward the target property range.

Standard conditional generation with MatterGen achieves 15% of samples in the target band-gap range, dropping to 6.5% when filtering for stable, unique, and novel candidates. Self-correcting search pushes forward the targeting-viability frontier at every conditioning strength tested, generating ~30% more viable candidates in the target range. The approach is general. Any property predictable from model internals becomes a viable steering signal.

We reproduced these results on a single GB10 GPU to understand the method's behavior across conditioning strengths and random seeds.

Two-panel chart showing MatterGen self-correcting search reproduction, targeting vs guidance strength and seed robustness at gamma=1.0
Our MatterGen reproduction across conditioning strengths and seeds. Stability is checked after relaxation; the plotted band-gap estimates are from before relaxation. Self-correction reaches 25% on the combined measure at gamma=1.5. N=32 per arm, MatGL surrogate, single GB10 GPU.

At gamma=1.5, self-correction doubled the share of candidates that passed both the stability and band-gap checks, from 12.5% to 25% when both were scored after relaxation. Band-gap targeting alone reached 44%. Results varied across conditioning strengths and seeds. A best-of-three variant with a minimum predicted band gap improved the weaker runs at lower conditioning strengths, while the original method gave the strongest result at gamma=1.5.

Model internals carry verification signal that meaningfully improves generation. This was the starting point for our work. The natural question was whether the same principle extends beyond conditional models. MatterGen receives the band-gap target as input at every conditioning strength. Can test-time verification steer a model that has no conditioning signal at all?

Probe-Gradient Guidance

Crystalite (Hadži Veljković et al.) is a ~67M-parameter Diffusion Transformer that generates crystal structures unconditionally. No property target enters the model. It denoises atom tokens, fractional coordinates, and lattice descriptors jointly from noise, using a subatomic tokenizer that encodes elements as continuous vectors derived from periodic table position and valence structure. Its training data is Alex-MP-20 (675,204 structures total, 540,162 in the training split), 97.9% of which are metals by our analysis. The trained model therefore has little experience with the wide-band-gap region we want to target.

We trained a small timestep-conditioned two-layer probe MLP on Crystalite's atom-mean hidden states to predict band gap. The probe has 256 hidden units and 197,890 parameters in total. The probe achieves 0.957 AUROC for identifying structures in the target band-gap window. The model represents whether a structure is metallic or insulating at every denoising step, despite never being trained to condition on that property. Kreiman et al. show this is a general property of atomistic transformers. Given sufficient data, standard attention learns interatomic structure without hard-coded equivariance, encoding physics the model was never explicitly trained on.

We first applied the same Metropolis accept/reject strategy that works on MatterGen. On an unconditional model, it failed across all 36 configurations tested. Every one produced 0% in-window and 97-100% metals. The accept/reject step selects among proposals the model generates, so it needs useful candidates to enter that pool. Gradient guidance gives us a way to change the proposals themselves.

Instead of accepting or rejecting proposals, we backpropagate through the probe at each denoising step to produce a gradient on the generation trajectory. The gradient pushes the denoising path toward structures with higher predicted band gap.

Without guidance, the model generated 96.5% metals and no candidates in the target window. At w=1, metals drop to 0.8% and 3.5% of structures hit the 4-6 eV range. At w=10, metals reach 0.0% and 24.2% of structures hit the target, with a mean band gap of 4.19 eV. The trained probe opens a target region that the fixed, unconditional generator rarely visits on its own.

Two guidance methods

Both methods use a model's internal representations to improve generation. MatterGen begins with a requested property and uses a probe to select proposals along its sampling path. Crystalite begins without a property target, and the probe supplies gradients that change that path. MatterGen uses an equivariant graph neural network, while Crystalite uses a diffusion transformer.

MatterGen with self-correction

25% passed both stability and band-gap checks at gamma=1.5, compared with 12.5% without self-correction, using estimates after relaxation. Band-gap targeting alone reached 44%.

The model receives the target property during generation. The probe helps it choose among proposals, with performance varying across conditioning strengths and seeds.

Crystalite with gradient guidance

24.2% reached the predicted band-gap window with the base generator fixed. A checkpoint trained on a balanced dataset reached 42.6%, with 100% lattice validity and 99.6% geometry validity.

The generator receives no property target. A trained probe supplies the steering signal, allowing a new target to be introduced without retraining the generator.

The checks answer different questions. MatterGen's combined result includes a stability test after relaxation. Crystalite's band-gap, lattice, and geometry results measure targeting and structural checks separately.

Generation speed also changes the search budget. In the recorded sweeps, guided Crystalite generated 1,024 structures in about 146 seconds, while MatterGen's gamma=1.5 self-correction run generated 32 in about 902 seconds. These timings describe the two sampling configurations. Crystalite's throughput makes broad guidance sweeps practical, and changing the property requires training a small probe rather than the generator. The probe is less than 0.3% of the generator’s size.

We then tested whether stronger guidance retained the variety of generated compositions.

Bar chart showing in-window rate climbing from 0.1% at w=0 to 33.7% at w=15, with error bars across 3 seeds. Annotations show compositional uniqueness held at 99.6-99.9% across all weights.
Mean targeting improves with guidance weight while compositional uniqueness remains 99.6–99.9% and mean novelty exceeds 99% at guided settings. The sweep covers 18,432 structures across six weights and three seeds, with 1,024 per batch.

Across this larger sweep, mean targeting reached 31.8% at w=10 and 33.7% at w=15, while mean compositional uniqueness remained 99.6–99.9% and mean novelty stayed above 99%. Guidance shifted the generated compositions without reducing the number of distinct chemical systems explored.

Training for valid structures

The base model generated diverse compositions, but only 5–16% passed structural checks in the reported sweeps. Training on a subset with 35% insulators improved the geometry of guided candidates. At w=3, the balanced checkpoint reached 42.6% in the target window, with 100% lattice validity and 99.6% geometry validity. Compositional uniqueness was 78%, compared with 99.7% for the base model. Training changed which structures the model could generate reliably, while guidance determined where it searched.

A formation-energy probe also achieved 0.990 AUROC. Energy above the convex hull requires comparison with competing compounds, so that check needs a consistent reference dataset. For compositional constraints, token masks can directly include or exclude elements, while probe gradients steer continuous properties such as band gap.

Hull-probe diagnostic

The recorded hull-probe diagnostic was 0.000 AUROC. The training code also used zero when the validation labels contained only one class, where AUROC is undefined. That record alone cannot establish whether model representations carry a useful hull-related signal. Computing energy above the hull requires consistent energies for the candidate and its competing reference phases.

Integrating Verification into Generation

Does the verification engine improve a conditional model that already has its own search procedure? We tested this on MatterGen with a fixed budget of 100 expensive validations per method. The three variants are self-correcting search alone, conditional generation with the engine replacing self-correction, and self-correcting search composed with the engine.

VariantYield @0.15 eVYield @0.10 eV
Self-correctionband_gap_only55
Self-correctionband_gap_rohs55
Conditional + engineband_gap_only5.75.3
Conditional + engineband_gap_rohs4.74.3
Self-correction + engineband_gap_only109
Self-correction + engineband_gap_rohs99

Validated hits per 100 expensive evaluations on MatterGen. e_above_hull thresholds in eV. band_gap_only selects on predicted band gap; band_gap_rohs adds a rule-of-thumb hardness screen. Self-correction and self-correction + engine are single-seed; conditional + engine is averaged across three seeds.

Replacing self-correction with the engine alone produces mixed results, slight gains on one variant and slight regression on the other. Composing the engine with self-correction doubles validated yield. On the same candidate pool, with the same evaluation budget, the composed pipeline produces 10 validated hits per 100 evaluations versus 5 for self-correction alone. At the stricter threshold, 9 versus 5. The pattern holds across both selection variants.

The benefit came from combining self-correction with selection, using the same expensive evaluation budget more effectively.

From Generation to Synthesis

Property estimates can improve both how candidates are generated and which candidates receive expensive computational checks. The fixed Crystalite generator reached a new band-gap region through probe guidance, balanced training improved structural validity, and combining selection with MatterGen's self-correction increased validated yield under a fixed evaluation budget.

The broader opportunity is to learn which proposed materials are worth testing under specified physical conditions. Synthesis records connect what was attempted, the process used, and what formed. They could provide supervision for models that estimate whether a candidate can be made and which experiment would resolve the remaining uncertainty.

This connects candidate generation to scientific judgment. An agent needs a useful model of the material and process, an estimate of what the next measurement could change, and a way to carry that evidence into its next decision. Computational verification helps allocate the search budget; physical experiments provide the feedback needed to improve that understanding.


Citation

@article{barnes2026verification,
  author  = {Barnes, Jarrod},
  title   = {Scaling Test-Time Verification for Novel Materials},
  journal = {Dynamical Systems},
  year    = {2026},
  url     = {https://dynamicalsystems.ai/blog/scaling-test-time-verification}
}

If you are running autonomous experiments, building generation-to-synthesis pipelines, or working on synthesizability prediction, we want to hear about it.