LEARNING & PHYSICAL TIME · UPDATED 16 SEPTEMBER 2026
What the experiments show
Results from explicit-solution model checks, combined-reference training and a faster diffusion model with a physical clock.
Loading verified results…
What adding water-containing data changed
Two runs started from the same model and used the same number of updates. One continued with the existing data; the other included 2,056 water-containing SPICE configurations. Evaluation used 512 separate configurations across 64 solute groups.
Water-force errors fell 13.2%; solute-force errors were essentially unchanged. Relative-energy error across solvent configurations fell 15.7%, but total-energy errors on the original molecular datasets rose.
| Quantity | Control | With solvent data |
|---|---|---|
| Solute force (eV/Å) | 0.10286 | 0.10293 |
| Water force (eV/Å) | 0.07508 | 0.06517 |
| Solvent relative energy (eV) | 0.25742 | 0.21705 |
| Solvent total energy (eV) | 0.69025 | 0.79417 |
| OMol total energy (eV) | 0.27524 | 0.33435 |
| Original SPICE total energy (eV) | 0.25788 | 0.39354 |
The incumbent model is unchanged. These results do not qualify a model for biological reaction forces, free-energy barriers or reaction-rate prediction.
Both runs completed about 14,500 updates in 43–44 minutes on one RTX 2000 Ada. All 4,508 evaluation records per run passed independent numerical reopening; GPU shutdown was verified. Counts refer to configurations, not distinct reactions.
Uncertainty, difficult cases and next check
When resampling the 64 solute groups, the conditional 95% interval for water-force error change is −17.7% to −8.5%. For solute forces it is −2.0% to +1.7%, so improvement is not demonstrated. Relative-energy change has interval −20.8% to −3.6%. These intervals describe this fixed comparison, not variation across training seeds or all biological chemistry.
A difficult phosphorus-containing training configuration improved from about 54,000 to 6.56 eV/Å force error. It was used for training and remains inaccurate; this is not evidence of generalisation. The original model also failed on it, and the source labels were retained.
Water labels refer to the source topology of finite quantum clusters. They do not establish bulk salt concentration, pH or physical reaction times. New evaluation charges are −1, 0 and +1; existing dianion checks remain separate. Original model pretraining exposure is possible.
The old molecular energy regressions mostly concern shifts between molecular groups, while within-group variation barely changed. Next, check energy calibration using training data only, with network weights and forces fixed. No correction has been applied and no candidate has been promoted.
Download the comparison, source metrics and uncertainty analysis (JSON)
Earlier experiments are retained below for context.
What shared learning contributes
Both experiments used the same batches, initial model, seed and loss settings. BMS trained the shared model in both. The control let OMol/SPICE train only their output layers; the joint run also let them update the shared representation.
| Quantity | Output-only control | Joint model |
|---|---|---|
| Solution QM force (eV/Å) | 0.145214 | 0.145417 |
| Solution solvent response (eV/Å) | 0.004882 | 0.004883 |
| OMol energy (eV) | 0.520560 | 0.274111 |
| OMol force (eV/Å) | 0.087566 | 0.086383 |
| SPICE energy (eV) | 0.515435 | 0.264416 |
| SPICE force (eV/Å) | 0.084261 | 0.083771 |
Shared learning reduced OMol/SPICE energy errors by about 47–49% relative to the control, and also lowered their force errors. Overall solution QM-force errors differed by 0.14%, with the control slightly better. This supports retaining the joint approach at the tested budget.
14,464 updates; 89,201 BMS configuration visits and 3,268 unique paired configurations per auxiliary reference. All 3,996 validation outputs passed independent native-reference reopening. Runtime 44.6 minutes, peak allocated memory 1.61 GB; GPU shutdown independently confirmed.
Scope, reproducibility and next decision
This is one seed and one fixed training budget, not a universal architecture ranking. Initial record identities, targets and schedule match; small float32 differences in recomputed predictions are quantified in the download. Tiny solution differences should not be treated as conclusive model superiority.
The three native quantum references retain separate energies and force targets. The control’s auxiliary losses were checked at every update to ensure they could not alter shared parameters. No new quantum calculations were generated.
Keep the existing joint model as the experimental incumbent. Next, measure actual coupled GPU simulation throughput and extend fixed-model solution validation before a larger training release. Reaction rates, unseen biological chemistry and production deployment remain unqualified.
Download the matched comparison and independent audit (JSON)
First paired solution dynamics check
Both frozen models simulated the same glycylglycine molecule in explicit water: 1,655 atoms. Using a smooth classical solvent potential, halving the timestep reduced energy deviation by approximately fourfold for both models.
| Model | 0.25 fs timestep | 0.125 fs timestep |
|---|---|---|
| Released AMP | 0.2792 | 0.06951 |
| Joint model | 0.2801 | 0.06974 |
Complete-system energy derivatives and the preset finite-state, energy-deviation and timestep checks passed. This establishes a short mechanics test. Longer-time stability, quantum-reference accuracy and biological reaction rates remain unqualified.
What changed in the solvent calculation
The earlier PME probes retained cutoff energy jumps. Charge-pair crossings explained the saved directional-derivative discrepancies. This comparison uses the authors’ smooth reaction-field helper, unchanged Amber/TIP3P parameters and explicitly recorded Lennard–Jones switching. The learned weights and starting state were unchanged.
This is a distinct classical Hamiltonian from the earlier PME tests and uses a different water model from the full published TIP4P-FB setup. The earlier failed convergence observations remain available. Only one starting solvent environment and 50 femtoseconds per run have been tested.
Download the paired dynamics metrics and independent audit (JSON)
How the trained model handles our solution cases
Compared the frozen joint model with the released AMP model on 40 existing configurations across eight reaction examples, four families and 16 solvent preparations. Requested salt settings are 0 and 0.15 M; actual ion conditions remain in the source records.
| Quantity | Released | Joint | Change |
|---|---|---|---|
| QM force (kcal/mol/Å) | 2.2148 | 2.1730 | -1.89% |
| Nearby solvent force (kcal/mol/Å) | 0.0987 | 0.1026 | +4.03% |
| Relative energy (kcal/mol) | 0.5178 | 0.5163 | -0.28% |
Reacting-molecule forces improved slightly; solvent-force agreement worsened. All 40 predictions were finite. These exposed development cases use a different quantum reference method from BMS training. This is a practical transfer comparison, not an untouched generalisation test. Relative energies use 36 configurations; four singleton solvent roots have no energy-difference score.
All three known short-contact challenges still fail for both models. Joint training did not resolve their invalid forces. Their original quantum labels and failure evidence are retained.
Simulation-interface checks and remaining limits
Both models passed the authors’ force-interface fixture: quantum and solvent energy derivatives, translation, periodic image changes, cubic-axis rotation, cached neighbour lists and scripted execution. This small fixture does not establish stable dynamics.
The full solvated-system CPU derivative probe failed. A classical-only check on the same saved state isolated finite-difference precision: the double-precision Reference backend passed the unchanged tolerance. The failed probe remains recorded; no reaction-rate or production-model qualification follows from these checks.
Deployment uses the authors’ reduced polarization and periodic electrostatic cutoff; native quantum-reference comparisons use the original reference convention. These settings are kept distinct. In the earlier PME setup, both models completed matched 50-femtosecond runs. The smaller timestep increased maximum energy deviation by about 49–52%, so those runs did not establish timestep convergence. The separately reported smooth reaction-field comparison addresses this mechanics issue; longer-time stability remains unqualified.
Download solution comparisons, challenge results and mechanics evidence (JSON)
Learning across solution and molecular reference data
One shared model completed a pass through 89,201 BMS solution configurations, alongside controlled reuse of 3,268 paired OMol/SPICE configurations. Validation molecules are protected across sources. These are configuration counts, not distinct reactions.
| Quantity | Before | After | Error change |
|---|---|---|---|
| Solution energy (eV) | 0.19391 | 0.18886 | ↓ 2.61% |
| Solution QM force (eV/Å) | 0.14613 | 0.14542 | ↓ 0.49% |
| Solvent response (eV/Å) | 0.004945 | 0.004883 | ↓ 1.25% |
| OMol energy (eV) | 0.50924 | 0.27411 | ↓ 46.17% |
| OMol force (eV/Å) | 0.09438 | 0.08638 | ↓ 8.48% |
| SPICE energy (eV) | 0.50306 | 0.26442 | ↓ 47.44% |
| SPICE force (eV/Å) | 0.08526 | 0.08377 | ↓ 1.75% |
Full-QM energy errors fell substantially; overall solution performance was broadly retained. The reaction subset’s QM-force error rose 0.24% and solvent-response error rose 0.46%. This checkpoint has not been released for simulation or reaction-rate prediction.
2,346 solution validation configurations and 825 per full-QM method. About 52 minutes of GPU-job runtime across the initial segment and continuation; peak allocated memory 1.62 GB on one RTX 2000 Ada. Shutdown is confirmed.
Uncertainty, subset results and training limits
The overall solution QM-force change is small: −0.49%, with a conditional 95% interval of −1.33% to +0.22% when resampling 116 linked validation groups. It is not a conclusive improvement. The solvent-response change is −1.25%, with interval −1.95% to −0.69%.
Reaction relative-energy error fell 7.29%, although reaction-force errors rose slightly. Small-molecule energy error was nearly unchanged (+0.03%). Tripeptide energy and force errors decreased. These tests do not measure free-energy barriers or physical reaction rates.
Native quantum methods, geometry and force targets remain separate. Source-linked validation groups are protected; 1,342 records are explicitly omitted from this experimental training view and retained in the source bank. The three known numerical failures and their surrounding components are preserved.
Missing native C6 supervision is explicitly disabled in this experiment. Released-model pretraining exposure remains; these are not globally unseen-chemistry tests. The earlier standalone calibration is retained separately, and its learned outputs were not used to initialize this joint run.
Download joint-training results, subset metrics and independent audit (JSON)
EXTERNAL RETENTION CHECK PASSED
All 899 author-reserved peptide and miniprotein configurations produced finite results. Relative to the original model, quantum-force error fell 1.15%, solvent-response error 1.36%, and within-molecule energy error 3.26%. Both biomolecular subsets improved on these metrics. No fitting was performed during this check.
One molecular representation, two quantum references
We trained separate energy outputs for OMol25 and SPICE using 3,271 paired molecular configurations, then evaluated both on 825 validation pairs. Related molecules stay in the same split. The two methods retain their own energy and force targets.
| Method / quantity | Before | After | Error change |
|---|---|---|---|
| OMol energy (eV) | 0.5093 | 0.4179 | ↓ 17.94% |
| OMol force (eV/Å) | 0.09438 | 0.08610 | ↓ 8.78% |
| SPICE energy (eV) | 0.5031 | 0.4304 | ↓ 14.44% |
| SPICE force (eV/Å) | 0.08526 | 0.08541 | ↑ 0.17% |
Energy predictions improved for both methods. Force improvement is limited to OMol. The unchanged BMS solution model remains the baseline; this checkpoint has not been promoted for simulation.
13.0 minutes on one RTX 2000 Ada. Peak allocated GPU memory: 234 MB. The controller collected the outputs and confirmed shutdown. The independent audit reproduced both error tables and the training-only energy offsets.
Data, energy conventions and limits
These are full-quantum PubChem molecular examples used to integrate public reference sources. They add no new solution reactions. The wider explicit-environment BMS25 data and its quantum/solvent-force components remain separate.
Both sources use their native geometries and labels. Numerical energy offsets depend on element counts and charge and were fitted only on training records; they are not isolated-atom or formation energies. Forces are energy derivatives and errors are per Cartesian component.
The experiment uses a frozen released representation, two method-specific outputs, three fixed epochs and no repeated tuning. Released-model pretraining exposure is not erased by the project split. Missing SPICE spin metadata is preserved. Reactive dynamics, salt response and biological reaction rates remain unqualified.
Earlier BMS-only trial regressed
Both models produced finite results on all 899 reserved configurations. That earlier training trial made the overall errors worse.
The stopped training run processed over 42,000 configurations. This comparison uses its saved step-3,900 checkpoint; it is not a completed training epoch.
A faster diffusion model with a clock
The Brownian model predicts sodium displacement statistics on four fresh test trajectories within the fixed limits.
90,540 configurations checked automatically
| Source | Finite predictions | Numerical failures |
|---|
These are configuration counts, not distinct reactions. The quantum reference labels and split ownership are preserved. Finite output alone does not establish accuracy; the new training entry point requires a complete, matching numerical check.
Inspect the three model failures
All three tripeptide configurations contain very short H–H contacts. Their model predictions have nonfinite molecular forces. They remain in the source data; they have not been deleted or relabelled as bad quantum references.
CLASSICAL SOLUTION → BROWNIAN DIFFUSION
Does the faster model predict the test trajectories?
Six independently seeded trajectories estimate the diffusion coefficient. Four new trajectories test it. The two earlier failed test trajectories remain separate and were never used to fit this result.
━ Brownian prediction ● Held-out molecular dynamics
Time is in picoseconds (ps), not animation frames. One ps is 10⁻¹² seconds.
Conditions, uncertainty and limits
TIP3P water with one sodium/chloride pair in a 3 nm periodic box; 298.15 K, approximately 0.0615 mol/L salt. Each trajectory has 100 ps equilibration and 2 ns production with a 2 fs integration step.
The checks cover sodium self-diffusion statistics at 5, 10, 20, 40 and 80 ps. Uncertainty resamples whole independent trajectories. The reference uses a classical potential and Langevin thermostat.
This is not a reaction-rate measurement or experimental diffusion validation. Protein motion, ion association, membranes and organelle-scale transfer remain untested.
899 AUTHOR-RESERVED BMS25 CONFIGURATIONS
Earlier BMS-only checkpoint (step 3,900)
714 peptide and 185 miniprotein configurations. Both models see the same test inputs; no test configuration was used for this additional fitting.
| Quantity | Released AMP | Step 3,900 | Change |
|---|
Force errors are root mean square per Cartesian component. Solvent forces here are the quantum embedding response, not total classical solvent forces. Relative energies are compared within each molecule across 895 configurations; four singleton configurations cannot supply an energy difference.
What the aggregate comparison does not show
Peptide relative-energy error improved, but force errors worsened for both peptides and miniproteins, and miniprotein relative-energy error increased. Separately, both models fail on an exposed high-energy tripeptide configuration. Finite results on these 899 examples do not establish numerical stability everywhere or validate reactive dynamics.
Inspectable evidence
Broader solution comparison: 40 existing configurations across eight reaction examples and four families produced finite predictions. Molecular-force disagreement was 2.21 kcal/mol/Å and within-root relative-energy disagreement 0.52 kcal/mol. The quantum methods differ, and these retain development ownership. Read the cross-reference comparison and audit (JSON).
Downloaded results and source identities were checked against the fixed inputs. Independent recalculation reproduced the model-error summaries. The diffusion analysis preserves the original failed assessment and unchanged accuracy limits.
Download results and provenance (JSON)Model comparison: 82 seconds, peak allocated GPU memory 1.15 GB. Eight additional diffusion trajectories: 22.9 minutes. Both jobs used one RTX 2000 Ada and have confirmed GPU shutdown.