DooABLe / docs /validation.md
pranamanam's picture
Upload 309 files
81ae663 verified
|
Raw
History Blame Contribute Delete
3.78 kB
# Completed validation
Validation used CPU execution with the versions in `environment.json`. Numerical results below are measured outputs from the included runs. The manuscript retains placeholders for its main scientific comparisons.
## End-to-end data checks
| Dataset | Clean molecules | Train | Test | MAE | R² |
|---|---:|---:|---:|---:|---:|
| CACO2 | 897 | 717 | 180 | 0.4454 | 0.4707 |
| BACE | 1513 | 1210 | 303 | 0.6474 | 0.5823 |
Caco-2 MAE uses log10(cm/s). BACE MAE uses pIC50. Both fits use scaffold groups, split seed 0, and 256 extra trees. Downloaded data were parsed and canonicalized before splitting. Source checksums, split assignments, test predictions, and fitted weights are included.
## Generative reference runs
| System | Nodes | Edges | Outcomes | Seeds | Updates | DooABLe endpoint TV | Conditional gap |
|---|---:|---:|---:|---:|---:|---:|---:|
| Multiplicity | 12 | 18 | 2 | 5 | 1200 | 5.6311e-05 | 1.6247e-07 |
| Lattice | 67 | 147 | 22 | 5 | 2000 | 0.0519513 | 0.148841 |
| String edits | 35 | 60 | 16 | 1 | 1200 | 0.00411287 | 0.00049674 |
| Reaction graph | 337 | 410 | 168 | 1 | 1200 | 0.000863913 | 1.01626e-05 |
Conditional gaps use route-cost units and temperature 0.7. Per-seed values, standard errors, and comparator measurements are stored with each run. The lattice and string runs have lower endpoint TV under uniform-backward trajectory balance at the recorded training budgets. DooABLe has lower conditional route divergence on those runs. These mixed outcomes are retained in the supplied CSV files.
The reaction graph contains 12 parents, 13 reagents, four directional templates, and a budget of two additional reactions. All 1,000 DooABLe samples replayed successfully, with 166 unique outcomes. Scoring evaluated all 168 canonical outcomes before learning. Mean reaction count among the sampled routes was 1.395. Property summaries in `results/chemistry_small/candidates.json` are model predictions.
The ablation run uses five seeds and tests uniform backward probabilities, zero execution cost, duplicated terminal outcomes, unnormalized backward weights, and exact prefix values. The exact sweep includes 50 multiplicity/temperature comparisons. The public preference sweep includes 15 exact reward settings. Budget-integration runs exercise the complete molecular runner at B = 0 and B = 1 with 400 updates and 100 samples.
## Software checks
The 24 tests cover exact path sums, endpoint normalization, cost/temperature behavior, deterministic error bounds, cycles, budgeted expansion, chemical replay, duplicate identifiers, scaffold isolation, serialization, checkpoint reload, and resumed-training equivalence.
The documented public workflow was executed from data download through property fitting, graph construction, neural training, sampling, replay, candidate evaluation, and plot export. Existing weights support immediate inference. New training checkpoints contain optimizer state for resumption.
Full licensed-catalog generation, held-out graph-instance generalization, LIT-PCBA docking, learned molecular encoders, and the original third-party baselines remain unexecuted. Their concrete dependencies are listed in `experiments.md`.
The final ZIP files were extracted into clean directories. The manuscript compiled to eight main pages and 20 total pages, with 72 references, zero unresolved citations, and zero overfull boxes. The repository was installed into a fresh virtual environment using the recorded installed dependencies. All 24 tests passed. CLI checks covered toy training, checkpoint resumption, sampling, chemistry graph construction, property scoring with the included weights, exact sampling, and replay of 50 chemical routes. Figure 1 is identical in the manuscript and repository.