# Trained machine-learning models for molecular absorption-spectrum prediction

This record provides five trained machine-learning models for predicting molecular absorption spectra and associated normalization constants from molecular atomic structure or three-dimensional ground-state electron density. The models were developed using QM7 molecules and electronic-structure calculations of ground-state electron densities and absorption spectra.

The release supports inference, reproducibility and comparison of density-based and structure-based approaches to molecular spectroscopy.

## Model files

|Filename|Model|Input|Prediction|
|-|-|-|-|
|`cnn\\\\\\\\\\\\\\\_model\\\\\\\\\\\\\\\_spectra.pt`|Three-dimensional convolutional neural network (CNN)|Ground-state electron-density grid|Normalized absorption spectrum|
|`cnn\\\\\\\\\\\\\\\_model\\\\\\\\\\\\\\\_normconst.pt`|Three-dimensional CNN regressor|Ground-state electron-density grid|Spectrum-normalization constant|
|`mace\\\\\\\\\\\\\\\_model\\\\\\\\\\\\\\\_spectra.pth`|MACE|Atomic species and Cartesian coordinates|Normalized absorption spectrum|
|`schnet\\\\\\\\\\\\\\\_model\\\\\\\\\\\\\\\_spectra.pt`|SchNet|Atomic species and Cartesian coordinates|Normalized absorption spectrum|
|`dimenetpp\\\\\\\\\\\\\\\_model\\\\\\\\\\\\\\\_spectra.pt`|DimeNet++|Atomic species and Cartesian coordinates|Normalized absorption spectrum|

The SchNet and DimeNet++ files are PyTorch checkpoint bundles containing trained weights, model identifiers, architecture settings and training configuration. The associated software reconstructs the models from these bundles. The bundles require compatible model definitions and software dependencies for inference.

## Software and usage

Installation instructions, model definitions, inference workflows, input requirements and preprocessing guidance are maintained in the associated GitHub repository:

[**QM7-Absorption-ML — repository and README**](https://github.com/stfc-ai4s/QM7-Absorption-ML#readme)

To use the models:

1. Download the required model files from this record.
2. Follow the repository README to install a compatible software environment.
3. Prepare molecular geometries or electron-density grids according to the repository's instructions for the selected model.
4. Use the documented inference workflow, supplying the downloaded checkpoint and any required configuration.

Refer to the repository for current commands, script locations and supported options. Use the software and configuration identified there as compatible with these released checkpoints.

The normalization CNN requires its matching training configuration to recover predictions in the original spectrum-sum units. This supporting configuration is maintained in the GitHub repository, separately from the weights record. Follow the repository instructions to locate and use it with the corresponding checkpoint.

## Spectral normalization

The study evaluates spectra over 900 energy bins from 0.0 to 0.4495 atomic units, with a spacing of 0.0005 atomic units. Spectral predictions are normalized to unit sum within this window.

The normalization CNN predicts the sum of the corresponding reference-spectrum values. Multiplying a unit-sum spectral prediction by the predicted normalization constant recovers the calculated spectrum's intensity scale, provided both predictions refer to the same molecule, energy window and target preprocessing convention.

The structure-based models predict spectral shape from atomic species and coordinates. Using the normalization CNN to recover the intensity scale additionally requires a ground-state electron density.

## Data and reproducibility

The associated ground-state densities, molecular geometries and calculated reference spectra are available online through the sources listed in the GitHub repository: https://github.com/stfc-ai4s/QM7-Absorption-ML

Reproduction requires the matching input representation, preprocessing, energy grid and normalization conventions. Validation-set reproduction also requires the corresponding data split and evaluation inputs described in the repository.

These models were developed for the chemical and numerical domain represented by the training data. Predictions outside that domain require separate validation. The conventional density CNN is not intrinsically rotationally invariant, so electron-density inputs must follow the training orientation and spatial-sampling conventions.

## 

