STRIKER Functions

run_STRIKER_model(input_file, output_file, log_fn=None, should_stop=None)

Correct unknown or inconsistent adducts in a spectral library using the pre-trained STRIKER model.

When to Use

Use this tool when your spectral library contains missing or incorrect adduct annotations and you want to correct them to avoid data loss during metadata normalization.

STRIKER Model

The STRIKER model is a pre-trained MLP classifier built on diverse, well-processed adduct datasets. It achieves approximately 98% accuracy and is recommended for most users.

Recommendation

Apply STRIKER as the first step in your spectral library curation workflow.

Parameters:
  • input_file (str) – Path to the spectral library in MSP, MGF, or JSON format containing a mix of valid, unknown, and blank adduct metadata fields.

  • output_file (str) – Directory where STRIKER will save the corrected library and reports.

Returns:

The function writes the following files to output_file:
  • corrected_library_STRIKER_model_correction_report.txt: A comma-separated text file listing unique corrections in your library. The first column contains the adduct, the second the correction, and the third the frequency of each correction.

  • corrected_library_STRIKER_model_number_of_unk_adducts.txt: Reports the number of blank or unknown adducts.

  • corrected_library_STRIKER_model.*: The corrected spectral library in the same format as the input.

Return type:

Files

Examples

>>> from striker import correct_adduct
>>> correct_adduct.run_STRIKER_model(
...     input_file=r"/example_data/tab1-correct-the-adduct/example1-predict-using-pre-trained-model/input/HMDB.msp",
...     output_file=r"/example_data/tab1-correct-the-adduct/example1-predict-using-pre-trained-model/output"
... )
run_another_model(input_file, output_file, model_path, log_fn=None, should_stop=None)

Correct unknown or inconsistent adducts in a spectral library using a pre-trained model.

This function is intended for expert users who have trained their own STRIKER models using the “Train the Model” workflow. Only pre-trained model files ending with _model.pkl are supported.

Parameters:
  • input_file (str) – Path to the spectral library file in MSP, MGF, or JSON format. The library may contain a mix of valid, unknown, or blank adduct metadata fields.

  • model_path (str) – Path to the pre-trained model file ending with _model.pkl.

  • output_file (str) – Directory where STRIKER will save the corrected library and related reports.

Returns:

The function writes the following output files to output_file:
  • <name>_corrected_library_<model_name>_model.*: Corrected spectral library.

  • <name>_number_of_unk_adducts.txt: Number of spectra with unknown adducts.

  • <name>_correction_report.txt: Summary of corrections applied to the library.

Return type:

Files

Examples

>>> from striker import correct_adduct
>>> correct_adduct.run_another_model(
...     input_file=r"/example_data/tab1-correct-the-adduct/example1-predict-using-pre-trained-model/input/HMDB.msp",
...     model_path=r"/your/path/to/your-model-name_model.pkl",
...     output_file=r"/example_data/tab1-correct-the-adduct/example1-predict-using-pre-trained-model/output"
... )
train_another_model(training, testing, modelName, outpath, log_fn=None, should_stop=None)

Train a custom STRIKER model for adduct correction using labeled training data.

This function is intended for users familiar with machine learning who wish to train or refine adduct correction models using CSV files of input and standardized adducts.

Parameters:
  • training (str) – Path to a CSV file containing two columns: - ‘input’: Variants of ion notations for the same adduct. - ‘corrected’: Standardized adduct representation.

  • testing (str) – Path to a CSV file used to evaluate the model’s performance on unseen adduct notations.

  • modelName (str) – Name of the model for later selection in the “Correct using pre-trained model” tab.

  • outpath (str) – Directory where the trained model, vectorizer, and reports will be saved.

Returns:

The function writes the following files to outpath:
  • {modelName}_model.pkl: Serialized trained model.

  • {modelName}_vectorizer.pkl: Serialized vectorizer.

  • {modelName}_testing_results.csv: Predictions on the testing set.

  • {modelName}_accuracy.txt: Accuracy score of the model on testing data.

  • {modelName}_classification_report.txt: Full classification report.

Return type:

Files

Examples

>>> from striker import correct_adduct
>>> correct_adduct.train_another_model(
...     training=r"/example_data/tab1-correct-the-adduct/example2-build-your-own-model/input/mlp_3_training_set_with_unknown.csv",
...     testing=r"/example_data/tab1-correct-the-adduct/example2-build-your-own-model/input/mlp_4_testing_set.csv",
...     modelName='MyCustomModel',
...     outpath=r"/example_data/tab1-correct-the-adduct/example2-build-your-own-model/output"
... )
predict_the_adduct(qLibrary, sLibrary, distanceType, selectedCriteria, the_tolerance, the_threshold, selectedMatchValues, selectedTestAccuracyResponse, outpath, log_fn=None, should_stop='')
When to Use?

After adduct correction, spectra are categorized as corrected, excluded, or unknown. This task evaluates the most appropriate similarity method for your library.

Required

Identifying the optimal similarity method relies on the presence of known adducts; therefore, ensure your library includes them.

Recommendation

Before using this function, apply a metadata normalization tool to ensure consistency in the spectral library. We highly recommend using the FragHub library, which integrates normalized spectra from multiple sources: https://zenodo.org/records/15356456

Parameters:
  • qLibrary (str) – Path to the query spectral library (.msp, .mgf, or .json). Spectra may contain unknown or known adducts depending on the mode.

  • sLibrary (str) – Path to the subject spectral library (.msp, .mgf, or .json). Must contain known adduct annotations in its metadata.

  • distanceType ({"OSA", "Entropy", "CosineGreedy"}) – Similarity metric used to compare spectra. - “OSA” : Numba-accelerated dynamic programming alignment. - “Entropy” : Shannon-entropy–based similarity (ms_entropy). - “CosineGreedy” : MatchMS cosine-greedy similarity.

  • selectedCriteria ({"InChIKey", "SMILES", "Molecular formula", "None"}) – Metadata key used for filtering candidate spectra in the subject library.

  • selectedMatchValues (bool) – Whether to use real intensities (True) or convert all intensities to 1 (False).

  • selectedTestAccuracyResponse (bool) – If True, evaluate prediction accuracy for spectra with known adducts in the query library. If False, predict adducts for unknown spectra and export updated spectral files.

  • outpath (str) – Directory where exported results, accuracy reports, predicted spectra, and missing spectra will be saved.

Returns:

The function writes the following output files to outpath:
  • Accuracy mode: *_accuracy.txt, *_match_report.csv,

*_match_report_summary.txt.

  • Prediction mode: predicted spectra file, missing spectra file, summary files.

Return type:

Files

Examples

>>> from striker import predict_adduct
>>> predict_the_adduct(
...     qLibrary=r"/example_data/tab2-predict-the-adduct/input/query-library/HMDB_Corrected_Library_STRIKER.msp",
...     sLibrary=r"/example_data/tab2-predict-the-adduct/input/subject-library/FragHub-NEG_LC.msp",
...     distanceType="Entropy",
...     selectedCriteria="InChIKey",
...     selectedMatchValues=True,
...     selectedTestAccuracyResponse=True,
...     outpath=r"/example_data/tab2-predict-the-adduct/output"
... )
parse_xml(xml_path, csv_file, cwd, log_fn=None, should_stop=None)

Build the HMDB spectral library from downloaded MS/MS XML files.

This function parses HMDB XML spectral files, merges metadata from a CSV file, and exports combined data into MSP, JSON, and summary reports. It is designed for experimental MS/MS spectra.

When to Use?

Use this option when you want to construct a comprehensive HMDB spectral library from the downloaded “MS-MS Spectra Files (XML)” and merge it with curated metadata.

Parameters:
  • xml_path (str) – Path to the directory containing HMDB XML files.

  • csv_file (str) – Path to a CSV file containing HMDB metadata to be merged with XML entries.

  • cwd (str) – Output directory where MSP, JSON, and summary files will be written.

Returns:

The function processes all .xml files in the given directory and produces:
  • STRIKER-HMDB.msp: MSP-formatted spectral library.

  • STRIKER-HMDB.json: JSON list of MSP entries.

  • STRIKER-missing-spectra.txt: List of XML files without usable spectra.

  • STRIKER-summary.txt: Summary statistics of processed spectra.

Return type:

Files

Examples

Process all XML files in a folder:

>>> from striker import build_hmdb_library
>>> parse_xml(
...     xml_path=r"/example_data/tab3-create-hmdb-library/input/hmdb_xml_files/",
...     csv_file=r"/example_data/tab3-create-hmdb-library/input/metabolites-2024-10-03.csv",
...     cwd=r"/example_data/tab3-create-hmdb-library/output/"
... )
subset_a_library(input_file: str, selected_field: str, output_path: str, log_fn=None, values_file: str = None, tolerance: float | None = None, range_vals: tuple | None = None, should_stop=None)

Subset a spectral library based on metadata fields, tolerance-based matching, or numeric ranges.

When to use?

Use this feature when you have a comprehensive spectral library and wish to extract a specific sub-library based on metadata criteria. This is particularly useful for filtering spectra by attributes such as ionization mode (e.g., positive or negative mode) or sample matrix (e.g., blood, urine). For example, you can extract all spectra associated with positive ionization mode, or retrieve entries that correspond to a predefined list of InChIKeys related to a specific biological matrix.

This function loads spectra from an input library and filters them based on one of three modes:

  1. Direct value matching.

  2. Numeric range filtering.

  3. Value matching with ± tolerance.

Supported metadata keys include:

INCHIKEY, SMILES, NAME, PRECURSORTYPE, PRECURSORMZ, EXACTMASS, IONMODE, FORMULA.

Parameters:
  • input_file (str) – Path to the input spectral library (.msp, .mgf, or .json).

  • selected_field (str) – Metadata field selection. May contain: - “KEY” (direct match) - “KEY with Tolerance” - “KEY Range”

  • output_path (str) – Directory where the filtered library will be saved.

  • log_fn (callable, optional) – Logger function. Defaults to print().

  • values_file (str, optional) – Path to text file containing metadata values used for filtering. Required for: - direct matching - tolerance-based matching

  • tolerance (float, optional) – Allowed deviation from reference values (± tolerance). Required when “with Tolerance” is used.

  • range_vals (tuple of float, optional) – Minimum and maximum values for range-based filtering. Required when “Range” is used.

Returns:

Writes:
  • filtered spectral library file

  • summary text file containing extracted value counts

Return type:

Files

Examples

Direct match:

>>> from striker import subset_library
>>> subset_a_library(input_file=r"example_data/tab4-extract-sub-library/input/library/STRIKER-HMDB.msp",
    ...                                      selected_field="INCHIKEY",
...                  output_path=r"example_data/tab4-extract-sub-library/output/",
...                  values_file=r"example_data/tab4-extract-sub-library/input/metadata value files/test3-inchikeys.txt"
    ...)

Range filter:

>>> from striker import subset_library
>>> subset_a_library(input_file="example_data/tab4-extract-sub-library/input/library/STRIKER-HMDB.msp",
    ...                                      selected_field="PRECURSORMZ Range",
...                  output_path="",
...                  range_vals=(100, 500)
    ...)

Tolerance filtering:

>>> from striker import subset_library
>>> subset_a_library(input_file="example_data/tab4-extract-sub-library/input/library/STRIKER-HMDB.msp",
    ...                                      selected_field="EXACTMASS with Tolerance",
...                  output_path="example_data/tab4-extract-sub-library/input/library/",
...                  values_file=r"example_data/tab4-extract-sub-library/input/metadata value files/test3-test5-exactmass.txt",
...                  tolerance=0.01
    ...)