Usage¶
The project should be executed in the following order:
Download and prepare the dataset.
Create the features and labels.
Train and evaluate the model.
Generate predictions.
Download and prepare the dataset¶
The dataset module downloads the dataset from Hugging Face, creates a
local sample, removes incomplete rows, and saves the cleaned dataset.
Run:
python -m asteroid_impact.dataset
By default, the project downloads a sample of up to 500 rows.
To download the dataset again, use:
python -m asteroid_impact.dataset --force-download
The generated files are:
data/raw/dataset.csvdata/processed/dataset.csv
Create features¶
The features module creates the input features and the target labels.
Run:
python -m asteroid_impact.features
This command creates:
data/processed/features.csvdata/processed/labels.csv
Train the model¶
The train module trains and evaluates the neural network.
Run:
python -m asteroid_impact.modeling.train
This command creates:
models/impact_probability_model.kerasmodels/feature_scaler.joblibreports/model_metrics.csvreports/training_history.csv
Generate predictions¶
The predict module loads the trained model and generates predictions.
Run:
python -m asteroid_impact.modeling.predict
The predictions are saved as a CSV file in the processed data directory.