Module Reference

Dataset

asteroid_impact.dataset.download_dataset(output_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/raw/dataset.csv'), sample_size: int = 500, force_download: bool = False) → DataFrame

Download a sample of the asteroid impact-risk dataset.

The dataset is downloaded from Hugging Face and stored locally as a CSV file. If the file already exists, the local copy is reused unless force_download is set to True.

Parameters

output_path:

Path where the downloaded dataset will be saved.

sample_size:

Maximum number of rows to keep in the local sample.

force_download:

Whether to download the dataset again if a local copy exists.

Returns

pandas.DataFrame

The downloaded dataset.

asteroid_impact.dataset.main(force_download: bool = <typer.models.OptionInfo object>)

Downloads and prepares the asteroid impact-risk dataset.

asteroid_impact.dataset.prepare_dataset(input_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/raw/dataset.csv'), output_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/processed/dataset.csv')) → DataFrame

Clean the raw dataset and save the processed version.

Rows containing missing values are removed before the processed dataset is written to disk.

Parameters

input_path:

Path to the raw dataset.

output_path:

Path where the cleaned dataset will be saved.

Returns

pandas.DataFrame

The cleaned dataset.

Features

Dataset acquisition and preprocessing utilities.

This module downloads the source dataset from Hugging Face, stores a local sample, and cleans the data before it is used in feature engineering and model training.

asteroid_impact.features.main(force_download: bool = <typer.models.OptionInfo object>)

Download and prepare the asteroid impact-risk dataset.

Parameters

force_downloadbool, default=False

If True, redownload the source dataset even when a local copy already exists.

Returns

None

This command runs the data acquisition and preprocessing pipeline.

asteroid_impact.features.prepare_dataset(input_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/raw/dataset.csv'), output_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/processed/dataset.csv')) → DataFrame

Clean the raw dataset and save the processed version.

Rows containing missing values are removed before the processed dataset is written to disk.

Parameters

input_pathPath, default=RAW_DATA_PATH

Path to the raw dataset.

output_pathPath, default=PROCESSED_DATA_PATH

Path where the cleaned dataset will be saved.

Returns

pandas.DataFrame

The cleaned dataset.

Configuration

Configuration settings for the asteroid impact-risk project.

This module defines the project root and the locations used for raw data, processed data, models, and generated reports. It also ensures the required folders exist when the package is imported.

Plots

Feature engineering utilities for the asteroid impact-risk model.

The module creates the feature matrix and transformed target variable used to train the regression model.

asteroid_impact.plots.create_features(input_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/processed/dataset.csv'), features_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/processed/features.csv'), labels_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/processed/labels.csv')) → tuple[DataFrame, Series]

Create model features and the transformed target variable.

Missing rows are removed and the impact probability is transformed using a base-10 logarithm. Numerical columns are used as model features.

Parameters

input_pathPath, default=DATASET_PATH

Path to the processed dataset.

features_pathPath, default=FEATURES_PATH

Path where the feature data will be saved.

labels_pathPath, default=LABELS_PATH

Path where the target labels will be saved.

Returns

tuple[pandas.DataFrame, pandas.Series]

The feature matrix and transformed target variable.

asteroid_impact.plots.main()

Generate the feature and label files.

Returns

None

Executes the feature creation pipeline.

Model training

Model training pipeline for the impact-probability regressor.

asteroid_impact.modeling.train.build_model(input_shape: int, learning_rate: float = 0.001) → Model
asteroid_impact.modeling.train.main()
asteroid_impact.modeling.train.train_model(features_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/processed/features.csv'), labels_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/processed/labels.csv'), model_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/models/impact_probability_model.keras'), scaler_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/models/feature_scaler.joblib'), metrics_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/reports/model_metrics.csv'), history_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/reports/training_history.csv'), epochs: int = 100, batch_size: int = 32, learning_rate: float = 0.001, sample_size: int = 0, progress_callback: Callable[[float, str], Any] | None = None) → dict[str, Any]

Model prediction

Prediction pipeline for the trained impact-probability model.

This module loads the saved model and scaler, transforms new feature values, and writes impact probability predictions to disk.

asteroid_impact.modeling.predict.main()

Generate predictions for the processed features.

Returns

None

Executes the prediction pipeline.

asteroid_impact.modeling.predict.predict(features_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/processed/features.csv'), model_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/models/impact_probability_model.keras'), scaler_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/models/feature_scaler.joblib'), predictions_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/processed/predictions.csv')) → DataFrame

Generate impact probability predictions.

The trained model and feature scaler are loaded from disk. The input features are scaled before being passed to the model.

Parameters

features_pathPath, default=FEATURES_PATH

Path to the processed feature data.

model_pathPath, default=MODEL_PATH

Path to the trained Keras model.

scaler_pathPath, default=SCALER_PATH

Path to the saved feature scaler.

predictions_pathPath, default=PREDICTIONS_PATH

Path where predictions will be saved.

Returns

pandas.DataFrame

A dataframe containing logarithmic and original-scale predictions.

Interactive application

asteroid_impact.app.create_interface() → Blocks

Create the main Gradio interface with all tabs.

asteroid_impact.app.launch_app()

Interactive UI

Data Exploration tab for the Gradio application.

asteroid_impact.ui.data_exploration.create_data_exploration_tab() → None

Create the Data Exploration tab with interactive visualizations.

asteroid_impact.ui.data_exploration.generate_interactive_plot(df, plot_type, x_col, y_col, bins)

Generate interactive plot based on selection.

asteroid_impact.ui.data_exploration.plot_correlations() → Figure

Plot correlation heatmap of numeric features.

asteroid_impact.ui.data_exploration.plot_distribution() → Figure

Plot impact probability distribution.

asteroid_impact.ui.data_exploration.plot_feature_distributions() → Figure

Plot distributions of all features.

Training Interface tab for the Gradio application.

asteroid_impact.ui.training.create_training_tab() → None

Create the Training tab.

Model Evaluation tab for the Gradio application.