Module Reference¶
Dataset¶
- asteroid_impact.dataset.download_dataset(output_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/raw/dataset.csv'), sample_size: int = 500, force_download: bool = False) DataFrame¶
Download a sample of the asteroid impact-risk dataset.
The dataset is downloaded from Hugging Face and stored locally as a CSV file. If the file already exists, the local copy is reused unless
force_downloadis set toTrue.Parameters¶
- output_path:
Path where the downloaded dataset will be saved.
- sample_size:
Maximum number of rows to keep in the local sample.
- force_download:
Whether to download the dataset again if a local copy exists.
Returns¶
- pandas.DataFrame
The downloaded dataset.
- asteroid_impact.dataset.main(force_download: bool = <typer.models.OptionInfo object>)¶
Downloads and prepares the asteroid impact-risk dataset.
- asteroid_impact.dataset.prepare_dataset(input_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/raw/dataset.csv'), output_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/processed/dataset.csv')) DataFrame¶
Clean the raw dataset and save the processed version.
Rows containing missing values are removed before the processed dataset is written to disk.
Parameters¶
- input_path:
Path to the raw dataset.
- output_path:
Path where the cleaned dataset will be saved.
Returns¶
- pandas.DataFrame
The cleaned dataset.
Features¶
Dataset acquisition and preprocessing utilities.
This module downloads the source dataset from Hugging Face, stores a local sample, and cleans the data before it is used in feature engineering and model training.
- asteroid_impact.features.main(force_download: bool = <typer.models.OptionInfo object>)¶
Download and prepare the asteroid impact-risk dataset.
Parameters¶
- force_downloadbool, default=False
If True, redownload the source dataset even when a local copy already exists.
Returns¶
- None
This command runs the data acquisition and preprocessing pipeline.
- asteroid_impact.features.prepare_dataset(input_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/raw/dataset.csv'), output_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/processed/dataset.csv')) DataFrame¶
Clean the raw dataset and save the processed version.
Rows containing missing values are removed before the processed dataset is written to disk.
Parameters¶
- input_pathPath, default=RAW_DATA_PATH
Path to the raw dataset.
- output_pathPath, default=PROCESSED_DATA_PATH
Path where the cleaned dataset will be saved.
Returns¶
- pandas.DataFrame
The cleaned dataset.
Configuration¶
Configuration settings for the asteroid impact-risk project.
This module defines the project root and the locations used for raw data, processed data, models, and generated reports. It also ensures the required folders exist when the package is imported.
Plots¶
Feature engineering utilities for the asteroid impact-risk model.
The module creates the feature matrix and transformed target variable used to train the regression model.
- asteroid_impact.plots.create_features(input_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/processed/dataset.csv'), features_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/processed/features.csv'), labels_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/processed/labels.csv')) tuple[DataFrame, Series]¶
Create model features and the transformed target variable.
Missing rows are removed and the impact probability is transformed using a base-10 logarithm. Numerical columns are used as model features.
Parameters¶
- input_pathPath, default=DATASET_PATH
Path to the processed dataset.
- features_pathPath, default=FEATURES_PATH
Path where the feature data will be saved.
- labels_pathPath, default=LABELS_PATH
Path where the target labels will be saved.
Returns¶
- tuple[pandas.DataFrame, pandas.Series]
The feature matrix and transformed target variable.
Model training¶
Model training pipeline for the impact-probability regressor.
- asteroid_impact.modeling.train.build_model(input_shape: int, learning_rate: float = 0.001) Model¶
- asteroid_impact.modeling.train.main()¶
- asteroid_impact.modeling.train.train_model(features_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/processed/features.csv'), labels_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/processed/labels.csv'), model_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/models/impact_probability_model.keras'), scaler_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/models/feature_scaler.joblib'), metrics_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/reports/model_metrics.csv'), history_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/reports/training_history.csv'), epochs: int = 100, batch_size: int = 32, learning_rate: float = 0.001, sample_size: int = 0, progress_callback: Callable[[float, str], Any] | None = None) dict[str, Any]¶
Model prediction¶
Prediction pipeline for the trained impact-probability model.
This module loads the saved model and scaler, transforms new feature values, and writes impact probability predictions to disk.
- asteroid_impact.modeling.predict.main()¶
Generate predictions for the processed features.
Returns¶
- None
Executes the prediction pipeline.
- asteroid_impact.modeling.predict.predict(features_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/processed/features.csv'), model_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/models/impact_probability_model.keras'), scaler_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/models/feature_scaler.joblib'), predictions_path: Path = PosixPath('/home/runner/work/asteroid_impact/asteroid_impact/data/processed/predictions.csv')) DataFrame¶
Generate impact probability predictions.
The trained model and feature scaler are loaded from disk. The input features are scaled before being passed to the model.
Parameters¶
- features_pathPath, default=FEATURES_PATH
Path to the processed feature data.
- model_pathPath, default=MODEL_PATH
Path to the trained Keras model.
- scaler_pathPath, default=SCALER_PATH
Path to the saved feature scaler.
- predictions_pathPath, default=PREDICTIONS_PATH
Path where predictions will be saved.
Returns¶
- pandas.DataFrame
A dataframe containing logarithmic and original-scale predictions.
Interactive application¶
- asteroid_impact.app.create_interface() Blocks¶
Create the main Gradio interface with all tabs.
- asteroid_impact.app.launch_app()¶
Interactive UI¶
Data Exploration tab for the Gradio application.
- asteroid_impact.ui.data_exploration.create_data_exploration_tab() None¶
Create the Data Exploration tab with interactive visualizations.
- asteroid_impact.ui.data_exploration.generate_interactive_plot(df, plot_type, x_col, y_col, bins)¶
Generate interactive plot based on selection.
- asteroid_impact.ui.data_exploration.plot_correlations() Figure¶
Plot correlation heatmap of numeric features.
- asteroid_impact.ui.data_exploration.plot_distribution() Figure¶
Plot impact probability distribution.
- asteroid_impact.ui.data_exploration.plot_feature_distributions() Figure¶
Plot distributions of all features.
Training Interface tab for the Gradio application.
- asteroid_impact.ui.training.create_training_tab() None¶
Create the Training tab.
Model Evaluation tab for the Gradio application.