Project Workflow¶
Dataset download¶
The project downloads the dataset from Hugging Face using the following source:
DATASET_URI = (
"hf://datasets/juliensimon/sentry-impact-risk/"
"data/sentry_impact_risk.parquet"
)
A sample of up to 500 rows is selected and stored in:
data/raw/dataset.csv
The dataset is then cleaned by removing rows containing missing values. The cleaned dataset is stored in:
data/processed/dataset.csv
Feature engineering¶
The features.py module reads the processed dataset and removes missing
values if necessary.
The target variable is transformed using a base-10 logarithm:
All numerical columns are selected as model features except:
impact_probabilitylog_impact_probability
The resulting files are:
features.csv: model input variables.labels.csv: transformed target variable.
Model architecture¶
The neural network is defined in modeling/train.py.
It contains:
An input layer.
A dense layer with 64 neurons and ReLU activation.
A dense layer with 32 neurons and ReLU activation.
A dense layer with 16 neurons and ReLU activation.
An output layer with one neuron.
The model uses:
The Adam optimizer.
A learning rate of
0.001.Mean squared error as the loss function.
Mean absolute error as a metric.
Data splitting and scaling¶
The dataset is divided into training and test data using an 80/20 split.
The features are standardized with StandardScaler before training:
scaler = StandardScaler()
The scaler is saved to:
models/feature_scaler.joblib
Training¶
The model is trained for a maximum of 100 epochs with a batch size of 32.
Early stopping is used to reduce overfitting. Training stops when the validation loss does not improve for 15 consecutive epochs.
Evaluation¶
The model is evaluated using:
MAE: Mean Absolute Error.
RMSE: Root Mean Squared Error.
R²: Coefficient of Determination.
The results are saved to:
reports/model_metrics.csv
The training history is saved to:
reports/training_history.csv
Prediction¶
The prediction module loads:
The trained Keras model.
The saved feature scaler.
The processed feature data.
The model produces a prediction in logarithmic form. The original scale is then recovered using: