More about forecasting in cienciadedatos.net
- ARIMA and SARIMAX models with python
- Time series forecasting with machine learning
- Forecasting time series with gradient boosting: XGBoost, LightGBM and CatBoost
- Forecasting time series with XGBoost
- Global Forecasting Models: Multi-series forecasting
- Global Forecasting Models: Comparative Analysis of Single and Multi-Series Forecasting Modeling
- Probabilistic forecasting
- Forecasting with deep learning
- Forecasting energy demand with machine learning
- Forecasting web traffic with machine learning
- Intermittent demand forecasting
- Modelling time series trend with tree-based models
- Bitcoin price prediction with Python
- Stacking ensemble of machine learning models to improve forecasting
- Interpretable forecasting models
- Mitigating the Impact of Covid on forecasting Models
- Forecasting time series with missing values
Introduction¶
Machine learning interpretability, also known as explainability, refers to the ability to understand, interpret, and explain the decisions or predictions made by machine learning models in a human-understandable way. It aims to shed light on how a model arrives at a particular result or decision.
Due to the complex nature of many modern machine learning models, such as ensemble methods, they often function as black boxes, making it difficult to understand why a particular prediction was made. Explainability techniques aim to demystify these models, providing insight into their inner workings and helping to build trust, improve transparency, and meet regulatory requirements in various domains. Improving model explainability not only helps to understand model behavior, but also helps to identify biases, improve model performance, and enable stakeholders to make more informed decisions based on machine learning insights.
The skforecast library is compatible with some of the most used interpretability methods: SHAP values, partial dependence plots, and model-specific methods.
Libraries¶
Libraries used in this document.
# Data manipulation
# ==============================================================================
import pandas as pd
import numpy as np
from skforecast.datasets import fetch_dataset
# Plotting
# ==============================================================================
import matplotlib.pyplot as plt
import shap
from skforecast.plot import set_dark_theme
# Modeling and forecasting
# ==============================================================================
import sklearn
import lightgbm
import skforecast
from sklearn.inspection import PartialDependenceDisplay
from lightgbm import LGBMRegressor
from skforecast.recursive import ForecasterRecursive
from skforecast.preprocessing import RollingFeatures
from skforecast.model_selection import backtesting_forecaster, TimeSeriesFold
color = '\033[1m\033[38;5;208m'
print(f"{color}Version skforecast: {skforecast.__version__}")
print(f"{color}Version scikit-learn: {sklearn.__version__}")
print(f"{color}Version lightgbm: {lightgbm.__version__}")
print(f"{color}Version shap: {shap.__version__}")
Version skforecast: 0.25.0 Version scikit-learn: 1.7.2 Version lightgbm: 4.7.0 Version shap: 0.52.0
Data¶
# Download data
# ==============================================================================
data = fetch_dataset(name="vic_electricity")
data.head(3)
╭──────────────────────────── vic_electricity ─────────────────────────────╮ │ Description: │ │ Half-hourly electricity demand for Victoria, Australia │ │ │ │ Source: │ │ O'Hara-Wild M, Hyndman R, Wang E, Godahewa R (2022).tsibbledata: Diverse │ │ Datasets for 'tsibble'. https://tsibbledata.tidyverts.org/, │ │ https://github.com/tidyverts/tsibbledata/. │ │ https://tsibbledata.tidyverts.org/reference/vic_elec.html │ │ │ │ URL: │ │ https://raw.githubusercontent.com/skforecast/skforecast- │ │ datasets/main/data/vic_electricity.csv │ │ │ │ Shape: 52608 rows x 4 columns │ ╰──────────────────────────────────────────────────────────────────────────╯
| Demand | Temperature | Date | Holiday | |
|---|---|---|---|---|
| Time | ||||
| 2011-12-31 13:00:00 | 4382.825174 | 21.40 | 2012-01-01 | True |
| 2011-12-31 13:30:00 | 4263.365526 | 21.05 | 2012-01-01 | True |
| 2011-12-31 14:00:00 | 4048.966046 | 20.70 | 2012-01-01 | True |
# Aggregation to daily frequency
# ==============================================================================
data = data.resample('D').agg({'Demand': 'sum', 'Temperature': 'mean'})
data.head(3)
| Demand | Temperature | |
|---|---|---|
| Time | ||
| 2011-12-31 | 82531.745918 | 21.047727 |
| 2012-01-01 | 227778.257304 | 26.578125 |
| 2012-01-02 | 275490.988882 | 31.751042 |
# Create calendar variables
# ==============================================================================
data['day_of_week'] = data.index.dayofweek
data['month'] = data.index.month
data.head(3)
| Demand | Temperature | day_of_week | month | |
|---|---|---|---|---|
| Time | ||||
| 2011-12-31 | 82531.745918 | 21.047727 | 5 | 12 |
| 2012-01-01 | 227778.257304 | 26.578125 | 6 | 1 |
| 2012-01-02 | 275490.988882 | 31.751042 | 0 | 1 |
# Split train-test
# ==============================================================================
end_train = '2014-12-01 23:59:00'
data_train = data.loc[: end_train, :]
data_test = data.loc[end_train:, :]
print(f"Dates train : {data_train.index.min()} --- {data_train.index.max()} (n={len(data_train)})")
print(f"Dates test : {data_test.index.min()} --- {data_test.index.max()} (n={len(data_test)})")
Dates train : 2011-12-31 00:00:00 --- 2014-12-01 00:00:00 (n=1067) Dates test : 2014-12-02 00:00:00 --- 2014-12-31 00:00:00 (n=30)
Forecasting model¶
A forecasting model is created to predict the energy demand using the past 7 values (last week), a rolling mean over the last 24 days as a window feature, and the temperature together with two calendar features (day of week and month) as exogenous variables.
# Create a recursive multi-step forecaster (ForecasterRecursive)
# ==============================================================================
window_features = RollingFeatures(stats=['mean'], window_sizes=24)
exog_features = ['Temperature', 'day_of_week', 'month']
forecaster = ForecasterRecursive(
estimator = LGBMRegressor(random_state=123, verbose=-1),
lags = 7,
window_features = window_features
)
forecaster.fit(
y = data_train['Demand'],
exog = data_train[exog_features],
)
forecaster
ForecasterRecursive
General Information
- Estimator: LGBMRegressor
- Lags: [1 2 3 4 5 6 7]
- Window features: ['roll_mean_24']
- Calendar features: None
- Window size: 24
- Series name: Demand
- Exogenous included: True
- Categorical features: auto
- Weight function included: False
- Differentiation order: None
- Drop NaN from series: False
- Creation date: 2026-09-24 13:16:32
- Last fit date: 2026-09-24 13:16:36
- Skforecast version: 0.25.0
- Python version: 3.13.14
- Forecaster id: None
Exogenous Variables
Temperature, day_of_week, month
Data Transformations
- Transformer for y: None
- Transformer for exog: None
Training Information
- Training range: [Timestamp('2011-12-31 00:00:00'), Timestamp('2014-12-01 00:00:00')]
- Training index type: DatetimeIndex
- Training index frequency: D
Estimator Parameters
-
{'boosting_type': 'gbdt', 'class_weight': None, 'colsample_bytree': 1.0, 'importance_type': 'split', 'learning_rate': 0.1, 'max_depth': -1, 'min_child_samples': 20, 'min_child_weight': 0.001, 'min_split_gain': 0.0, 'n_estimators': 100, 'n_jobs': None, 'num_leaves': 31, 'objective': None, 'random_state': 123, 'reg_alpha': 0.0, 'reg_lambda': 0.0, 'subsample': 1.0, 'subsample_for_bin': 200000, 'subsample_freq': 0, 'verbose': -1}
Fit Kwargs
-
{}
Model-specific feature importances¶
Feature importance in machine learning determines the relevance or importance of each feature (or variable) in a model's prediction. In other words, it measures how much each feature contributes to the model's output.
Feature importance can be used for several purposes, such as identifying the most relevant features for a given prediction, understanding the behavior of a model, and selecting the best set of features for a given task. It can also help to identify potential biases or errors in the data used to train the model. It is important to note that feature importance is not a definitive measure of causality. Just because a feature is identified as important does not necessarily mean that it caused the outcome. Other factors, such as confounding variables, may also be at play.
The method used to calculate feature importance may vary depending on the type of machine learning model used. Different models can have different assumptions and characteristics that affect the importance calculation. For example, decision tree-based models, such as Random Forest and Gradient Boosting, typically use methods that measure the reduction of impurities or the effect of permutations. Linear regression models typically use coefficients. The magnitude of the coefficient reflects the strength and direction of the relationship between the predictor and the target variable.
The importance of the predictors included in a forecaster can be obtained using the method get_feature_importances(). This method accesses the coef_ and feature_importances_ attributes of the internal estimator.
⚠ Warning
Theget_feature_importances() method will only return values if the forecaster's estimator has either the coef_ or feature_importances_ attribute, which is the default in scikit-learn.
# Extract feature importance
# ==============================================================================
importance = forecaster.get_feature_importances()
importance
| feature | importance | |
|---|---|---|
| 8 | Temperature | 568 |
| 0 | lag_1 | 426 |
| 1 | lag_2 | 291 |
| 6 | lag_7 | 248 |
| 4 | lag_5 | 243 |
| 2 | lag_3 | 236 |
| 7 | roll_mean_24 | 227 |
| 10 | month | 224 |
| 5 | lag_6 | 200 |
| 3 | lag_4 | 177 |
| 9 | day_of_week | 160 |
According to this ranking, Temperature is the most relevant feature, followed by lag_1 and lag_2. These values should be read with caution: for LGBMRegressor, the feature_importances_ attribute counts by default the number of times a feature is used to split the data (importance_type='split'), not the improvement in accuracy that those splits produce. A feature can participate in many splits with little effect, or in a few very influential ones. Setting importance_type='gain' when instantiating the model, or using the SHAP values presented in the next section, usually gives a more faithful picture of each feature's real contribution.
SHAP explanations for skforecast models¶
SHAP (SHapley Additive exPlanations) values are a widely adopted method for explaining machine learning models. They provide both visual and quantitative insights into how features and their values impact the model. SHAP values serve two primary purposes:
Global Interpretability: SHAP values help identify how each feature influenced the model during training. By averaging SHAP values across the dataset, one can rank features by their overall importance and gain insight into the model’s decision-making process.
Local Interpretability: SHAP values also explain individual predictions by indicating how much each feature contributed to a specific output. This enables a breakdown of single predictions to understand the role each feature played in the outcome.
SHAP value explanations can be generated for skforecast models using two essential components:
The internal estimator of the forecaster, accessible via
forecaster.estimator.The internal matrices used for fitting, backtesting, and predicting with the forecaster. These matrices are accessible through the method
create_predict_X()and by setting the argumentreturn_predictors=Truein thebacktesting_forecaster()function.
By leveraging these elements, users can produce clear and interpretable explanations for their forecasting models. These explanations can be used to assess model reliability, identify the most influential features, and better understand the relationships between input variables and the target variable.
⚠ Warning
SHAP values and partial dependence plots quantify the association between the features and the predictions of the model, they do not measure causality. Furthermore, when these methods perturb or combine feature values, they can generate scenarios that are unrealistic for autocorrelated data such as time series: the lags of a series are strongly correlated with each other (for example, a very highlag_1 is unlikely to occur together with a very low roll_mean_24). The explanations should therefore be interpreted as a description of how the model uses the features, not as properties of the underlying process.
SHAP feature importance in the overall model¶
Averaging the SHAP values across the data set used to train the model, it is possible to obtain an estimation of the contribution (magnitude and direction) of each feature in the model. The higher the absolute value of the SHAP value, the more important the feature is for the model. The sign of the SHAP value indicates whether the feature has a positive or negative impact on the prediction.
First, the training matrices used to fit the model are created with the method create_train_X_y().
# Training matrices used by the forecaster to fit the internal estimator
# ==============================================================================
X_train, y_train = forecaster.create_train_X_y(
y = data_train['Demand'],
exog = data_train[exog_features],
)
display(X_train.head(3)) # Features
display(y_train.head(3)) # Target
| lag_1 | lag_2 | lag_3 | lag_4 | lag_5 | lag_6 | lag_7 | roll_mean_24 | Temperature | day_of_week | month | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Time | |||||||||||
| 2012-01-24 | 280188.298774 | 239810.374218 | 207949.859910 | 225035.325476 | 240187.677944 | 247722.494256 | 292458.685446 | 222658.202570 | 26.611458 | 1 | 1 |
| 2012-01-25 | 287474.816646 | 280188.298774 | 239810.374218 | 207949.859910 | 225035.325476 | 240187.677944 | 247722.494256 | 231197.497184 | 19.759375 | 2 | 1 |
| 2012-01-26 | 239083.684380 | 287474.816646 | 280188.298774 | 239810.374218 | 207949.859910 | 225035.325476 | 240187.677944 | 231668.556646 | 20.038542 | 3 | 1 |
Time 2012-01-24 287474.816646 2012-01-25 239083.684380 2012-01-26 214239.304588 Freq: D, Name: y, dtype: float64
✎ Note
Although the data starts on 2011-12-31, the training matrix begins on 2012-01-24. Since the forecaster includes a rolling mean over a 24-day window (roll_mean_24), its window size is 24: the first 24 observations of the series are needed to compute the predictors of the first training row, so they appear as features but not as target values. This is why X_train and y_train contain 24 fewer rows than the original training series.
Then, the SHAP values are calculated with the shap library. Calling the explainer with the training data returns an Explanation object that stores the SHAP values. If the data set is large, it is recommended to use only a random sample.
# Create SHAP explainer (tree-based model)
# ==============================================================================
explainer = shap.TreeExplainer(forecaster.estimator)
# Sample 50% of the data to speed up the calculation
rng = np.random.default_rng(seed=785412)
sample = rng.choice(X_train.index, size=int(len(X_train)*0.5), replace=False)
X_train_sample = X_train.loc[sample, :]
y_train_sample = y_train.loc[sample]
explanation = explainer(X=X_train_sample, y=y_train_sample)
shap_values = explanation.values
✎ Note
The SHAP library has several explainers, each designed for a different type of model. Theshap.TreeExplainer explainer is used for tree-based models, such as the LGBMRegressor used in this example. For more information, see the SHAP documentation.
Once the SHAP values are calculated, several plots can be generated to visualize the results.
SHAP Summary Plot¶
The SHAP summary plot typically displays the feature importance or contribution of each feature to the model's output across multiple data points. It shows how much each feature contributes to pushing the model's prediction away from a base value (often the model's average prediction). By examining a SHAP summary plot, one can gain insights into which features have the most significant impact on predictions, whether they positively or negatively influence the outcome, and how different feature values contribute to specific predictions.
# Shap summary plot (top 10)
# ==============================================================================
shap.initjs()
shap.summary_plot(shap_values, X_train_sample, max_display=10, show=False)
fig, ax = plt.gcf(), plt.gca()
ax.set_title('SHAP summary plot')
ax.tick_params(labelsize=8)
fig.set_size_inches(6, 3)
In this plot, each point is one observation of the training sample. Its position on the horizontal axis is its SHAP value (impact on the prediction) and its color indicates whether the feature value is high (pink) or low (blue) in that observation. Several patterns emerge:
High values of
lag_1(demand of the previous day) push the prediction up and low values push it down: demand is strongly persistent from one day to the next.High values of
day_of_week(weekends) reduce the predicted demand, consistent with the lower electricity consumption of Saturdays and Sundays.Temperatureshows large positive SHAP values at its highest values (cooling demand), while intermediate values have a negative impact.
# Shap summary plot (bar)
# ==============================================================================
shap.summary_plot(shap_values, X_train_sample, plot_type='bar', plot_size=(6, 3))
The bar plot summarizes the mean absolute SHAP value of each feature: lag_1, day_of_week and Temperature dominate the model output. It is worth comparing this ranking with the one returned by get_feature_importances(): day_of_week is last according to the number of splits, but second according to SHAP. This means that the model uses this feature in relatively few splits, but those splits change the predictions considerably. When the two rankings disagree, SHAP values (or the gain of the splits) are usually a more reliable measure of each feature's actual influence.
SHAP Dependence Plots¶
SHAP dependence plots are visualizations used to understand the relationship between a feature and the model output by displaying how the value of a single feature affects predictions made by the model while considering interactions with other features. These plots are particularly useful for examining how a certain feature impacts the model's predictions across its full range of values.
# Dependence plot for Temperature
# ==============================================================================
fig, ax = plt.subplots(figsize=(6, 3))
scatter = shap.plots.scatter(explanation[:, "Temperature"], show=False, ax=ax)
ax.set_title("Temperature dependence plot")
ax.set_ylabel("SHAP value for the 'Temperature' feature");
The dependence plot reveals a clear U-shaped relationship: demand increases at low temperatures (heating) and, more sharply, at high temperatures (air conditioning), with its minimum in the mild range of around 15-20 degrees Celsius, where the contribution of Temperature to the prediction is negative. Capturing this type of non-linear relationship is one of the main advantages of tree-based models over linear models.
SHAP Explanations for Individual Predictions¶
SHAP values not only allow for interpreting the general behavior of the model (Global Interpretability) but also serve as a powerful tool for analyzing individual predictions (Local Interpretability). This is especially useful when trying to understand why a model made a specific prediction for a given instance.
To carry out this analysis, it is necessary to access the predictor values (lags and exogenous variables) at the time of the prediction. This can be achieved by using the create_predict_X method or by enabling the return_predictors=True argument in the backtesting_forecaster function.
SHAP values of predict() output
Suppose the forecaster is employed to predict the next 30 values of the series, and a specific prediction corresponding to the date '2014-12-28' requires explanation.
# Forecasting next 30 days
# ==============================================================================
predictions = forecaster.predict(steps=30, exog=data_test[exog_features])
set_dark_theme()
fig, ax = plt.subplots(figsize=(6, 2.5))
data_test['Demand'].plot(ax=ax, label='Test')
predictions.plot(ax=ax, label='Predictions', linestyle='--')
ax.set_xlabel(None)
ax.legend();
The method create_predict_X is used to create the input matrix used internally by the forecaster's predict method. This matrix is then used to generate SHAP values for the forecasted values.
# Create input matrix used to forecast the next 30 steps
# ==============================================================================
X_predict = forecaster.create_predict_X(steps=30, exog=data_test[exog_features])
X_predict.head(3)
| lag_1 | lag_2 | lag_3 | lag_4 | lag_5 | lag_6 | lag_7 | roll_mean_24 | Temperature | day_of_week | month | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 2014-12-02 | 237812.592388 | 234970.336660 | 189653.758108 | 202017.012448 | 214602.854760 | 218321.456402 | 214318.765210 | 211369.709659 | 19.833333 | 1 | 12 |
| 2014-12-03 | 230878.900870 | 237812.592388 | 234970.336660 | 189653.758108 | 202017.012448 | 214602.854760 | 218321.456402 | 212777.981610 | 19.616667 | 2 | 12 |
| 2014-12-04 | 230782.656189 | 230878.900870 | 237812.592388 | 234970.336660 | 189653.758108 | 202017.012448 | 214602.854760 | 214485.198829 | 21.702083 | 3 | 12 |
# SHAP values for the predictions
# ==============================================================================
explanation_predict = explainer(X_predict)
# Waterfall plot for a single prediction
# ==============================================================================
predicted_date = '2014-12-28'
iloc_predicted_date = X_predict.index.get_loc(predicted_date)
shap.plots.waterfall(explanation_predict[iloc_predicted_date], show=False)
fig = plt.gcf()
ax = fig.axes[0]
fig.set_size_inches(6, 3.5)
ax.tick_params(labelsize=10)
plt.show()
The waterfall plot illustrates how different features pushed the model’s output higher (shown in red) or lower (shown in blue), relative to the average model prediction.
lag_1had the largest negative impact, reducing the prediction by over 22,000 units.Temperaturecontributed positively, increasing the prediction by around 7,685 units.Other features like
day_of_week,month, and various lag values also influenced the prediction to a lesser degree.
The model prediction (f(x)) was 200,849.86, while the expected value (E[f(x)]), the average model output over the training data used as background, is 224,358.59. This means the specific inputs for this prediction led the model to forecast a value lower than average, largely due to the strong negative impact of lag_1.
The same insights can be obtained using the shap.force_plot function.
# Force plot for a single prediction
# ==============================================================================
shap.force_plot(
base_value = explainer.expected_value,
shap_values = explanation_predict.values[iloc_predicted_date],
features = X_predict.iloc[iloc_predicted_date, :]
)
Have you run `initjs()` in this notebook? If this notebook was from another user you must also trust this notebook (File -> Trust notebook). If you are viewing this notebook on github the Javascript has been stripped for security. If you are using JupyterLab this error is because a JupyterLab extension has not yet been written.
# Force plot for the 30 predictions
# ==============================================================================
shap.force_plot(
base_value = explainer.expected_value,
shap_values = explanation_predict.values,
features = X_predict
)
Have you run `initjs()` in this notebook? If this notebook was from another user you must also trust this notebook (File -> Trust notebook). If you are viewing this notebook on github the Javascript has been stripped for security. If you are using JupyterLab this error is because a JupyterLab extension has not yet been written.
The force plot shows the same information as the waterfall plot, but condensed into a horizontal bar: features pushing the prediction above the base value appear in red and those pushing it below appear in blue. In the version that combines the 30 predictions, each vertical slice corresponds to one date (the individual plots are rotated 90 degrees and stacked side by side). This makes it easy to detect periods in which the same features consistently drive the forecasts, such as the weekly pattern introduced by day_of_week. The plot is interactive: hovering over each region shows the feature and its contribution.
SHAP values of backtesting_forecaster() output
The analysis of individual predictions using SHAP values can also be applied to predictions made in a backtesting process. For that, the return_predictors=True argument must be set in the backtesting_forecaster function. This will return a DataFrame with the predicted value ('pred'), the partition it belongs to ('fold'), and the value of the lags and exogenous variables used to make each prediction.
In this scenario, a backtesting process is employed to train the model using data up to '2014-12-01 23:59:00'. The model then generates predictions in folds of 24 steps. SHAP values are subsequently computed for the forecast corresponding to the date '2014-12-16'.
# Backtesting returning the predictors
# ==============================================================================
cv = TimeSeriesFold(
steps = 24,
initial_train_size = len(data.loc[:'2014-12-01 23:59:00'])
)
_, predictions = backtesting_forecaster(
forecaster = forecaster,
y = data['Demand'],
exog = data[exog_features],
cv = cv,
metric = 'mean_absolute_error',
return_predictors = True,
)
predictions.head(3)
| fold | pred | lag_1 | lag_2 | lag_3 | lag_4 | lag_5 | lag_6 | lag_7 | roll_mean_24 | Temperature | day_of_week | month | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2014-12-02 | 0 | 230878.900870 | 237812.592388 | 234970.336660 | 189653.758108 | 202017.012448 | 214602.854760 | 218321.456402 | 214318.765210 | 211369.709659 | 19.833333 | 1 | 12 |
| 2014-12-03 | 0 | 230782.656189 | 230878.900870 | 237812.592388 | 234970.336660 | 189653.758108 | 202017.012448 | 214602.854760 | 218321.456402 | 212777.981610 | 19.616667 | 2 | 12 |
| 2014-12-04 | 0 | 237992.220195 | 230782.656189 | 230878.900870 | 237812.592388 | 234970.336660 | 189653.758108 | 202017.012448 | 214602.854760 | 214485.198829 | 21.702083 | 3 | 12 |
# Waterfall for a single prediction generated during backtesting
# ==============================================================================
predictions = predictions.astype(data[exog_features].dtypes) # Ensure same dtypes
iloc_predicted_date = predictions.index.get_loc('2014-12-16')
explanation_backtesting = explainer(predictions.iloc[:, 2:])
shap.plots.waterfall(explanation_backtesting[iloc_predicted_date], show=False)
fig = plt.gcf()
ax = fig.axes[0]
fig.set_size_inches(6, 3.5)
ax.tick_params(labelsize=8)
plt.show()
Scikit-learn partial dependence plots¶
Partial dependence plots (PDPs) are a useful tool for understanding the relationship between a feature and the target outcome in a machine learning model. In scikit-learn, partial dependence plots can be created with PartialDependenceDisplay.from_estimator. This function visualizes the effect of one or two features on the predicted outcome, while marginalizing the effect of all other features.
The resulting plots show how changes in the selected feature(s) affect the predicted outcome while holding other features constant on average. Remember that these plots should be interpreted in the context of your model and data. They provide insight into the relationship between specific features and the model's predictions.
In addition to the average curve, the argument kind='both' used in the next cell also draws the Individual Conditional Expectation (ICE) curves: one thin line per observation showing how its prediction changes as the value of the feature varies, keeping all other features fixed. The partial dependence curve is the average of all ICE curves. Roughly parallel ICE curves indicate that the feature affects all observations in a similar way; curves that cross or diverge are a sign of interactions with other features.
A more detailed description of partial dependence plots can be found in the scikit-learn User Guide.
# Scikit-learn partial dependence plots
# ==============================================================================
fig, ax = plt.subplots(figsize=(8, 3))
PartialDependenceDisplay.from_estimator(
estimator = forecaster.estimator,
X = X_train,
features = ["Temperature", "lag_1"],
kind = 'both',
ax = ax,
)
ax.set_title("Partial Dependence Plot")
fig.tight_layout()
plt.show()
Session information¶
import session_info
session_info.show(html=False)
----- lightgbm 4.7.0 matplotlib 3.10.9 numpy 2.4.6 pandas 2.3.3 session_info v1.0.1 shap 0.52.0 skforecast 0.25.0 sklearn 1.7.2 ----- IPython 9.15.0 jupyter_client 8.9.1 jupyter_core 5.9.1 ----- Python 3.13.14 | packaged by conda-forge | (main, Jun 12 2026, 09:44:26) [MSC v.1944 64 bit (AMD64)] Windows-11-10.0.26200-SP0 ----- Session information updated at 2026-09-24 13:16
Citation¶
How to cite this document
If you use this document or any part of it, please acknowledge the source, thank you!
Interpretable and explainable forecasting models by Joaquín Amat Rodrigo and Javier Escobar Ortiz, available under a Attribution-NonCommercial-ShareAlike 4.0 International at https://www.cienciadedatos.net/documentos/py57-interpretable-forecasting-models.html
How to cite skforecast
If you use skforecast for a publication, we would appreciate it if you cite the published software.
Zenodo:
Amat Rodrigo, Joaquin, & Escobar Ortiz, Javier. (2026). skforecast (v0.25.0). Zenodo. https://doi.org/10.5281/zenodo.8382788
APA:
Amat Rodrigo, J., & Escobar Ortiz, J. (2026). skforecast (Version 0.25.0) [Computer software]. https://doi.org/10.5281/zenodo.8382788
BibTeX:
@software{skforecast, author = {Amat Rodrigo, Joaquin and Escobar Ortiz, Javier}, title = {skforecast}, version = {0.25.0}, month = {09}, year = {2026}, license = {BSD-3-Clause}, url = {https://skforecast.org/}, doi = {10.5281/zenodo.8382788} }
Did you like the article? Your support is important
Your contribution will help me to continue generating free educational content. Many thanks! 😊
This work by Joaquín Amat Rodrigo and Javier Escobar Ortiz is licensed under a Attribution-NonCommercial-ShareAlike 4.0 International.
Allowed:
-
Share: copy and redistribute the material in any medium or format.
-
Adapt: remix, transform, and build upon the material.
Under the following terms:
-
Attribution: You must give appropriate credit, provide a link to the license, and indicate if changes were made. You may do so in any reasonable manner, but not in any way that suggests the licensor endorses you or your use.
-
NonCommercial: You may not use the material for commercial purposes.
-
ShareAlike: If you remix, transform, or build upon the material, you must distribute your contributions under the same license as the original.
