Примечание
Перейти к концу для загрузки полного кода примера или для запуска этого примера в вашем браузере через JupyterLite или Binder
Заполнение пропущенных значений с помощью вариантов IterativeImputer
Класс IterativeImputer очень гибкий — он может использоваться с различными оценщиками для выполнения регрессии в режиме круговой очереди, рассматривая каждую переменную поочерёдно как выходную.
В этом примере мы сравниваем некоторые оценщики для заполнения пропущенных признаков с помощью IterativeImputer:
-
BayesianRidge: регрессия с регуляризацией линейной модели -
RandomForestRegressor: регрессия леса случайных деревьев -
make_pipeline(Nystroem,Ridge): конвейер с расширением полиномиального ядра степени 2 и регрессией с регуляризацией линейной модели -
KNeighborsRegressor: аналогично другим подходам KNN заполнения
Особый интерес представляет возможность IterativeImputer имитировать поведение пакета missForest, популярного пакета заполнения для R.
Обратите внимание, что KNeighborsRegressor отличается от заполнения KNN, которое обучается на образцах с пропущенными значениями, используя метрику расстояния, учитывающую пропущенные значения, а не заполняя их.
Цель состоит в сравнении различных оценщиков, чтобы определить, какой из них лучше всего подходит для IterativeImputer при использовании оценщика BayesianRidge на наборе данных о жилье Калифорнии с случайным удалением одного значения из каждой строки.
Для этой конкретной схемы пропущенных значений мы видим, что BayesianRidge и RandomForestRegressor дают лучшие результаты.
Следует отметить, что некоторые оценщики, такие как HistGradientBoostingRegressor, могут напрямую обрабатывать пропущенные признаки и часто рекомендуются вместо построения конвейеров с сложными и дорогостоящими стратегиями заполнения пропущенных значений.

# Authors: The scikit-learn developers
# SPDX-License-Identifier: BSD-3-Clause
import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
from sklearn.datasets import fetch_california_housing
from sklearn.ensemble import RandomForestRegressor
# To use this experimental feature, we need to explicitly ask for it:
from sklearn.experimental import enable_iterative_imputer # noqa
from sklearn.impute import IterativeImputer, SimpleImputer
from sklearn.kernel_approximation import Nystroem
from sklearn.linear_model import BayesianRidge, Ridge
from sklearn.model_selection import cross_val_score
from sklearn.neighbors import KNeighborsRegressor
from sklearn.pipeline import make_pipeline
N_SPLITS = 5
rng = np.random.RandomState(0)
X_full, y_full = fetch_california_housing(return_X_y=True)
# ~2k samples is enough for the purpose of the example.
# Remove the following two lines for a slower run with different error bars.
X_full = X_full[::10]
y_full = y_full[::10]
n_samples, n_features = X_full.shape
# Estimate the score on the entire dataset, with no missing values
br_estimator = BayesianRidge()
score_full_data = pd.DataFrame(
cross_val_score(
br_estimator, X_full, y_full, scoring="neg_mean_squared_error", cv=N_SPLITS
),
columns=["Full Data"],
)
# Add a single missing value to each row
X_missing = X_full.copy()
y_missing = y_full
missing_samples = np.arange(n_samples)
missing_features = rng.choice(n_features, n_samples, replace=True)
X_missing[missing_samples, missing_features] = np.nan
# Estimate the score after imputation (mean and median strategies)
score_simple_imputer = pd.DataFrame()
for strategy in ("mean", "median"):
estimator = make_pipeline(
SimpleImputer(missing_values=np.nan, strategy=strategy), br_estimator
)
score_simple_imputer[strategy] = cross_val_score(
estimator, X_missing, y_missing, scoring="neg_mean_squared_error", cv=N_SPLITS
)
# Estimate the score after iterative imputation of the missing values
# with different estimators
estimators = [
BayesianRidge(),
RandomForestRegressor(
# We tuned the hyperparameters of the RandomForestRegressor to get a good
# enough predictive performance for a restricted execution time.
n_estimators=4,
max_depth=10,
bootstrap=True,
max_samples=0.5,
n_jobs=2,
random_state=0,
),
make_pipeline(
Nystroem(kernel="polynomial", degree=2, random_state=0), Ridge(alpha=1e3)
),
KNeighborsRegressor(n_neighbors=15),
]
score_iterative_imputer = pd.DataFrame()
# iterative imputer is sensible to the tolerance and
# dependent on the estimator used internally.
# we tuned the tolerance to keep this example run with limited computational
# resources while not changing the results too much compared to keeping the
# stricter default value for the tolerance parameter.
tolerances = (1e-3, 1e-1, 1e-1, 1e-2)
for impute_estimator, tol in zip(estimators, tolerances):
estimator = make_pipeline(
IterativeImputer(
random_state=0, estimator=impute_estimator, max_iter=25, tol=tol
),
br_estimator,
)
score_iterative_imputer[impute_estimator.__class__.__name__] = cross_val_score(
estimator, X_missing, y_missing, scoring="neg_mean_squared_error", cv=N_SPLITS
)
scores = pd.concat(
[score_full_data, score_simple_imputer, score_iterative_imputer],
keys=["Original", "SimpleImputer", "IterativeImputer"],
axis=1,
)
# plot california housing results
fig, ax = plt.subplots(figsize=(13, 6))
means = -scores.mean()
errors = scores.std()
means.plot.barh(xerr=errors, ax=ax)
ax.set_title("California Housing Regression with Different Imputation Methods")
ax.set_xlabel("MSE (smaller is better)")
ax.set_yticks(np.arange(means.shape[0]))
ax.set_yticklabels([" w/ ".join(label) for label in means.index.tolist()])
plt.tight_layout(pad=1)
plt.show()
Общее время выполнения скрипта: (0 минут 6,489 секунды)
Связанные примеры
© 2007–2025 The scikit-learn developers
Licensed under the 3-clause BSD License.
https://scikit-learn.org/1.6/auto_examples/impute/plot_iterative_imputer_variants_comparison.html