Spec-Zone.ru › scikit-learn

Примечание

Перейти к концу для загрузки полного кода примера или для запуска этого примера в вашем браузере через JupyterLite или Binder

Заполнение пропущенных значений с помощью вариантов IterativeImputer

Класс IterativeImputer очень гибкий — он может использоваться с различными оценщиками для выполнения регрессии в режиме круговой очереди, рассматривая каждую переменную поочерёдно как выходную.

В этом примере мы сравниваем некоторые оценщики для заполнения пропущенных признаков с помощью IterativeImputer:

  • BayesianRidge: регрессия с регуляризацией линейной модели
  • RandomForestRegressor: регрессия леса случайных деревьев
  • make_pipeline (Nystroem, Ridge): конвейер с расширением полиномиального ядра степени 2 и регрессией с регуляризацией линейной модели
  • KNeighborsRegressor: аналогично другим подходам KNN заполнения

Особый интерес представляет возможность IterativeImputer имитировать поведение пакета missForest, популярного пакета заполнения для R.

Обратите внимание, что KNeighborsRegressor отличается от заполнения KNN, которое обучается на образцах с пропущенными значениями, используя метрику расстояния, учитывающую пропущенные значения, а не заполняя их.

Цель состоит в сравнении различных оценщиков, чтобы определить, какой из них лучше всего подходит для IterativeImputer при использовании оценщика BayesianRidge на наборе данных о жилье Калифорнии с случайным удалением одного значения из каждой строки.

Для этой конкретной схемы пропущенных значений мы видим, что BayesianRidge и RandomForestRegressor дают лучшие результаты.

Следует отметить, что некоторые оценщики, такие как HistGradientBoostingRegressor, могут напрямую обрабатывать пропущенные признаки и часто рекомендуются вместо построения конвейеров с сложными и дорогостоящими стратегиями заполнения пропущенных значений.

California Housing Regression with Different Imputation Methods
# Authors: The scikit-learn developers
# SPDX-License-Identifier: BSD-3-Clause

import matplotlib.pyplot as plt
import numpy as np
import pandas as pd

from sklearn.datasets import fetch_california_housing
from sklearn.ensemble import RandomForestRegressor

# To use this experimental feature, we need to explicitly ask for it:
from sklearn.experimental import enable_iterative_imputer  # noqa
from sklearn.impute import IterativeImputer, SimpleImputer
from sklearn.kernel_approximation import Nystroem
from sklearn.linear_model import BayesianRidge, Ridge
from sklearn.model_selection import cross_val_score
from sklearn.neighbors import KNeighborsRegressor
from sklearn.pipeline import make_pipeline

N_SPLITS = 5

rng = np.random.RandomState(0)

X_full, y_full = fetch_california_housing(return_X_y=True)
# ~2k samples is enough for the purpose of the example.
# Remove the following two lines for a slower run with different error bars.
X_full = X_full[::10]
y_full = y_full[::10]
n_samples, n_features = X_full.shape

# Estimate the score on the entire dataset, with no missing values
br_estimator = BayesianRidge()
score_full_data = pd.DataFrame(
    cross_val_score(
        br_estimator, X_full, y_full, scoring="neg_mean_squared_error", cv=N_SPLITS
    ),
    columns=["Full Data"],
)

# Add a single missing value to each row
X_missing = X_full.copy()
y_missing = y_full
missing_samples = np.arange(n_samples)
missing_features = rng.choice(n_features, n_samples, replace=True)
X_missing[missing_samples, missing_features] = np.nan

# Estimate the score after imputation (mean and median strategies)
score_simple_imputer = pd.DataFrame()
for strategy in ("mean", "median"):
    estimator = make_pipeline(
        SimpleImputer(missing_values=np.nan, strategy=strategy), br_estimator
    )
    score_simple_imputer[strategy] = cross_val_score(
        estimator, X_missing, y_missing, scoring="neg_mean_squared_error", cv=N_SPLITS
    )

# Estimate the score after iterative imputation of the missing values
# with different estimators
estimators = [
    BayesianRidge(),
    RandomForestRegressor(
        # We tuned the hyperparameters of the RandomForestRegressor to get a good
        # enough predictive performance for a restricted execution time.
        n_estimators=4,
        max_depth=10,
        bootstrap=True,
        max_samples=0.5,
        n_jobs=2,
        random_state=0,
    ),
    make_pipeline(
        Nystroem(kernel="polynomial", degree=2, random_state=0), Ridge(alpha=1e3)
    ),
    KNeighborsRegressor(n_neighbors=15),
]
score_iterative_imputer = pd.DataFrame()
# iterative imputer is sensible to the tolerance and
# dependent on the estimator used internally.
# we tuned the tolerance to keep this example run with limited computational
# resources while not changing the results too much compared to keeping the
# stricter default value for the tolerance parameter.
tolerances = (1e-3, 1e-1, 1e-1, 1e-2)
for impute_estimator, tol in zip(estimators, tolerances):
    estimator = make_pipeline(
        IterativeImputer(
            random_state=0, estimator=impute_estimator, max_iter=25, tol=tol
        ),
        br_estimator,
    )
    score_iterative_imputer[impute_estimator.__class__.__name__] = cross_val_score(
        estimator, X_missing, y_missing, scoring="neg_mean_squared_error", cv=N_SPLITS
    )

scores = pd.concat(
    [score_full_data, score_simple_imputer, score_iterative_imputer],
    keys=["Original", "SimpleImputer", "IterativeImputer"],
    axis=1,
)

# plot california housing results
fig, ax = plt.subplots(figsize=(13, 6))
means = -scores.mean()
errors = scores.std()
means.plot.barh(xerr=errors, ax=ax)
ax.set_title("California Housing Regression with Different Imputation Methods")
ax.set_xlabel("MSE (smaller is better)")
ax.set_yticks(np.arange(means.shape[0]))
ax.set_yticklabels([" w/ ".join(label) for label in means.index.tolist()])
plt.tight_layout(pad=1)
plt.show()

Общее время выполнения скрипта: (0 минут 6,489 секунды)

Launch binder
Launch JupyterLite

Download Jupyter notebook: plot_iterative_imputer_variants_comparison.ipynb

Download Python source code: plot_iterative_imputer_variants_comparison.py

Download zipped: plot_iterative_imputer_variants_comparison.zip

Связанные примеры

Заполнение пропущенных значений перед построением оценщика

Основные моменты выпуска для scikit-learn 0.22

Отображение оценщиков и сложных конвейеров

Объединение предикторов с использованием стекинга

© 2007–2025 The scikit-learn developers
Licensed under the 3-clause BSD License.
https://scikit-learn.org/1.6/auto_examples/impute/plot_iterative_imputer_variants_comparison.html

Spec-Zone.ru

Настройки Оффлайн Что нового Помощь О нас
Spec-Zone .ru
спецификации, руководства, описания, API