Nearest Neighbor Sampling for Covariate Shift Adaptation

François Portier; Lionel Truquet; Ikko Yamane

Pré-Publication, Document De Travail Année : 2024

Nearest Neighbor Sampling for Covariate Shift Adaptation

(1, 2) , (1, 2) , (1, 2)

1
2

François Portier

Fonction : Auteur
PersonId : 1400246

Ecole Nationale de la Statistique et de l'Analyse de l'Information [Bruz]

Centre de Recherche en Economie et Statistique [Bruz]

Lionel Truquet

Fonction : Auteur
PersonId : 974950

Ecole Nationale de la Statistique et de l'Analyse de l'Information [Bruz]

Centre de Recherche en Economie et Statistique [Bruz]

Ikko Yamane

Fonction : Auteur
PersonId : 1400247

Ecole Nationale de la Statistique et de l'Analyse de l'Information [Bruz]

Centre de Recherche en Economie et Statistique [Bruz]

Résumé

Many existing covariate shift adaptation methods estimate sample weights given to loss values to mitigate the gap between the source and the target distribution. However, estimating the optimal weights typically involves computationally expensive matrix inversion and hyper-parameter tuning. In this paper, we propose a new covariate shift adaptation method which avoids estimating the weights. The basic idea is to directly work on unlabeled target data, labeled according to the k-nearest neighbors in the source dataset. Our analysis reveals that setting $k=1$ is an optimal choice. This property removes the necessity of tuning the only hyper-parameter $k$ and leads to a running time quasi-linear in the sample size. Our results include sharp rates of convergence for our estimator, with a tight control of the mean square error and explicit constants. In particular, the variance of our estimators has the same rate of convergence as for standard parametric estimation despite their non-parametric nature. The proposed estimator shares similarities with some matching-based treatment effect estimators used, e.g., in biostatistics, econometrics, and epidemiology. Our experiments show that it achieves drastic reduction in the running time with remarkable accuracy.

Mots clés

Nearest Neighbor Covariate Shift Machine Learning

Domaines

Machine Learning [stat.ML]

Fichier principal

2312.09969v2.pdf (1.31 Mo)

Origine	Fichiers produits par l'(les) auteur(s)

Ikko Yamane : Connectez-vous pour contacter le contributeur

https://hal.science/hal-04645530

Soumis le : jeudi 11 juillet 2024-17:47:50

Dernière modification le : mardi 16 juillet 2024-03:09:07

Dates et versions

hal-04645530 , version 1 (11-07-2024)

Licence

Paternité

Identifiants

HAL Id : hal-04645530 , version 1

Citer

François Portier, Lionel Truquet, Ikko Yamane. Nearest Neighbor Sampling for Covariate Shift Adaptation. 2024. ⟨hal-04645530⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

GENES CREST ENSAI

78 Consultations

60 Téléchargements

Nearest Neighbor Sampling for Covariate Shift Adaptation

Résumé

Mots clés

Domaines

Dates et versions

Licence

Identifiants

Citer

Exporter

Collections

Partager