Agnostically Learning Multi-Index Models with Queries
Abstract
Abstract. We study the power of query access for the fundamental task of agnostic learning under the Gaussian distribution. In the agnostic model, no assumptions are made on the labels of the examples, and the goal is to compute a hypothesis that is competitive with the best-fit function in a known class; i.e., it achieves error [Formula: see text], where [Formula: see text] is the error of the best function in the class. We focus on a general family of multi-Index models (MIMs), which are [Formula: see text]-variate functions that depend only on a few relevant directions, i.e., have the form [Formula: see text] for an unknown link function [Formula: see text] and a [Formula: see text] matrix [Formula: see text]. MIMs cover a wide range of commonly studied function classes, including real-valued function classes, such as constant-depth neural networks with ReLU activations, and Boolean concept classes, such as intersections of halfspaces. Our main result shows that query access gives significant runtime improvements over random examples for agnostically learning both real-valued and Boolean-valued MIMs. Under standard regularity assumptions for the link function (namely, bounded variation or surface area), we give an agnostic query learner for MIMs with running time [Formula: see text]. In contrast, algorithms that rely only on random labeled examples inherently require [Formula: see text] samples and runtime, even for the basic problem of agnostically learning a single ReLU or a halfspace. As special cases of our general approach, we obtain the following results: [Formula: see text] For the class of depth-[Formula: see text], width-[Formula: see text] ReLU networks on [Formula: see text], our agnostic query learner runs in time [Formula: see text]. This bound qualitatively matches the runtime of an algorithm by Chen, Klivans, and Meka [ Learning deep ReLU networks is fixed-parameter tractable, 2022] for the realizable PAC setting with random examples. [Formula: see text] For the class of arbitrary intersections of [Formula: see text] halfspaces on [Formula: see text], our agnostic query learner runs in time [Formula: see text]. Prior to our work, no improvement over the agnostic PAC model complexity (without queries) was known, even for the case of a single halfspace. In both these settings, we provide evidence that the [Formula: see text] runtime dependence is required for proper query learners, even for agnostically learning a single ReLU or halfspace. Our algorithmic result establishes a strong computational separation between the agnostic PAC and the agnostic PAC + Query models under the Gaussian distribution for a range of natural function classes. Prior to our work, no such separation was known for any natural concept class, even for the case of a single halfspace, for which it was an open problem posed by Feldman [ On the power of membership queries in agnostic learning, 2008]. Our results are enabled by a general dimension-reduction technique that leverages query access to estimate gradients of (a smoothed version of) the underlying label function.
Many UC-authored scholarly publications are freely available on this site because of the UC's open access policies. Let us know how this access is important for you.