arXiv ScienceSearch

arXiv subjects

Ritabrata Ray

Publications and source records attributed to Ritabrata Ray.

5 recordsLinked to original sources

Generalization Guarantees on Data-Driven Tuning of Gradient Descent with Langevin Updates

We study learning to learn through the lens of hyperparameter tuning. We propose the Langevin Gradient Descent Algorithm (LGD), which approximates the mean of the posterior distribution defined by the loss function and regularizer of a regression task with convex objective. For classification tasks, the LGD algorithm estimates the posterior probabilities of each class on the test set. We prove the existence of an optimal hyperparameter configuration for which the LGD algorithm achieves the Bayes' optimal solution for squared loss on regression tasks, and for which LGD closely approximates the posterior probabilities for well-specified classification tasks. Subsequently, we study generalization guarantees on meta learning optimal hyperparameters for the LGD algorithm from a given set of tasks in the data-driven setting. For a number of parameters $d$ and hyperparameter dimension $h$, we show a pseudo-dimension bound of $O(dh)$, up to logarithmic terms under mild assumptions on LGD. This matches the dependence of the bounds on number of parameters obtained in prior work for linear regression using the elastic net, which only allows for $h=2$ hyperparameters, and extends their bounds to regression on convex loss. Compared to bounds on regularized logistic regression that allow for only $h=1$ hyperparameter, our bounds improve greatly on the dependence on samples per task at the cost of worse dependence on the number of parameters by accounting for hardware-aware procedures. Finally, we show empirical evidence of the success of LGD and the meta learning procedure for few-shot learning on linear and logistic regression using synthetically created datasets.

cs.LG

Eigenfunction Extraction for Ordered Representation Learning

Recent advances in representation learning reveal that widely used objectives, such as contrastive and non-contrastive, implicitly perform spectral decomposition of a contextual kernel, induced by the relationship between inputs and their contexts. Yet, these methods recover only the linear span of top eigenfunctions of the kernel, whereas exact spectral decomposition is essential for understanding feature ordering and importance. In this work, we propose a general framework to extract ordered and identifiable eigenfunctions, based on modular building blocks designed to satisfy key desiderata, including compatibility with the contextual kernel and scalability to modern settings. We then show how two main methodological paradigms, low-rank approximation and Rayleigh quotient optimization, align with this framework for eigenfunction extraction. Finally, we validate our approach on synthetic kernels and demonstrate on real-world image datasets that the recovered eigenvalues act as effective importance scores for feature selection, enabling principled efficiency-accuracy tradeoffs via adaptive-dimensional representations.

cs.LG

Sample-Optimal Zero-Violation Safety For Continuous Control

In this paper, we study the problem of ensuring safety with a few shots of samples for partially unknown systems. We first characterize a fundamental limit when producing safe actions is not possible due to insufficient information or samples. Then, we develop a technique that can generate provably safe actions and recovery behaviors using a minimum number of samples. In the performance analysis, we also establish Nagumos theorem - like results with relaxed assumptions, which is potentially useful in other contexts. Finally, we discuss how the proposed method can be integrated into a policy gradient algorithm to assure safety and stability with a handful of samples without stabilizing initial policies or generative models to probe safe actions.

eess.SY

Modular and fractional L-intersecting families of vector spaces

In the first part of this paper, we prove a theorem which is the $q$-analogue of a generalized modular Ray-Chaudhuri-Wilson Theorem shown in [Alon, Babai, Suzuki, J. Combin. Theory Series A, 1991]. It is also a generalization of the main theorem in [Frankl and Graham, European J. Combin. 1985] under certain circumstances. In the second part of this paper, we prove $q$-analogues of results on a recent notion called \emph{fractional $L$-intersecting family} for families of subspaces of a given vector space. We use the above theorem to obtain a general upper bound to the cardinality of such families. We give an improvement to this general upper bound in certain special cases.

math.CO

Fractional cross intersecting families

Let $\mathcal{A}=\{A_{1},...,A_{p}\}$ and $\mathcal{B}=\{B_{1},...,B_{q}\}$ be two families of subsets of $[n]$ such that for every $i\in [p]$ and $j\in [q]$, $|A_{i}\cap B_{j}|= \frac{c}{d}|B_{j}|$, where $\frac{c}{d}\in [0,1]$ is an irreducible fraction. We call such families "$\frac{c}{d}$-cross intersecting families". In this paper, we find a tight upper bound for the product $|\mathcal{A}||\mathcal{B}|$ and characterize the cases when this bound is achieved for $\frac{c}{d}=\frac{1}{2}$. Also, we find a tight upper bound on $|\mathcal{A}||\mathcal{B}|$ when $\mathcal{B}$ is $k$-uniform and characterize, for all $\frac{c}{d}$, the cases when this bound is achieved.

math.CO