1 INTRODUCTION

The advancement of neural networks has significantly driven the development of optimization methods [1–3]. Nevertheless, a notable challenge lies in the extensive number of parameters, which can result in numerous potential local and global minima of the loss function [4–7]. Many studies have endeavored to identify the most optimal minima from diverse perspectives [8–10].

Flatter minima are associated with superior generalization properties. This issue has been explored theoretically and empirically in many studies, including [11–14]. The process of locating a flat minimum is not straightforward. In such contexts, the Hessian matrix, which represents the second-order derivative of the loss function, is frequently utilized. This matrix is crucial for understanding the local behavior of the function around the point of extremum.

Existing research [15–22] has addressed the question of Hessian spectra in typical neural network architectures, revealing a characteristic pattern. The spectrum typically comprises a significant number of eigenvalues close to zero, referred to as bulk eigenvalues, and a smaller number of distinct eigenvalues, known as outliers [15]. Specifically, in a \(K\)-label classification task, precisely \(K\) non-zero eigenvalues can be observed [16, 17]. Theoretical estimations of these non-vanishing eigenvalues enable the establishment of an upper bound on the local rate of growth of the loss function.

However, the issue of altering the landscape of the loss function when adding new objects to the sample remains unresolved. In this paper, we meticulously investigate how the loss surface transforms as the sample size increases. Specifically, we examine the absolute value of the loss function changes when adding another object. An overview of our observations is presented in Fig. 1.

Fig. 1.
Fig. 1.
Full size image

Overview of our observations. Part (a) shows the loss function landscape, which is a surface in the parameters space. Part (b) shows the losses difference. It arises, when one more object is added to the dataset. Here we exhibit the behavior for dimension equals 2. Near the minimum θ*, the mean loss value for \(k + 1\) objects \({{\mathcal{L}}_{{k + 1}}}(\boldsymbol\theta )\) tends to be similar to the same for k objects \({{\mathcal{L}}_{k}}(\boldsymbol\theta )\).

We obtain theoretical estimates for the convergence of this difference in a fully connected neural network as the sample size approaches infinity. These results are derived through the analysis of the Hessian spectrum. These estimates allow us to determine the dependence of this difference on the structure of the neural network, including the size of the hidden layer and the number of layers. We empirically verify these theoretical results by examining the behavior of the loss surface on various datasets. The obtained plots substantiate the validity of the theoretical calculations.

Contributions. Our contributions can be summarized as follows:

• We present a theoretical analysis of the convergence of the loss landscape in a fully connected neural network as the sample size increases, deriving upper bounds for the difference in loss function values when adding a new object to the sample.

• We demonstrate the validity of our theoretical results through empirical studies on various datasets, showing that the loss function surface exhibits convergence for image classification tasks.

• We highlight the implications of our findings for understanding the local geometry of neural loss landscapes and for the development of sample size determination techniques, addressing a previously unexplored issue in the field.

Outline. The rest of the paper is organized as follows. Section 2 divides existing results into some topics, highlighting their key contributions and findings. Section 3 considers general notation and some preliminary calculations. In Section 4, we provide theoretical bounds on the hessian and losses difference norms. Empirical study of the obtained results is given in Section 5. We summarize and present the results in Sections 6 and 7. Additional experiments and proofs of theorems are included in the Appendix A.

2 RELATED WORK

Understanding the geometry of neural network loss landscapes. The geometry of neural network loss landscapes, particularly through the Hessian matrix, has been extensively studied. [14] identifies key properties including that in the multi-label classificaion problem the landscape exhibits exactly \(K\) directions of high positive curvature, where \(K\) is the number of classes. [23] and [15] use random matrix theory and spectral analysis to understand the loss surface dynamics and optimization. [24] introduces a model for understanding linear mode connectivity (LMC) by analyzing the topography of the loss landscape. [25] provides a theoretical understanding of the double descent phenomenon in finite-width neural networks, leveraging influence functions to derive expressions for the population loss. [26] characterizes the instabilities of gradient descent during training with large learning rates, observing landscape flattening and shift. However, nowhere is the question raised that this geometry becomes unchanged with an increase in the number of objects.

Hessian-based generalization and optimization. The Hessian matrix is crucial for studying generalization and optimization in neural networks. In [27] a Hessian-based distance for fine-tuning against label noise was introduced, which can match the scale of the observed generalization gap of fine-tuned models in practice. [28] optimizes for wider local minima using training data, while concurrently maintaining low loss values on validation data to improve generalization. Authors of [29] ties loss curvature to input-output behavior, explaining progressive sharpening and flat minima. All these studies provide insights into optimizing and generalizing neural networks using Hessian-based methods. However, none of these papers address the impact of changing sample size.

Spectral analysis and structural insights. Spectral analysis of the Hessian matrix offers insights into neural network structure and properties. There is a variety of works [15–17] underlying that the Hessian matrix of typical loss surface exhibits a spectrum composed of two parts: a bulk centered near zero and outliers away from the bulk. [18] shows that depending on the data properties, the nonlinear response model, and the loss function, the Hessian can have qualitatively different spectral behaviors. [19] highlights the Hessian’s role in sharpness regularization. [20] reveals class/cross-class structure in Hessian spectra. [17] and [21] analyze Hessian dynamics and hierarchical structure. [22] unifies low-dimensional observations, explaining Hessian spectra and gradient descent alignment. [30] and [16] analyze the Hessian’s eigenvalue distribution, highlighting over-parametrization and data dependency. [31] discovers and mathematically models the power-law Hessian spectrum, providing a maximum entropy interpretation and a framework for spectral analysis in deep learning. However, these low-rank approximations have not been sufficiently investigated in terms of the convergence of the corresponding spectra.

Decomposing and analyzing the hessian matrix. Decomposing the Hessian matrix provides insights into neural network training and generalization. [32] proposes a decoupling conjecture to analyze Hessian properties, decomposing it as the Kronecker product of two smaller matrices. [33] explores CNN Hessian maps, revealing the Hessian rank grows as the square root of the number of parameters. [34] provides exact formulas and tight upper bounds for the Hessian rank of deep linear networks. [35] simplifies high-dimensional derivative calculations using tensor calculus. However, the limit properties of the Hessian with increasing sample size have not been sufficiently investigated.

Sample size determination. A significant amount of research has been conducted on the topic of determining a sufficient sample size. Generally, these studies can be divided into several categories: statistical, heuristic, and Bayesian. We observe the greatest consistency between our results and articles [36, 37]. In the first work, a method is proposed to determine the sample size that accounts for the Kullback–Leibler divergence between the posterior distributions of parameters across subsamples. The second paper is a successful synthesis of a wide range of existing methods, unifying them and implementing them in a comprehensive library. However, none of these previous studies on sample size determination have considered complex models such as neural networks. We believe that our findings can be applied to such models, and we intend to explore this further in future research.

3 PRELIMINARIES

3.1 General Notation

In this section, we introduce the general notation used in the rest of the paper and the basic assumptions. Consider a conditional probability \(p({\mathbf{y}}|{\mathbf{x}})\), that maps the given unobserved variable \({\mathbf{x}} \in \mathcal{X}\) to the corresponding output \({\mathbf{y}} \in \mathcal{Y}\), and which we try to approximate using neural network \({{f}_{{\boldsymbol{\theta }}}}( \cdot )\) with parameters \({\boldsymbol{\theta }} \in {{\mathbb{R}}^{P}}\).

Given the dataset

$$\mathfrak{D} = \left\{ {({{{\mathbf{x}}}_{i}},{{{\mathbf{y}}}_{i}})} \right\},\quad i = 1, \ldots ,m,$$

the empirical loss function calculated for all the given dataset of size \(m\) is

$$\mathcal{L}({\boldsymbol{\theta }}) = \frac{1}{m}\sum\limits_{i = 1}^m \ell ({{f}_{\boldsymbol\theta }}({{{\mathbf{x}}}_{i}}),{{{\mathbf{y}}}_{i}}) \approx {{\mathbb{E}}_{{({\mathbf{x}},{\mathbf{y}}) \sim p({\mathbf{x}},{\mathbf{y}})}}}\left[ {\ell ({{f}_{{\boldsymbol{\theta }}}}({\mathbf{x}}),{\mathbf{y}})} \right],$$

so it is an approximation of general loss function, which is calculated using the joint distribution \(p({\mathbf{x}},{\mathbf{y}})\). Here we denote the per-object loss function as \(\ell ({\mathbf{z}},{\mathbf{y}})\), e.g., cross-entropy loss \({\text{CE}}({\text{softmax}}({\mathbf{z}}),{\mathbf{y}})\) in multi-label classification. If we fix first k samples, corresponding loss function is

$${{\mathcal{L}}_{k}}({\boldsymbol{\theta }}) = \frac{1}{k}\sum\limits_{i = 1}^k \ell ({{f}_{{\boldsymbol{\theta }}}}({{{\mathbf{x}}}_{i}}),{{{\mathbf{y}}}_{i}}).$$

Difference between losses for sample sizes \(k + 1\) and k is

$${{\mathcal{L}}_{{k + 1}}}({\boldsymbol{\theta }}) - {{\mathcal{L}}_{k}}({\boldsymbol{\theta }}) = \frac{1}{{k + 1}}\left( {\ell ({{f}_{{\boldsymbol{\theta }}}}({{{\mathbf{x}}}_{{k + 1}}}),{{{\mathbf{y}}}_{{k + 1}}}) - {{\mathcal{L}}_{k}}({\boldsymbol{\theta }})} \right).$$
(1)

The further investigation in the work is aimed at studying exactly this difference, which occurs when adding another object to the dataset. We are especially interested in limiting properties when the sample size tends to infinity.

Assumption 1. Let \({\boldsymbol{\theta }}{\kern 1pt} \text{*}\) be the local minimum of both empirical loss functions \({{\mathcal{L}}_{k}}({\boldsymbol{\theta }})\) and \({{\mathcal{L}}_{{k + 1}}}({\boldsymbol{\theta }})\), i.e., \(\nabla {{\mathcal{L}}_{k}}({\boldsymbol{\theta }}{\kern 1pt} \text{*})\, = \,\nabla {{\mathcal{L}}_{{k + 1}}}({\boldsymbol{\theta }}{\kern 1pt} \text{*})\) = 0.

This assumption will allow us to investigate the behavior of the loss function landscape, limiting ourselves to considering just one point.

3.2 Second-Order Approximation

Let us use second-order Taylor approximation for mentioned above loss functions at \({\boldsymbol{\theta }}{\kern 1pt} \text{*}\). We suppose that decomposition to the second order will be sufficient to study local behavior. The first-order term vanishes because the gradients \(\nabla {{\mathcal{L}}_{k}}({\boldsymbol{\theta }}{\kern 1pt} \text{*})\) and \(\nabla {{\mathcal{L}}_{{k + 1}}}({\boldsymbol{\theta }}{\kern 1pt} \text{*})\) zeroes:

$${{\mathcal{L}}_{k}}({\boldsymbol{\theta }}) \approx {{\mathcal{L}}_{k}}({\boldsymbol{\theta }}{\kern 1pt} \text{*}) + \frac{1}{2}{{({\boldsymbol{\theta }} - {\boldsymbol{\theta }}{\kern 1pt} \text{*})}^{{\text{T}}}}{{{\mathbf{H}}}^{{(k)}}}({\boldsymbol{\theta }}{\kern 1pt} \text{*})({\boldsymbol{\theta }} - {\boldsymbol{\theta }}{\kern 1pt} \text{*}),$$
(2)

where we denoted the Hessian of \({{\mathcal{L}}_{k}}({\boldsymbol{\theta }})\) w.r.t. parameters θ at \({\boldsymbol{\theta }}{\kern 1pt} \text{*}\) as \({{{\mathbf{H}}}^{{(k)}}}({\boldsymbol{\theta }}{\kern 1pt} \text{*}) \in {{\mathbb{R}}^{{P \times P}}}\). Moreover, the total Hessian can be written as the average value of the Hessians of the individual terms of the empirical loss function:

$$\begin{gathered} {{{\mathbf{H}}}^{{(k)}}}({\boldsymbol{\theta }}) = \nabla _{{\boldsymbol{\theta }}}^{2}{{\mathcal{L}}_{k}}({\boldsymbol{\theta }}) \\ = \frac{1}{k}\sum\limits_{i = 1}^k \nabla _{{\boldsymbol{\theta }}}^{2}\ell ({{f}_{{\boldsymbol{\theta }}}}({{{\mathbf{x}}}_{i}}),{{{\mathbf{y}}}_{i}}) = \frac{1}{k}\sum\limits_{i = 1}^k {{{\mathbf{H}}}_{i}}({\boldsymbol{\theta }}). \\ \end{gathered} $$

Therefore, using the obtained second-order approximation (2), formula for the difference of losses (1) becomes

$$\begin{gathered} {{\mathcal{L}}_{{k + 1}}}({\boldsymbol{\theta }}) - {{\mathcal{L}}_{k}}({\boldsymbol{\theta }}) \\ \, = \frac{1}{{k + 1}}\left( {\ell ({{f}_{{{\boldsymbol{\theta }}{\kern 1pt} \text{*}}}}({{{\mathbf{x}}}_{{k + 1}}}),{{{\mathbf{y}}}_{{k + 1}}}) - \frac{1}{k}\sum\limits_{i = 1}^k \ell ({{f}_{{{\boldsymbol{\theta }}{\kern 1pt} \text{*}}}}({{{\mathbf{x}}}_{i}}),{{{\mathbf{y}}}_{i}})} \right) \\ \, + \frac{1}{{k + 1}}{{({\boldsymbol{\theta }} - {\boldsymbol{\theta }}{\kern 1pt} \text{*})}^{{\text{T}}}}\left( {{{{\mathbf{H}}}_{{k + 1}}}({\boldsymbol{\theta }}{\kern 1pt} \text{*}) - \frac{1}{k}\sum\limits_{i = 1}^k {{{\mathbf{H}}}_{i}}({\boldsymbol{\theta }}{\kern 1pt} \text{*})} \right)({\boldsymbol{\theta }} - {\boldsymbol{\theta }}{\kern 1pt} \text{*}). \\ \end{gathered} $$

After that, using triangle inequality, we can derive the following:

$$\begin{gathered} \left| {{{\mathcal{L}}_{{k + 1}}}({\boldsymbol{\theta }}) - {{\mathcal{L}}_{k}}({\boldsymbol{\theta }})} \right| \\ \leqslant \;\frac{1}{{k + 1}}\left| {\ell ({{f}_{{{\boldsymbol{\theta }}{\kern 1pt} \text{*}}}}({{{\mathbf{x}}}_{{k + 1}}}),{{{\mathbf{y}}}_{{k + 1}}}) - \frac{1}{k}\sum\limits_{i = 1}^k \ell ({{f}_{{{\boldsymbol{\theta }}{\kern 1pt} \text{*}}}}({{{\mathbf{x}}}_{i}}),{{{\mathbf{y}}}_{i}})} \right| \\ \, + \frac{1}{{k + 1}}\left\| {{\boldsymbol{\theta }} - {\boldsymbol{\theta }}{\kern 1pt} \text{*}} \right\|_{2}^{2}{{\left\| {{{{\mathbf{H}}}_{{k + 1}}}({\boldsymbol{\theta }}{\kern 1pt} \text{*}) - \frac{1}{k}\sum\limits_{i = 1}^k {{{\mathbf{H}}}_{i}}({\boldsymbol{\theta }}{\kern 1pt} \text{*})} \right\|}_{2}}. \\ \end{gathered} $$

So the problem of the boundedness and convergence of the losses difference is reduced to the analysis of the two terms:

• Difference of the loss functions at optima for new object and previous ones:

$$\left| {\ell ({{f}_{{{\boldsymbol{\theta }}{\kern 1pt} \text{*}}}}({{{\mathbf{x}}}_{{k + 1}}}),{{{\mathbf{y}}}_{{k + 1}}}) - \frac{1}{k}\sum\limits_{i = 1}^k \ell ({{f}_{{{\boldsymbol{\theta }}{\kern 1pt} \text{*}}}}({{{\mathbf{x}}}_{i}}),{{{\mathbf{y}}}_{i}})} \right|,$$

• Difference of the Hessians at optima for new object and previous ones:

$${{\left\| {{{{\mathbf{H}}}_{{k + 1}}}({\boldsymbol{\theta }}{\kern 1pt} \text{*}) - \frac{1}{k}\sum\limits_{i = 1}^k {{{\mathbf{H}}}_{i}}({\boldsymbol{\theta }}{\kern 1pt} \text{*})} \right\|}_{2}}.$$

It should be mentioned that the first term can be easily upper-bounded by a constant, since the loss function itself takes limited values. However, the expression with Hessians is not so easy to evaluate. The rest of the work is devoted to a thorough analysis of this difference. Thus, we analyze the local convergence of the landscape of the loss function using its Hessian.

3.3 Fully Connected Neural Network

The main results of this work are obtained for K-label classification problem with a cross-entropy loss function. In this case the input vector is \({\mathbf{x}} \in {{\mathbb{R}}^{n}}\) and the output \({\mathbf{y}} \in {{\mathbb{R}}^{K}}\), which is a one-hot vector, with all components equal to 0 except a single component \({{y}_{k}}\) if and only if k is the correct class label for input x. Consider a L-layer fully connected network \({{f}_{\boldsymbol\theta }}( \cdot )\) with ReLU activation function after each linear layer. With \(\sigma ({\mathbf{x}}) = \left[ {{\mathbf{x}}\; \geqslant \;{\mathbf{0}}} \right]{\mathbf{x}}\) as the Rectified Linear Unit (ReLU) function, the output of this network is a vector of logits \({\mathbf{z}} \in {{\mathbb{R}}^{K}}\). They are computed recursively as

$${{{\mathbf{z}}}^{{(p)}}} = {{{\mathbf{W}}}^{{(p)}}}{{{\mathbf{x}}}^{{(p)}}} + {{{\mathbf{b}}}^{{(p)}}},$$
$${{{\mathbf{x}}}^{{(p + 1)}}} = \sigma ({{{\mathbf{z}}}^{{(p)}}}).$$

Here we denote the input and output of the p-th layer as \({{{\mathbf{x}}}^{{(p)}}}\) and \({{{\mathbf{z}}}^{{(p)}}}\), and set \({{{\mathbf{x}}}^{{(1)}}} = {\mathbf{x}}\), z = \({{f}_{{\boldsymbol{\theta }}}}({\mathbf{x}})\, = \,{{{\mathbf{z}}}^{{(L)}}}\). Also we denote θ = \({\text{col}}({{{\mathbf{w}}}^{{(1)}}},{{{\mathbf{b}}}^{{(1)}}}, \ldots ,{{{\mathbf{w}}}^{{(L)}}},{{{\mathbf{b}}}^{{(L)}}}) \in {{\mathbb{R}}^{P}}\) the parameters of the network.

For the p-th layer, \({{{\mathbf{w}}}^{{(p)}}}\) is the flattened weight matrix \({{{\mathbf{W}}}^{{(p)}}}\) and \({{{\mathbf{b}}}^{{(p)}}}\) is its corresponding bias vector. We denote \({\mathbf{p}} = {\text{softmax}}({\mathbf{z}}) \in {{\mathbb{R}}^{K}}\) as the output confidence, i.e.

$${{p}_{i}} = {\text{softmax}}{{({\mathbf{z}})}_{i}} = \frac{{\exp ({{z}_{i}})}}{{\sum\nolimits_{j = 1}^K {\exp ({{z}_{j}})} }} \in (0;1).$$

The loss function is cross-entropy loss:

$$\ell ({\mathbf{z}},{\mathbf{y}}) = {\text{CE}}({\mathbf{p}},{\mathbf{y}}) = - \sum\limits_{k = 1}^K {{y}_{k}}\log {{p}_{k}} \in {{\mathbb{R}}^{ + }}.$$

3.4 Hessian Decomposition

It is quite well known [16] that, via the chain rule [35], the Hessian can be decomposed as a sum of the following two matrices:

$${{{\mathbf{H}}}_{i}}({\boldsymbol{\theta }}) = \underbrace {{{\nabla }_{{\boldsymbol{\theta }}}}{{{\mathbf{z}}}_{i}}\frac{{{{\partial }^{2}}\ell ({{{\mathbf{z}}}_{i}},{{{\mathbf{y}}}_{i}})}}{{\partial {\mathbf{z}}_{i}^{2}}}{{\nabla }_{{\boldsymbol{\theta }}}}{\mathbf{z}}_{i}^{{\text{T}}}}_{{\text{G-term}}} + \underbrace {\sum\limits_{k = 1}^K \frac{{\partial \ell ({{{\mathbf{z}}}_{i}},{{{\mathbf{y}}}_{i}})}}{{\partial {{z}_{{ik}}}}}\nabla _{{\boldsymbol{\theta }}}^{2}{{z}_{{ik}}}}_{{\text{H-term}}},$$

where \({{\nabla }_{{\boldsymbol{\theta }}}}{{{\mathbf{z}}}_{i}} \in {{\mathbb{R}}^{{P \times K}}}\) is the Jacobian of the neural network function and \(\frac{{{{\partial }^{2}}\ell ({{{\mathbf{z}}}_{i}},{{{\mathbf{y}}}_{i}})}}{{\partial {\mathbf{z}}_{i}^{2}}}\) is the Hessian of the loss with respect to the network function, at the i-th sample. As it was mentioned [15–17], the Hessian spectrum consists of bulk which is concentrated around zero (corresponds to the H-term), and the edges which are scattered away from zero (G-term). We will focus on the top eigenspace, so we can approximate our full Hessian using only G-term, as

$${{{\mathbf{H}}}_{i}}({\boldsymbol{\theta }}) \approx {{\nabla }_{{\boldsymbol{\theta }}}}{{{\mathbf{z}}}_{i}}\frac{{{{\partial }^{2}}\ell ({{{\mathbf{z}}}_{i}},{{{\mathbf{y}}}_{i}})}}{{\partial {\mathbf{z}}_{i}^{2}}}{{\nabla }_{{\boldsymbol{\theta }}}}{\mathbf{z}}_{i}^{{\text{T}}}.$$

Moreover, recent works exploring the Neural Tangent Kernel [38, 39] assume that the logits z depend only linearly on the weights θ, which implies that the logit curvatures \(\nabla _{{\boldsymbol{\theta }}}^{2}{{z}_{{ik}}}\), and therefore the H-term are identically zero.

4 CONVERGENCE OF THE LOSS DIFFERENCE

In this section, we present the results obtained regarding the bounding the Hessian norm. After that, we apply the obtained upper bound on the Hessian to get a rate of the losses difference convergence. Despite the fact that this is only an upper bound and it is not always achieved, in the Section 5 we analyze the general trends for practical application.

Authors of [40] derived the formula of G-term in the fully connected neural network, so in our case we use the following approximation: \({{{\mathbf{H}}}_{i}}({\boldsymbol{\theta }}) \approx {\mathbf{F}}_{i}^{{\text{T}}}{{{\mathbf{A}}}_{i}}{{{\mathbf{F}}}_{i}}\). Here the following is denoted (we omit the index i for simplicity):

• Matrix representation of the ReLU activation function:

$${{{\mathbf{D}}}^{{(p)}}} = {\text{diag}}([{{{\mathbf{z}}}^{{(p)}}}\; \geqslant \;{\mathbf{0}}]),$$

• The partial derivative of logits w.r.t. logits at p-th layer:

$${{{\mathbf{G}}}^{{(p)}}} = \frac{{\partial {\mathbf{z}}}}{{\partial {{{\mathbf{z}}}^{{(p)}}}}} = {{{\mathbf{W}}}^{{(L)}}}{{{\mathbf{D}}}^{{(L - 1)}}}{{{\mathbf{W}}}^{{(L - 1)}}}{{{\mathbf{D}}}^{{(L - 2)}}} \cdot \ldots \cdot {{{\mathbf{D}}}^{{(p)}}},$$

• Its stacked version:

$${{{\mathbf{F}}}^{{\text{T}}}} = \left( {\begin{array}{*{20}{c}} {{{{({{{\mathbf{G}}}^{{(1)}}})}}^{{\text{T}}}} \otimes {{{\mathbf{x}}}^{{(1)}}}} \\ {{{{({{{\mathbf{G}}}^{{(1)}}})}}^{{\text{T}}}}} \\ \vdots \\ {{{{({{{\mathbf{G}}}^{{(L)}}})}}^{{\text{T}}}} \otimes {{{\mathbf{x}}}^{{(L)}}}} \\ {{{{({{{\mathbf{G}}}^{{(L)}}})}}^{{\text{T}}}}} \end{array}} \right),$$

• And the Hessian of the loss function w.r.t. logits (according to [41]):

$${\mathbf{A}} = \nabla _{{\mathbf{z}}}^{2}\ell ({\mathbf{z}},{\mathbf{y}}) = {\text{diag}}({\mathbf{p}}) - {{\mathbf{p}}{{\mathbf{p}}}^{{\text{T}}}}.$$

To the best of our knowledge, no one had received an estimate for the norms of such an approximation of the Hessian previously. Next, we present a Theorem 1 that contains an upper bound on the spectral norm of the Hessian in a fully connected neural network.

4.1 Boundedness of the Hessian

Below there is a Theorem 1, a detailed proof of which is provided in Appendix A.2.

Theorem 1. Consider a \(L\)-layer fully connected neural network with ReLU activation function and without bias terms, applied to solve a K-label classification problem. Suppose the following is satisfied: \({\text{||}}{{{\mathbf{W}}}^{{(p)}}}{\text{|}}{{{\text{|}}}_{2}}\;\leqslant \;{{M}_{{\mathbf{W}}}}\) and \({\text{||}}{{{\mathbf{x}}}_{i}}{\text{|}}{{{\text{|}}}_{2}}\;\leqslant \;{{M}_{{\mathbf{x}}}}\) for all layers \(p = 1, \ldots ,L\) in network and for all objects \(i = 1, \ldots ,m\) in the dataset. Then, for any object \(i = 1, \ldots ,m\) the following inequality holds:

$${{\left\| {{{{\mathbf{H}}}_{i}}({\boldsymbol{\theta }})} \right\|}_{2}}\;\leqslant \;L\sqrt 2 M_{{\mathbf{x}}}^{2}M_{{\mathbf{W}}}^{{2L}} + \sqrt 2 \frac{{M_{{\mathbf{W}}}^{2}(M_{{\mathbf{W}}}^{{2L}} - 1)}}{{M_{{\mathbf{W}}}^{2} - 1}}.$$

This theorem allows us to understand the dependence of the Hessian norm on the structure of the neural network: the size of the hidden layer and the number of layers. So, the next Lemma 1 exhibits the dependence on the hidden size.

Lemma 1. If each model parameter is bounded by a constant \(M > 0\), that is \({\text{|}}w_{{ij}}^{{(p)}}{\text{|}}\;\leqslant \;M\) for all \(i,j = 1, \ldots ,h\) and for all layers \(p = 1, \ldots ,L\), then, under the conditions of Theorem 1, the following is true:

$${{\left\| {{{{\mathbf{H}}}_{i}}({\boldsymbol{\theta }})} \right\|}_{2}}\;\leqslant \;L\sqrt 2 M_{{\mathbf{x}}}^{2}{{(hM)}^{{2L}}} + \sqrt 2 \frac{{{{{(hM)}}^{2}}((hM{{)}^{{2L}}} - 1)}}{{{{{(hM)}}^{2}} - 1}}.$$

So, the following proportionality holds:

$${{\left\| {{{{\mathbf{H}}}_{i}}(\boldsymbol\theta )} \right\|}_{2}} \propto L{{(hM)}^{{2L}}}.$$

Obtained results allows us to claim that the Hessian norm is a power function of the size of the hidden layer \(h\) and an exponential function of the number of layers L. Although it may seem that the estimate received is too high, this is actually not the case. The fact is that if we choose \(h\) to be large, then the limiting constant \(M\) will most likely be very small. Because of this, the number under the power of \(2L\) will probably be less than one. Further, we use this upper bound to get an inequality for the loss function difference.

4.2 Losses Difference Convergence

Below there is a Lemma 2, a detailed proof of which is provided in Appendix A.4.

Lemma 2. Let θ be chosen as \({\text{||}}{\boldsymbol{\theta }} - {\boldsymbol{\theta }}{\kern 1pt} \text{*}{\text{||}}_{2}^{2}\;\leqslant \;{{R}^{2}}\) for some \(R > 0\). If there exist a non-negative constant \({{M}_{\ell }}\) such \(\left| {\ell ({{f}_{{{\boldsymbol{\theta }}{\kern 1pt} \text{*}}}}({{{\mathbf{x}}}_{i}}),{{{\mathbf{y}}}_{i}})} \right|\;\leqslant \;{{M}_{\ell }}\) for all objects \(i = 1, \ldots ,m\) in the dataset, then, under the conditions of Theorem 1, the following holds:

$$\begin{gathered} \left| {{{\mathcal{L}}_{{k + 1}}}({\boldsymbol{\theta }}) - {{\mathcal{L}}_{k}}({\boldsymbol{\theta }})} \right| \\ \leqslant \frac{2}{{k + 1}}\left( {{{M}_{\ell }}\, + \,\left( {L\sqrt 2 M_{{\mathbf{x}}}^{2}M_{{\mathbf{W}}}^{{2L}}\, + \,\sqrt 2 \frac{{M_{{\mathbf{W}}}^{2}(M_{{\mathbf{W}}}^{{2L}}\, - \,1)}}{{M_{{\mathbf{W}}}^{2} - 1}}} \right){{R}^{2}}} \right)\, \to \,0 \\ {\text{as}}\quad k \to \infty . \\ \end{gathered} $$

So, the following proportionality is true:

$$\left| {{{\mathcal{L}}_{{k + 1}}}({\boldsymbol{\theta }}) - {{\mathcal{L}}_{k}}({\boldsymbol{\theta }})} \right| \propto \frac{{L{{{(hM)}}^{{2L}}}{{R}^{2}}}}{k}.$$
(3)

Based on the estimates obtained, the following conclusions can be drawn. Firstly, the convergence of the loss function surface is affected by the distance from the extremum point. The farther we are from this point, the slower the convergence may be. Secondly, the situation regarding the number of layers and layer size is similar to that of the Hessian matrix. An increase in the number of layers \(L\) will lead to a higher convergence score, while an increase in layer size h does not necessarily have a negative impact. Again, this is due to the constant that evaluates the magnitude of each element in the weight matrix. Finally, the resulting estimate is inversely proportional to the sample size, exhibiting a sublinear rate of convergence. In the next section, we demonstrate the typical behavior of the loss function surface in practice and compare it with the theoretical one obtained.

5 EXPERIMENTS

To verify the theoretical estimates obtained, we conducted a detailed empirical study. In this section, we present the results from training a fully connected neural network for the Image Classification task. The experiments can be easily reproduced by following the instructions provided in our GitHub repository: https://github.com/kisnikser/landscape-hessian.

The primary objective of these experiments is to empirically confirm the convergence of the loss landscape as the sample size increases. To achieve this, we trained a fully connected neural network on the entire dataset and obtained the corresponding parameters \({{\hat {\boldsymbol\theta }}}\) as a point near the minimum. Subsequently, we examined the relationship between the average loss difference and the available sample size.

We utilized the pytorch library [42] as the Python framework for neural network training. The architecture employed is consistent with that described in Section 3, consisting of several linear layers with a ReLU activation function after each layer, except the final one. The size h was fixed for all hidden layers L. The network was trained over numerous epochs using the Adam optimizer [43] with a constant learning rate of 10–3. Various Image Classification datasets available in the torchvision library were used. While the graphs presented below are based on a single dataset, additional results for each dataset can be found in Appendix A.1. For the training process, a batch size of 64 was selected.

5.1 Direct Image Classification

In this experiment, we utilized the pixel values of images as inputs. Figure 2 displays the results obtained from an analysis of 10  000 objects from the MNIST dataset [44], with the network trained over 10 epochs. The corresponding input size is 784, while the output size is 10. The plots on the left were generated by fixing the number of layers at \(L = 5\) in the network. The hidden size across all layers was varied from 4 to 64. Concurrently, the figure on the right illustrates the behavior of the loss difference as the number of hidden layers changes from 1 to 10, with the hidden size h fixed at 16. This sequence was repeated 100 times for averaging. An exponential moving average with a smoothing factor of 0.99 was applied to the obtained results.

Fig. 2.
Fig. 2.
Full size image

The dependence of the absolute value of the loss function difference on the available sample size, direct image classification. The graphs on the left show a decrease in values as the dimension of the hidden layer increases. The graphs on the right show an increase in values as the number of layers increases.

Fig. 3.
Fig. 3.
Full size image

The dependence of the absolute value of the loss function difference on the available sample size, image features extraction. The graphs on the left show a decrease in values as the dimension of the hidden layer increases. The graphs on the right show an increase in values as the number of layers increases.

Fig. 4.
Fig. 4.
Full size image

The dependence of the absolute value of the loss function difference on the available sample size, direct image classification. The graphs on the left show a decrease in values as the dimension of the hidden layer increases. The graphs on the right show an increase in values as the number of layers increases. Results on different datasets: FashionMNIST, CIFAR10, and CIFAR100.

Fig. 5.
Fig. 5.
Full size image

The dependence of the absolute value of the loss function difference on the available sample size, image features extraction. The graphs on the left show a decrease in values as the dimension of the hidden layer increases. The graphs on the right show an increase in values as the number of layers increases. Results on different datasets: FashionMNIST, CIFAR10, and CIFAR100.

From the dependencies observed, it is evident that although the change is not substantial, adding more layers leads to a greater difference in the loss functions (see the right side). Conversely, increasing the hidden size results in a smaller difference between the loss functions (see the left side).

For readers unfamiliar with the topic, this may be surprising, as the estimation we have made (3) suggests a power dependence on the layer size h, and the value of \(L\) is clearly not less than 1. However, readers are referred to Section 4.2 for a more detailed discussion of this phenomenon. Additionally, in practice, the constant \(M\) that limits the magnitude of weights is found to be relatively small. Furthermore, since the MNIST dataset classification task is considered relatively straightforward, a shallow yet wide neural network can produce good classification results. Consequently, the values of the loss function have been observed to be lower for larger h values, and therefore their difference is also lower.

5.2 Image Features Extraction

In contrast to the previous experiment, this part employs a pre-trained image feature extractor. The fully connected network is utilized as a multi-label classification head. We selected the Vision Transformer (ViT) [45] from Google.

We similarly selected 10  000 objects randomly from the MNIST dataset and varied the hidden size and the number of layers. The results are consistent with those observed in direct image classification. This consistency confirms that the convergence presented does not depend on the nature of the space of the original objects \(\mathcal{X}\). Boundedness of this space is sufficient to observe the convergence of the loss function landscape.

Our experiment corroborates the convergence proved in Lemma 2. Additionally, the upper bound on the rate of this convergence holds true. Indeed, altering the parameters of the neural network, such as the number of layers and the layer size, leads to a slight change in the difference of the loss functions. We remind the reader that a larger number of graphs can be found in Appendix A.1.

6 DISCUSSION

The results of this study provide insights into the loss landscape convergence as the dataset size increases. Our theoretical analysis shows that the absolute difference between the average loss function values when adding one more object to the sample tends to zero, as the number of available objects tends to infinity. This was achieved by proving the upper-bound theorem for the Hessian norm in a fully connected neural network. Empirical study allows us to confirm our results practically. In particular, we claim that the loss function surface exhibits convergence for the Image Classification task, both as a direct classifier of initial representations and as a multi-label classification head after the pre-trained feature extractor.

The results of our research are highly connected to the problem of the local geometry of neural loss landscapes. Despite the fact that a large number of studies have been devoted to this issue, the change in dataset size has remained a significant gap. In this paper, we have tried to take the first steps in this direction.

Nevertheless, this study has potential limitations. First, our theoretical analysis was deterministic and not probabilistic. Although this may bring certain clarifications to the estimates, we do not think it can have a serious impact in practice. Second, the subject of our research is a fully connected neural network. We plan to extend our results to other architectures in future work. Third, using Assumption 1, we suppose existing such a point, which will be a minimum, starting with a certain sample size. Finally, using the triangle inequality in the proof of Lemma 2 yields a rough upper bound for the loss difference, so it may be improved in future work.

We believe that our findings will contribute to the development of more precise studies of the behavior of the loss landscape when changing the training sample size. We also believe that our results can be used to develop modern sample-size determination techniques. We expect this because the convergence of the loss function landscape can be considered as the sign that the training sample size is sufficient for the selected model. In future work, we hope to apply our observations to this field.

7 CONCLUSIONS

In this paper, we have presented a comprehensive study of the convergence of the loss landscape in a fully connected neural network as the sample size increases. Our theoretical analysis and empirical results demonstrate that the absolute difference between the average loss function values when adding one more object to the sample tends to zero as the number of available objects tends to infinity. These findings provide valuable insights into the local geometry of neural loss landscapes and address a previously unexplored issue in the field. We believe that our results will contribute to the development of more precise studies of the loss landscape behavior and have implications for the development of sample size determination techniques. Future work will focus on extending our results to other architectures and improving the upper bounds for the loss difference.