arXiv:2608.06155v1
Abstract
Conditional expectation operators (CEOs) and their associated conditional mean embeddings (CMEs) play a central role across applied mathematics and machine learning, appearing in nonparametric regression, Bayesian inverse problems, and Koopman operator theory. A fundamental question is when a CEO maps a function space on into a prescribed function space on , particularly a reproducing kernel Hilbert space (RKHS). We show that such mapping properties are characterized by the regularity of the Radon--Nikodym density of the conditional law, and establish a simple, verifiable sufficient condition under which the CEO is bounded and Hilbert--Schmidt. For RKHSs norm-equivalent to Sobolev spaces, this condition reduces to Sobolev regularity of the conditional density. The result yields a direct route to validate CME representations and error bounds for Galerkin-type and CME-based estimators. We verify the regularity condition in three settings: nonparametric regression, Bayesian inverse problems, and Koopman operator theory for stochastic dynamical systems. We show in each case that classical regularity results on the underlying probabilistic model imply the required mapping properties. The resulting framework offers a unified perspective on conditional expectation operators across probability, operator theory, kernel methods, and stochastic dynamics.
AI-generated audit
Audit summary
Not a correctness certificate. A “Correct” result may include yellow typos or minor formal corrections that do not affect substantive soundness. It means this audit found no unresolved substantive error under the stated criteria; it does not replace expert scrutiny or formal verification.
Current report
Detailed mathematical audit
01Statements4 reported findingsContains unsupported statements
The Hilbert--Schmidt regularity criterion, the abstract projection-error bounds, and the three application theorems are verified. The general learning-rate theorem is also verified under its explicit covariance-eigenvalue hypothesis. Its final claim that the Sobolev setting automatically gives the exponent is not established for an arbitrary input law : the proof substitutes approximation numbers for the Lebesgue embedding into an operator whose codomain is .
Regular conditional densities yield Hilbert--Schmidt CEOs and CMEs
Pages 7–11 · Proposition 3.1, Theorem 3.3, and Corollary 3.8 · arXiv:2608.06155v1
If belongs to for almost every and its squared RKHS norm is integrable, the Bochner kernel lies in . Proposition 3.1 therefore gives a Hilbert--Schmidt integral operator. The conditional-expectation identity verifies that this operator is the CEO, and applying it to gives the stated CME representation. The boundedness, compactness, and trace-class consequences follow from standard Hilbert--Schmidt operator identities.
Projection and learning rates under an assumed eigenvalue bound
Pages 12–17 · Theorems 3.12 and 3.16 · arXiv:2608.06155v1
The range condition gives the required factorization through the fractional covariance power. The projection residual is then controlled by the corresponding power-function or interpolation estimate. Under the separately stated decay , the effective-dimension calculation and the sampling estimate yield the displayed rate after the chosen regularization balance. These conclusions do not depend on the automatic Sobolev-exponent clause discussed separately below.
The exponent is not automatic for an arbitrary input distribution
Pages 16–17 · final paragraph of Theorem 3.16 and its proof · arXiv:2608.06155v1
The theorem assumes only that on a bounded Lipschitz domain and then states that the covariance eigenvalues satisfy . But , where is the embedding determined by the arbitrary probability law . The proof cites the classical approximation-number rate for with Lebesgue measure and does not supply a comparison between these two codomains. Consequently the claimed automatic exponent, and hence the unconditional specialization of the learning rate, is not verified under the printed assumptions. A verified sufficient repair for the upper eigenvalue bound actually used in the rate is to assume has an essentially bounded density with respect to Lebesgue measure; a two-sided claim additionally needs an appropriate lower comparison. No counterexample to the learning-rate conclusion itself is asserted.
Regression, Bayesian inverse, and elliptic-diffusion verification results
Pages 18–32 · Section 4 · arXiv:2608.06155v1
In the regression and Bayesian settings, the stated differentiability and integrability hypotheses permit the required Sobolev estimates for the conditional density. For the uniformly elliptic diffusion, the cited transition-density regularity and the compact-state-space bounds give the required square-integrable Sobolev norm. Substitution into Theorem 3.3 then yields the advertised Hilbert--Schmidt CEO and CME conclusions in each application.
02Proofs5 reported findingsContains incorrect or incomplete proofs
The core conditional-density, projection, and application proofs are correct and complete. The final Sobolev-rate specialization contains an unresolved change from to Lebesgue , so that portion of Theorem 3.16 is incomplete. The finite-sample projection formulas also print ordinary inverses where the already-defined Moore--Penrose pseudoinverse is required; this is a harmless notation typo.
Hilbert--Schmidt kernel and conditional-expectation argument
Pages 7–11 · Section 3.1 · arXiv:2608.06155v1
The paper verifies strong measurability and square integrability of the Hilbert-space-valued density kernel, identifies its integral operator with conditional expectation by testing against functions of , and then evaluates the operator on kernel sections. The norm and trace identities used downstream are valid for a Hilbert--Schmidt operator.
Fractional-range factorization and Galerkin error estimate
Pages 12–15 · Theorem 3.12 and proof · arXiv:2608.06155v1
The Douglas-type range factorization, spectral calculus for , and projection-residual estimate are applied with compatible domains. The separate regularity regimes in the theorem match the exponents used in the interpolation bounds, and the finite-rank estimator inherits the displayed operator-norm control.
Lebesgue approximation numbers are used for an embedding
Pages 16–17 · last step of the proof of Theorem 3.16 · arXiv:2608.06155v1
For the covariance operator, the relevant singular values are those of . The proof instead invokes and immediately concludes . Without a measure-comparison hypothesis this does not follow. Downstream dependency: only the final automatic choice and its specialized rate; the theorem under an assumed covariance-eigenvalue bound remains verified. Repair classification: verified sufficient repair for the needed upper bound. Add , so and the approximation-number upper rate transfers. Require a corresponding lower density bound if retaining the printed two-sided equivalence.
Gram-matrix inverses should be pseudoinverses
Pages 12–14 · Notation 3.10, Table 1, and the displayed empirical projection formulas · arXiv:2608.06155v1
The points are not assumed distinct and the kernels are not assumed strictly positive definite, so and need not be invertible. The paper has already defined the Moore--Penrose pseudoinverse. Replace and in these projection formulas by and . This is the standard orthogonal-projection formula and leaves every abstract result unchanged.
Verification of the regularity hypotheses in the three applications
Pages 18–32 · proofs of Theorems 4.2, 4.5, and 4.11 · arXiv:2608.06155v1
Each application derives the claimed conditional-density regularity from its explicit model assumptions, checks the needed integrability uniformly over the conditioning variable, and then invokes the abstract criterion with matching RKHS/Sobolev norms. No missing case or unsupported implication was found in these proof chains.
03Novelty0 reported findingsNo non-novelty findings
No non-novelty findings.