arXiv:2509.24654v2

Limiting behaviour of pattern counts in biased binary strings

Jon V. Kogan, Nicolò Paviato

math.PRmath.DS60G5528A3537A40

Abstract

For p(0,1)p \in (0,1), sample a binary sequence from the infinite product measure of Bernoulli(p)(p) distributions. It is known that for p=1/2p=1/2, almost every binary sequence is Poisson generic in the sense of Peres and Weiss, a property that reflects a specific statistical pattern in the frequency of finite substrings. However, this behaviour is highly exceptional: it fails for any p1/2p \ne 1/2. In these other cases, we show that the frequency of substrings of almost every sequence has either trivial or peculiar behaviour. Nevertheless, the Poisson limiting regime can be recovered if one restricts attention to substrings with a fixed number of successes in the Bernoulli(p)(p) trials.

AI-generated audit

Audit summary

Audited against arXiv v2

Not a correctness certificate. A “Correct” result may include yellow typos or minor formal corrections that do not affect substantive soundness. It means this audit found no unresolved substantive error under the stated criteria; it does not replace expert scrutiny or formal verification.

Current report

Detailed mathematical audit

Generated August 18, 2026
01Statements3 reported findingsCorrect

The phase diagram and limiting-law statements are supported. Theorem 1.1 gives the quenched zero, concentration, and critical regimes for occurrences of a random biased pattern. Theorem 1.3 gives the conditional Poisson approximation at critical scaling. Corollary 1.2's subsequence/nonconvergence alternatives follow from the possible limits of the conditional mean and exhaust the normalizations listed there.

Theorem 1.1Correct

Quenched phase diagram for pattern counts in biased binary strings

Pages 3–5 · Theorems 1.1, 1.3 and Corollary 1.2 · arXiv:2509.24654v2

Theorem 1.1 gives the almost-sure phase diagram for the number of occurrences of a random pattern in an independent biased binary string. The regimes are determined by the exponential scale of the conditional pattern probability relative to the observation-window length, exactly as stated. In the subcritical regime the count vanishes, in the supercritical regime the normalized count concentrates, and at the critical scale the theorem records the nondegenerate behavior. The bias endpoints and pattern-length growth restrictions match those used in the overlap estimates.

Theorem 1.3Correct

Conditional Poisson limit at critical scaling

Pages 4–5 and 20–27 · Theorem 1.3 · arXiv:2509.24654v2

Given the random pattern, the mean occurrence count is the number of admissible starting positions times its exact Bernoulli word probability. When this mean converges to λ\lambda, the dependency-neighbourhood bound for overlapping windows tends to zero almost surely, so Chen–Stein gives total-variation convergence to Poisson(λ)\operatorname{Poisson}(\lambda). Boundary starting positions contribute o(1)o(1), and the theorem's growth hypothesis makes the estimate uniform in the stated bias range.

Corollary 1.2Correct

Subsequence alternatives and failure of a global limit

Pages 3–4 and 27–31 · Corollary 1.2 · arXiv:2509.24654v2

The logarithm of the conditional mean is expressed through the empirical symbol frequency of the pattern. Its law of large numbers and iterated fluctuations produce subsequences tending to zero, infinity, or finite positive values in exactly the parameter cases listed. Theorem 1.1 controls the first two and Theorem 1.3 the finite case, showing that no single limiting distribution can exist where the corollary claims nonconvergence.

02Proofs3 reported findingsCorrect

Concentration, overlap estimates, and subsequence analysis. Conditioning on the sampled pattern makes occurrence indicators independent except when their windows overlap. The autocorrelation/overlap sum is bounded using the pattern's typical symbol frequencies and is negligible on the critical scale. Chen–Stein then compares the conditional count with a Poisson law whose parameter is the window length times the pattern probability. Away from criticality, first/second moment and exponential concentration bounds give the zero or law-of-large-numbers regimes. The exceptional probabilities are summable on the selected lengths and monotonic interpolation covers all intermediate indices.

Annealed-to-quenched transferCorrect and complete

Concentration, overlap estimates, and subsequence analysis

Pages 9–31 · Sections 3–5 · arXiv:2509.24654v2

Conditioning on the sampled pattern makes occurrence indicators independent except when their windows overlap. The autocorrelation/overlap sum is bounded using the pattern's typical symbol frequencies and is negligible on the critical scale. Chen–Stein then compares the conditional count with a Poisson law whose parameter is the window length times the pattern probability. Away from criticality, first/second moment and exponential concentration bounds give the zero or law-of-large-numbers regimes. The exceptional probabilities are summable on the selected lengths and monotonic interpolation covers all intermediate indices.

Sections 3–4Correct and complete

Overlap estimates and Chen–Stein approximation

Pages 9–23 · Sections 3–4 · arXiv:2509.24654v2

Occurrence indicators have dependency neighbourhood of radius equal to the pattern length. For each possible overlap, simultaneous occurrence forces a border relation in the pattern; the probability of atypically many such borders is summable. On the typical event, the two Chen–Stein error sums are o(1). Conditioning and then applying Borel–Cantelli makes the bound quenched for almost every pattern.

Section 5Correct and complete

Concentration and subsequence analysis

Pages 23–31 · Section 5 · arXiv:2509.24654v2

Truncating to nonoverlapping starting positions gives independent Bernoulli blocks and exponential concentration around the conditional mean. The discarded positions are controlled by the same overlap sum. Large-deviation bounds for the pattern's symbol frequency are summable along geometric scales, and monotonicity fills the gaps. The subsequence choices are then made from the exact logarithmic mean, so all phase alternatives are covered.

03Novelty0 reported findingsNo non-novelty findings

No non-novelty findings.

Detailed audit reportFull reasoning, manuscript locations, and references.
Open report PDF ↗

Author response

Challenge an audit finding

Local workflow preview

A listed author may submit formal evidence that an audit is inaccurate. The response would be considered in a fresh AI re-evaluation; it would not edit the audit automatically.

Paper
arXiv:2509.24654v2
Authors listed
Jon V. Kogan, Nicolò Paviato
Audit date
August 18, 2026
  1. 01Establish identityMatch an authenticated scholarly identity to this paper.
  2. 02Submit evidenceIdentify the finding and give a formal mathematical response.
  3. 03Re-evaluateA separate agent checks the response and records a disposition.
Recommended production method

Authenticate with ORCID, then require an exact arXiv match

MathAudit should accept the identity only when ORCID OAuth authenticates the claimant's iD and this exact arXiv paper appears in arXiv's public authority feed for that iD. A matching name alone is not sufficient.

ORCID OAuth and arXiv authority-record lookup are not connected in this local prototype.

Email fallback for papers without a linked ORCID

A production fallback could send a one-time link only when the submitted address matches an independently maintained author-contact allowlist for this paper. MathAudit must return the same message for every address so the form cannot reveal which contacts are on that list.

This demonstration does not send, store, or compare the address.

Structured response preview

This form remains unavailable until production identity verification succeeds. Nothing entered here is submitted.

This panel never establishes authorship in the local prototype. A production result should be described narrowly as an authenticated ORCID match or control of a separately allowlisted author-contact mailbox.