This note introduces the main proof in Chapters 3–6. The goal is to explain why a convex viscosity solution $u$ on an open convex domain $\Omega$ satisfying
The proof joins the dimension of supporting contact sets, the short axes of John ellipsoids, volume decay of strict sections, and Mooney’s covering theorem into one closed chain. The critical contact dimension is $q=n-k+1$.
For $x\in\Omega$ and $p\in\partial u(x)$, let $\ell_{x,p}$ be the corresponding supporting affine function. Write $d_u(x,p)$ for the stabilized affine dimension, at small scales, of the local contact set $\{u=\ell_{x,p}\}$. Define
$$C_m(u)=\left\{x:d_u(x,p)\ge m\ \text{for every }p\in\partial u(x)\right\}.$$
The words “for every supporting slope” are essential because the covering theorem in Chapter 6 requires volume decay for every section at the point.
Chapter 3: if one supporting slope has contact dimension at most $n-k$, then $u$ is smooth nearby.
Chapter 4: positive $k$-Hessian density controls the product of the shortest $k$ axes of the John ellipsoid of a small sublevel set.
Chapter 5: $m$-dimensional contact and quadratic localization together give $|S_h|\lesssim h^{(m+k)/2}$.
2. First engine: small contact sets imply local smoothness
The approximation, interior second-derivative estimate, and Evans–Krylov steps are standard; the device that actually starts the argument is a saddle surface. After subtracting a supporting affine function, normalize $u(0)=0$ and $u\ge0$. Suppose $\{u=0\}$ lies in a subspace $Z$ of dimension at most $n-k$ in a small ball. Split $\mathbb R^n=Y\oplus Z$. Then $\dim Y\ge k$; choose $k$ directions in $Y$ and construct
$$Q(y,z)=A|y|^2-|z|^2+\alpha r^2.$$
The two families of directions have opposite but complementary jobs. At least $k$ strongly positive curvatures in $Y$ ensure $D^2Q\in\Gamma_k$, while the negative curvature along the contact space $Z$ pushes $Q$ below $u$ on the spherical boundary. Choosing $A$ and $\alpha$, then scaling, gives both $Q<u$ on the boundary and $Q>u$ near the center. A purely convex paraboloid cannot provide this geometry.
For smooth $k$-admissible CNS approximations $u_j$, choose a regular value $t_j$ and take the component containing the origin of
$$D_j=\{Q-u_j>t_j\}.$$
The boundary/center sign change traps $D_j$ inside the small ball, and the regular value gives a smooth moving boundary. The localized Chou–Wang estimate then yields a Hessian bound independent of $j$ on $D_j$. Evans–Krylov for the concave operator $F=\sigma_k^{1/k}$, followed by Schauder estimates, yields local smoothness of the limit $u$.
Uniform convergence $u_j\to u$ does not by itself produce a $C^2$ limit. The genuine regularizing step is the moving domain cut out by the saddle surface together with the Chou–Wang Hessian estimate.
图1。 马鞍面的两个方向:接触方向中的负曲率负责封住边界,至少 $k$ 个横向正曲率负责保持 $k$-admissibility。图面可用鼠标拖动旋转,并用滚轮缩放。Figure 1. The two directions of the saddle surface: negative curvature in the contact directions closes the boundary, while at least $k$ positive transverse curvatures preserve $k$-admissibility. Drag to rotate and use the wheel to zoom.
3. Second engine: John ellipsoids and an interior viscosity test
The strict section $K_h=\{u\le h\}\cap\overline B_R$ is convex, but its boundary may be complicated, its principal directions may rotate with scale, and its directional sizes may be wildly different. The proof does not assume that the section itself is an ellipsoid. John’s theorem compresses this geometric complexity into one inscribed ellipsoid, hence $n$ semiaxes and their principal directions:
Order the semiaxes of $A_h$ as $a_1\ge\cdots\ge a_n$. Up to the fixed dimensional factor $n$, the full convex-geometric complexity of the section has now become bookkeeping for these $n$ ordered axes. Construct the test function
$$P_h(x)=h\left|A_h^{-1}(x-z_h)\right|^2.$$
Set $\mu_h=\min(P_h-u)$. The translate $P_h-\mu_h$ touches $u$ from above at an auxiliary interior point $y_h$. Since $D^2P_h=2hA_h^{-2}>0$, the viscosity inequality gives
This is an upper bound for the product of the shortest $k$ John axes. The equation controls their product, not each axis separately.
图2。 第4章内部黏性测试的几何结构。抬高平面显示 graph cap;$K_h$ 与 $E_h$ 位于自变量空间,测试点 $y_h$ 由最小化 $P_h-u$ 自动产生。Figure 2. The geometry of the interior viscosity test in Chapter 4. The lifted plane displays the graph cap; $K_h$ and $E_h$ lie in the variable space, and $y_h$ is generated by minimizing $P_h-u$.
4. 第三个引擎:接触维数与 section 体积
这一节的核心不是把 $K_h$ 想象成某个规则形状,而是把上一节得到的有序 John 轴分成三段。记 $q=n-k+1$。测试函数与黏性方程给出最短 $k$ 根轴的乘积上界
4. Third engine: contact dimension and section volume
The point is not to imagine $K_h$ as a regular shape, but to split the ordered John axes from the previous section into three ranges. Put $q=n-k+1$. The test function and viscosity equation give the product bound for the shortest $k$ axes,
$$a_q\cdots a_n\lesssim \lambda^{-1/2}h^{k/2}.$$
On the other hand, if the supporting contact set contains a relative $m$-ball of radius $\rho$, the upper John inclusion and a width min-max argument give $a_m\gtrsim\rho$. Since $a_1\ge\cdots\ge a_n$, this means
$$a_1,\ldots,a_m\gtrsim\rho.$$
When $m\ge q$, the two index ranges genuinely overlap: $a_q,\ldots,a_m$ belong both to the contact-lower-bound range and to the PDE-controlled product. Dividing these overlap factors out of the short-axis product yields
This forces rapid collective decay of the shortest $n-m$ axes after the overlap. A gap remains: the first $m$ axes have lower bounds but may still be arbitrarily long, so the last display alone does not control the volume of the whole section. The repair is to set
$$v(x)=u(x)+\frac{|x|^2}{2}.$$
Adding this smooth quadratic leaves the singular set unchanged: $u$ is locally $C^2$ at a point if and only if $v$ is. The $k$-Hessian equation is not applied to $v$; every PDE product estimate still comes from $u$. The quadratic term is used only for geometric localization. For the corresponding strict section, $v(x)\le h$ implies $|x|^2/2\le h$, hence
$$S_h^v\subset K_h\cap B_{\sqrt{2h}}.$$
This smooth lift truncates all potentially long first $m$ contact directions to $O(\sqrt h)$, contributing $h^{m/2}$; the remaining short-axis product after cancellation contributes $h^{k/2}$. Combining the two pieces gives
图3。 半轴按 $a_1\ge\cdots\ge a_n$ 排列。上括号覆盖 $q\le i\le n$,下括号覆盖 $1\le i\le m$,所以 $q\le i\le m$ 是两种控制的重合区间;着色的 $m+1\le i\le n$ 正是消去重合后要处理的短轴。Figure 3. Order the semiaxes as $a_1\ge\cdots\ge a_n$. The upper brace spans $q\le i\le n$ and the lower brace spans $1\le i\le m$, so $q\le i\le m$ is their overlap; the shaded range $m+1\le i\le n$ contains the short axes left after cancellation.图4。 Proposition 5.1 的几何顺序:接触片、$K_h$、John 轴下界以及 $S_h^v\subset K_h\cap B_{\sqrt{2h}}$。Figure 4. The geometric order in Proposition 5.1: the contact patch, $K_h$, the John-axis lower bound, and $S_h^v\subset K_h\cap B_{\sqrt{2h}}$.
The covering theorem allows the constants and height thresholds to depend on $(x,p)$, so the absence of a uniform lower bound for the contact radius $\rho$ does not obstruct the argument. Full-dimensional contact, on the other hand, would produce an upper quadratic test of arbitrarily small curvature, contradicting $\sigma_k\ge\lambda$. Hence $C_n(u)=\varnothing$.
Chapter 3 gives $\operatorname{Sing}(u)\subset C_q(u)$. Applying the stratification result at $m=q=n-k+1$ yields $\mathcal H^{n-1}_{\mathrm{loc}}(\operatorname{Sing}(u))=0$.
Here $q=2$. On the critical stratum $m=2$, contact geometry gives $a_2\gtrsim\rho$, while the PDE gives $a_2a_3a_4\lesssim h^{3/2}$. After adding the quadratic term,
图5。 $n=4$、$k=3$ 时,$m=2$ 与 $m=3$ 两层的 John 轴乘积、section 体积指数和 Hausdorff 指数。Figure 5. For $n=4$ and $k=3$, the John-axis products, section-volume exponents, and Hausdorff exponents on the $m=2$ and $m=3$ strata.
$k=n$: $q=1$ and the conclusion recovers the Monge–Ampère stratification $\mathcal H^{n-m}(C_m)=0$.
$k=2$: this method gives codimension-one nullity for the critical contact stratum; full smoothness needs the additional input of strict $2$-convexity.
Axis control: the equation controls only a product of short axes. Principal directions may rotate and redistribute with scale, so the conclusion is Hausdorff nullity rather than a uniform Minkowski packing estimate.
关于本文。 这是一篇介绍 Xiyu Hu, Sharp Hausdorff Bounds for the Interior Singular Set of Convex k-Hessian Solutions 第3–6章的证明导读 note。论文第8章的 mean-value 公式提供概念动机,但不作为上述主证明链的输入。
About this note. This is a proof-guide note to Chapters 3–6 of Xiyu Hu, Sharp Hausdorff Bounds for the Interior Singular Set of Convex k-Hessian Solutions. The mean-value formula in Chapter 8 provides conceptual motivation but is not an input to the proof chain above.
Two extremal problems can share a proof architecture without sharing a proof. Benjamin Bedert’s 2025 breakthrough on large sum-free subsets is additive and Fourier-analytic. A July 2026 manuscript on strongly $2$-primitive sets is multiplicative and hypergraph-theoretic. Their common language is not a transferable sieve, but a four-step design: localize, build a strong object in each local block, prevent leakage between blocks, and add the gains. That comparison points toward the right stability question – and also reveals why the most naive version of stability is false.
Status and scope. Bedert’s theorem is the February 2025 arXiv preprint cited below. The $27/2$ theorem is contained in a five-page manuscript by Przemek Chojecki supplied for this post in July 2026; no public preprint or peer-reviewed version was found, and the Erdős Problems page for #793 still labels the problem open at the time of writing. The written proof was checked here for internal coherence, but that is not independent peer review. The factor-graph and normalized-stability results later in this article are the additional analysis developed in the accompanying long note. Conditional and open statements are labelled explicitly.
Here is the headline before the details. The supplied manuscript claims
where $F(n)$ is the largest size of a strongly $2$-primitive subset of $[1,n]$. One might expect every near-extremizer to be close to the construction behind this formula: almost all primes, plus products of three primes near $n^{1/3}$ arranged as a nearly saturated linear $3$-uniform hypergraph. That literal statement is false. There is, however, an exact and useful replacement: after contracting harmless private-prime lifts, every near-extremizer has a canonical prime layer and a small exceptional hub carrying the entire second-order excess. If the residual hub is mostly made of squarefree triples, then pair saturation and $n^{1/3}$-scale stability follow. Proving – or refuting – that triple-core hypothesis is the remaining inverse problem.
1. Two problems that look similar only from far away
1.1 The sum-free extraction problem
A set $B\subset\mathbb Z$ is sum-free if it contains no $x,y,z$, not necessarily distinct, with
$$x+y=z.$$
For a finite set $A$ of integers, write
$$
S(A)=\max\{|B|:B\subseteq A,\ B\text{ is sum-free}\},
$$
and, for $N$-element sets of positive integers,
$$
S(N)=\min_{|A|=N}S(A).
$$
This is an extraction problem. An adversary gives us an arbitrary host set $A$, and we must find a large structured subset inside it. Erdős’s middle-third argument gives $S(A)\ge |A|/3$. Alon and Kleitman improved this to $(N+1)/3$, and Jean Bourgain proved $S(N)\ge (N+2)/3$ in 1997. The long-standing qualitative question was whether the additive improvement can tend to infinity:
Bedert answered yes, proving that an absolute $c\gt0$ exists such that
$$
S(A)\ge \frac{|A|}{3}+c\log\log |A|.
$$
1.2 The strongly $2$-primitive packing problem
A set $A\subseteq[1,n]$ is strongly $2$-primitive when
$$
a\nmid bc
\qquad
(a,b,c\in A,\ a\notin\{b,c\}),
$$
where $b=c$ is allowed. The word strongly matters. Under a more recent convention, a $2$-primitive set only forbids witnesses $b,c$ that are distinct. For example, $\{4,5,6\}$ passes that weaker test but fails the strong one because $4\mid6^2$.
Now define
$$
F(n)=\max\{|A|:A\subseteq[1,n],\ A\text{ is strongly }2\text{-primitive}\}.
$$
This is a packing problem. The host interval is fixed, and we directly construct the largest possible forbidden-divisibility family. The all-primes set gives the leading term $\pi(n)$. In 1938 Erdős proved upper and lower bounds of the form
and later asked whether the second-order term has an asymptotic constant. This is the modern Erdős Problem #793. The supplied 2026 manuscript proposes that the constant is $27/2$; it is important not to say that Erdős himself conjectured this numerical value.
Figure 1. The quantifiers already separate the two questions. Sum-free theory extracts a subset from an arbitrary host; the multiplicative problem packs a forbidden-divisibility family into a fixed interval.
1.3 Equality versus order
The relation $a\nmid bc$ is not ordinary product-freeness. The set $\{6,10,15\}$, for instance, has no internal equality $xy=z$, yet $6\mid10\cdot15$. Prime valuations expose the real geometry:
$$
a\mid bc
\quad\Longleftrightarrow\quad
v_p(a)\le v_p(b)+v_p(c)
\quad\text{for every prime }p.
$$
Thus the forbidden relation is a coordinatewise domination inequality in the divisor lattice. For squarefree integers, if $E_a=\{p:p\mid a\}$, it becomes
$$E_a\subseteq E_b\cup E_c.$$
So the natural combinatorial object is a $2$-cover-free family, not the solution set of a linear equation. Fourier characters are superb at detecting equations such as $x+y=z$. They do not come with an evident contractive projection that detects the partial order $\nu(a)\le\nu(b)+\nu(c)$. This is the first reason Bedert’s proof cannot simply be copied into the multiplicative setting.
2. The $27/2$ upper bound: every element needs a private factor
The upper bound begins with a small lemma that contains more information than the inequality it proves.
Private-factor lemma. Let $\mathcal B$ be a set of positive integers, and choose for every $a\in A$ a factorization $a=u_av_a$ with $u_a,v_a\in\mathcal B$. If $A$ is strongly $2$-primitive, then $|A|\le|\mathcal B|$.
Think of $E_a=\{u_a,v_a\}$ as a two-element multiset. If $u_a\ne v_a$ and neither coordinate is private to $a$, another chosen pair contains $u_a$ and another contains $v_a$; the corresponding two elements have product divisible by $a$. If $u_a=v_a=x$, failure of privacy would give another pair containing $x$ twice, hence another element equal to $x^2=a$. Therefore every $a$ has a private coordinate, and these coordinates are distinct. That is the injection $A\hookrightarrow\mathcal B$.
The manuscript chooses the multiplicative $2$-basis
Every $m\le n$ can be factored into two elements of $\mathcal B=\mathcal B_0\cup\mathcal B_1\cup\mathcal B_2\cup\mathcal B_3$. The case split is elementary but carefully tuned. Small $m$ can be balanced into two factors below $n^{3/5}$, a prime factor above $n^{2/5}$ can be split off, and the remaining difficult case groups two of the three largest prime factors into an element of $\mathcal B_2$ or $\mathcal B_3$.
The main prime layer comes from $\mathcal B_1$. The genuinely second-order counts are
The private-factor injection then gives the upper bound. A caution that becomes crucial for stability: $\mathcal B_2$ and $\mathcal B_3$ are labels in an upper-bound certificate. Near-saturation of these labels does not immediately say that the original elements of $A$ are themselves products of two or three primes.
3. The lower bound: turn properly coloured edges into prime triples
The lower bound lives in a linear $3$-uniform hypergraph. Let $\mathcal H$ be a family of triples of distinct primes such that any two triples share at most one prime and every edge product is at most $n$. Define
$$
A_{\mathcal H}
=\{p\le n:p\text{ prime and }p\notin V(\mathcal H)\}
\cup
\left\{\prod_{p\in E}p:E\in\mathcal H\right\}.
$$
Why is this strongly $2$-primitive? A target triple product has three distinct prime coordinates. Each other hyperedge supplies at most one of them, so two other elements supply at most two. The same argument still works when the two witnesses coincide. Primes outside the vertex set remain singleton elements and cannot divide any other chosen element.
3.1 Logarithmic prime bins
Fix a small mesh $h\gt0$ and divide primes near $y=n^{1/3}$ into bins
put $k=-i-j-3$. Then $i\le j\lt k$, and any $p\in P_i$, $q\in P_j$, $r\in P_k$ obeys $pqr\le n$.
If $i\lt j$, properly edge-colour the complete bipartite graph between $P_i$ and $P_j$, using primes of $P_k$ as colours. When $i=j$, do the same with the complete graph on $P_i$. The lower pair $\{p,q\}$, coloured by $r$, becomes the hyperedge $\{p,q,r\}$.
Figure 2. Within a cell, every colour class is a matching, so two produced triples never share a lower pair. Across cells, the sorted signature $(i,j,k)$ has fixed sum $-3$; two shared bin indices force the third and hence force the same cell.
The proper colouring is the local no-interference mechanism. The constant-sum signature is the global one. Together they make the union over all cells a linear hypergraph.
3.2 The cell weight
For any fixed finite collection of bins, the prime number theorem gives
$$|P_i|=(3+o(1))M\Delta_i.$$
Off the diagonal, a cell contributes approximately $9M^2\Delta_i\Delta_j$ edges. A diagonal cell contributes half as many unordered pairs. The exact geometric-series identity is
The number of vertices used is only $o(\Sigma_n)$. Replacing those primes by the triple products therefore yields the matching lower bound.
Figure 3. The same constant is visible in the canonical feasible pair space. The labels $9/2$ and $9$ are prime-counting contributions, not Euclidean areas of the drawing.
4. From Erdős’s middle third to Bedert’s $c\log\log N$
We now return to the additive problem. Let $\mathbb T=\mathbb R/\mathbb Z$, and let $\phi=\mathbf1_{(1/3,2/3)}$. The middle third of the circle is sum-free: two points in it cannot add, modulo $1$, to another point in it. Therefore, for every $x\in\mathbb T$,
$$A_x=\{a\in A:ax\pmod1\in(1/3,2/3)\}$$
is sum-free. Averaging $|A_x|$ over $x$ gives $|A|/3$. The problem is to force a positive fluctuation above that mean.
After a harmless normalization, the relevant Fourier series has the form
where $\chi$ is the non-principal real character modulo $3$. A lower bound for $\|F_A\|_1$ gives a one-sided large value and hence a large sum-free subset.
4.1 Where Bourgain’s sieve loses a logarithm
The historical correction is worth making explicit. The relevant predecessor is Jean Bourgain’s 1997 paper, not “John Bogan, 1993.” Also, the Littlewood $L^1$ conjecture had already been proved in 1981, independently by McGehee-Pigno-Smith and Konyagin. Bourgain’s argument combines Fourier analysis with a Möbius operation that filters harmonic indices.
Schematically, if $R_Q$ denotes the positive $Q$-rough integers, then
under the triangle inequality. If $Q$ is large enough to make the rough remainder negligible directly, this factor consumes the logarithm delivered by a Littlewood-type lower bound. This explains why merely “using the Littlewood theorem harder” does not produce the desired unbounded gain.
4.2 Bedert’s split: medium primes are sieved, small primes are projected
Bedert chooses
$$Q_1=(\log N)^{1/2},\qquad Q=(\log N)^{20}.$$
He Möbius-sieves only the medium primes $Q_1\le p\le Q$. Their reciprocal sum is $O(1)$, so the norm loss is a constant rather than a logarithm. The small primes $p\le Q_1$ are handled by a different operation.
For every small prime, record the exact valuation $\nu_p(a)$ and the unit residue after removing that prime power. Chinese remaindering packages all this data into a joint fibre $A(r,\nu)$. Fourier projection to a residue class,
At the main valuation level, the projection removes small-prime harmonic contamination without paying the full Möbius product. After truncation, one sees a large exponential sum plus an $L^2$-small error, to which a robust McGehee-Pigno-Smith test function applies.
At a general valuation level, lower fibres can still leak into the projected channel. Bedert proves an isolation-or-descent alternative: either the desired $L^1$ lower bound is already present, or a large fibre has a strictly lower fibre losing at most a controlled polylogarithmic factor. Iterating and reversing this descent produces a chain
Let $g_i$ be the normalized indicator of the $i$-th residue block and put $Q_i(x)=\exp(-|\widehat g_i(x)|)$. The non-Archimedean MPS test function is assembled from
The decisive observation is that $|\widehat g_i|$ is $1/q_i$-periodic, hence $\widehat Q_i$ is supported on $q_i\mathbb Z$. Multiplying an earlier block by a later $Q_k$ shifts its frequencies only by multiples of $q_k$. Since $q_i\mid q_k$, the earlier block cannot escape its residue class modulo $q_i$. Main inner products contribute one controlled unit per block, while geometric fibre growth makes the cross terms summable.
That is Bedert’s genuine no-leakage mechanism. It is more precise than the slogan “use many congruence classes and patch them.”
One further qualification prevents a common misstatement. Bedert obtains a suitable $F_4$-isomorphic model $B$ of the original set and proves the required $L^1$ bound there. The norm $\|F_A\|_1$ is not claimed to be invariant under an $F_4$-isomorphism. What is preserved is the four-term additive information needed for sum-freeness, so $S(A)=S(B)$.
Bedert’s inverse output is also stronger than the numerical lower bound. If $S(A)\le N/3+C$, his results force low additive dimension, a dense $F_4$-model, large additive energy in every substantial subset, and a “99% Structure Theorem” decomposing all but a small exceptional set into large small-doubling pieces. This is a genuine inverse theorem for host sets resisting sum-free extraction, but not an edit-distance classification by one canonical extremizer.
5. The real bridge: localize, build, prevent leakage, sum
Figure 4. The analogy is architectural. Bedert preserves Fourier residue lanes through nested moduli; the multiplicative construction preserves low codegrees through matching colours and unique scale signatures.
The dictionary is now clean:
Role
Bedert’s sum-free proof
Strongly $2$-primitive proof
Underlying relation
Linear equation in an abelian group
Coordinatewise domination of prime valuations
Local coordinates
Exact small-prime valuations and unit residues
Logarithmic sizes of prime factors
Local block
Joint $p$-adic fibre and MPS block
Scale cell and properly coloured graph
No interference
Nested moduli preserve Fourier support
Colour matchings and unique cell signatures preserve linearity
Accumulation
One controlled inner product per fibre
One triple per admissible lower pair
Inverse output
Additive dimension, energy, small-doubling pieces
Private factors and a cover-free exponent core
This is a substantial connection, but it is a connection of proof design. Bedert’s Fourier projection has no direct analogue for the divisor partial order. The most plausible transfer is therefore not a line-by-line proof but a research program: find a multiplicative localization, an isolation-or-descent alternative, and a no-leakage invariant that survives repeated prime powers and larger supports.
6. The first stability guess is false
The lower construction suggests a tempting statement: every set with
should differ in only $o(\Sigma_n)$ places from almost all primes plus a near-optimal linear family of triple products. Two examples show why this is too rigid.
6.1 A private-prime lift
Start with a near-optimal linear triple family $\mathcal H_n$ using odd primes below $n/2$. Keep primes above $n/2$, replace every unused prime $p\le n/2$ by $2p$, and retain the triple products. Each $2p$ has the private prime divisor $p$; no other chosen element contains that prime. The triple products are still protected by linearity. The new set remains strongly $2$-primitive and has the same second-order asymptotic size.
But essentially every prime below $n/2$ has been replaced by a composite. The symmetric-difference distance from the all-primes model is
So literal edit-distance stability fails by much more than the scale of the second-order term.
6.2 A nonlinear sunflower at almost no cost
Even if we insist on squarefree triple products, linearity need not hold in the original representation. Add a sunflower
$$\{\{2,3,r\}:r\in R\}$$
for many private petals $r$. Remove the singleton primes $2,3,r$ and add the products $6r$. The net cardinality loss is only $2$, while the support family contains $\asymp n/\log n$ edges sharing the pair $\{2,3\}$. It is $2$-cover-free because every target has its own private petal, but any linear subfamily contains at most one sunflower edge.
The two examples have the same cause: degree-one prime coordinates carry a huge amount of neutral decoration. Stability can only become true after quotienting out this freedom.
7. Near equality in the upper bound is an exact star forest
Fix the chosen factorizations $a=u_av_a$ from the multiplicative basis proof. Make a graph $G_A$ with vertex set $\mathcal B$, one edge $\{u_a,v_a\}$ for each $a\in A$, and allow loops.
Exact factor-graph theorem. Every nonloop edge has a degree-one endpoint, and every loop is an isolated component. Thus $G_A$ is a disjoint union of stars and isolated loops. If $U$ is the number of unused vertices and $c(G_A)$ is the number of nonloop star components, then
$$|\mathcal B|-|A|=U+c(G_A).$$
The proof is the private-factor lemma read without discarding information. If a nonloop edge $\{u,v\}$ had both endpoints in other edges, the corresponding two elements would have product divisible by $uv=a$. If a loop at $u$ met another edge, then $u^2$ would divide the square of the other element.
Figure 5. The upper-bound defect is counted exactly: an unused basis coordinate costs one, and each nonloop star component costs one.
then the defect $D_n=|\mathcal B|-|A|$ is at most $(\delta_n+o(1))\Sigma_n$. Hence almost every secondary coordinate in both $\mathcal B_2$ and $\mathcal B_3$ occurs in exactly one chosen factor pair. This is a rigorous stability theorem for the upper-bound certificate. It is not yet a classification of the original integers.
8. Private-prime compression gives the correct normal form
Suppose a prime $p$ divides $a\in A$ and divides no other element of $A$. Replacing $a$ by $p$ preserves the cardinality and strong $2$-primitivity. The new target $p$ cannot divide a product of two other elements because neither contains $p$; for any unchanged target, replacing $a$ by a divisor only decreases the product available to cover it.
For primes $p\gt n^{3/5}$, the basis factorization can be chosen to expose $p$ in every multiple. Thus a degree-one large-prime coordinate is genuinely a globally private prime and can be contracted. Contract all such coordinates simultaneously.
Normalized stability modulo private-prime lifts. From every strongly $2$-primitive $A\subseteq[1,n]$ satisfying the near-extremal bound above, private-prime contractions produce a strongly $2$-primitive set $\widetilde A$ of the same size with
The geometry is now transparent. Before compression, the leading $\pi(n)$ coordinates may be represented by arbitrary private multiples. After compression, almost every prime appears literally. Every residual composite is supported entirely on the small exceptional prime hub $V$, and its excess over the missing primes is exactly the second-order gain.
Figure 6. Literal stability fails on the left. The quotient by private-prime contractions produces the exact normal form on the right.
There is one more unconditional consequence. At most $|V|$ members of $C$ have support of size at most two. Indeed, any one- or two-prime residual element must have a valuation coordinate in which it is a strict global maximum; assign the element to such a prime. Two elements cannot receive the same prime. Therefore, in a normalized near-extremizer, all but $o(\Sigma_n)$ residual composites have at least three distinct prime divisors.
This is close to the hoped-for triple picture, but it does not prove that the typical element has exactly three prime factors, that it is squarefree, or that its prime factors lie near $n^{1/3}$.
9. Conditional stability when the residual core is made of triples
Now impose an additional hypothesis:
Squarefree-triple hypothesis. All but $o(\Sigma_n)$ elements of $C$ are products of three distinct primes.
Let $\mathcal H$ be the support family of those triples. Strong $2$-primitivity says precisely that it is $2$-cover-free:
$$
E\nsubseteq F\cup G
\qquad(E,F,G\in\mathcal H,\ E\notin\{F,G\}),
$$
with $F=G$ allowed.
9.1 Pruning to a linear core
If two triples $E,F$ share a pair and $x$ is the third vertex of $E$, then $x$ has degree one. Otherwise another edge $G$ containing $x$, together with $F$, would cover $E$. Delete such a petal edge whenever a repeated pair occurs. Every deletion removes at least one vertex as well as one edge, so the excess $|\mathcal H|-|V(\mathcal H)|$ does not decrease.
The result is a linear subfamily $\mathcal L$ with
and only $o(\Sigma_n)$ edges are lost in the normalized near-extremal setting. Notice the nuance: linearity is recovered after pruning private petals; it is not asserted for every cover-free representation.
9.2 Saturating the feasible pair space
Define
$$
\mathcal D_n=
\{(p,q):p\lt q\text{ prime and }pq^2\le n\}.
$$
Sort an edge of $\mathcal L$ as $p\lt q\lt r$ and map it to $(p,q)$. Linearity makes this map injective. Since $r\gt q$ and $pqr\le n$, the image lies in $\mathcal D_n$. Direct counting gives
The range $q\le n^{1/3}$ contributes $(9/2+o(1))\Sigma_n$; the range $n^{1/3}\lt q\le n^{2/5}$ contributes $(9+o(1))\Sigma_n$; the tail is negligible. Since the linear core already has $(27/2-o(1))\Sigma_n$ edges, its lower-pair map misses only $o(\Sigma_n)$ feasible pairs.
Moreover, for every fixed $\varepsilon\gt0$, all but $o(\Sigma_n)$ edges satisfy
This is the desired scale stability. It does not imply uniqueness. Different one-factorizations or different proper edge-colourings can change $\Theta(\Sigma_n)$ triples while preserving the same pair occupancy and extremal count. The stable object is the saturated pair space, not an individual colouring.
10. The remaining inverse problem is weighted and cover-free
After normalization, write every $c\in C$ as an exponent vector
and that almost every vector has support at least three. What remains unknown is the sharper assertion
$$
\#\{c\in C:c\text{ is not a squarefree product of three primes}\}
=o(\Sigma_n).
$$
This gap may conceal genuine alternative near-extremizers. If a triple $pqr$ has product slack, one can contemplate replacing it by $2pqr$ while deleting the singleton prime $2$. For a linear support family, the underlying three private coordinates still prevent coverage. It is not known whether a positive density of such bounded-hub multiplier layers can coexist globally near the $27/2$ threshold, or whether cross-layer collisions force them to be negligible.
Open normalized stability problem. Classify weighted $2$-cover-free exponent families under the product constraint $c\le n$ whose excess is $(27/2-o(1))\Sigma_n$. Decide whether the residual core is predominantly squarefree of support three, or whether bounded-hub multiplier layers produce genuinely different normalized near-extremizers.
10.1 What a Bedert-style descent would need
The additive proof suggests three design requirements, not three ready-made lemmas.
A local order projection. One needs to isolate a valuation or support layer while controlling contamination from lower exponent vectors. A divisor-lattice zeta or Möbius transform is a possible language, but no analogue of Bedert’s $L^1$-contractive residue projection is presently available.
An isolation-or-descent alternative. If the secondary basis slots do not have predominantly prime cofactors, the argument should descend to a structured hub of repeated factors with a quantitative gain or a summable loss.
A no-leakage invariant. Low codegree and scale signatures work for squarefree triples. A general invariant must survive repeated exponents and supports of size at least four.
The star-forest theorem is already a first inverse statement of this kind: near equality forces almost every basis coordinate to be a leaf attached to a small collection of hubs. The hard step is to turn that certificate-level hub structure into a classification of the original weighted exponent vectors.
11. What is proved, conditional, and open
Published/preprint additive result: Bedert proves $S(A)\ge |A|/3+c\log\log|A|$, together with inverse information involving additive dimension, dense $F_4$-models, energy and a $99\%$ decomposition into large small-doubling pieces.
Supplied 2026 manuscript: the claimed $27/2$ asymptotic follows from the multiplicative-basis upper bound and the logarithmic-cell hypergraph construction described above. The status caveat at the beginning remains in force.
Unconditional stability developed in the accompanying note: literal edit stability is false; the upper-bound factor graph is an exact star forest; private-prime compression gives the normal form $(\mathbb P\setminus V)\cup C$; and almost every residual composite has at least three distinct prime divisors.
Conditional stability: if almost all residual composites are squarefree triples, pruning gives a near-maximal linear core whose lower pairs saturate $\mathcal D_n$, and almost every prime factor lies at scale $n^{1/3+o(1)}$.
Open: prove that the normalized residual core is predominantly squarefree of support three, or construct a different near-extremal weighted cover-free core.
The conceptual moral is simple but useful. Bedert’s $p$-adic fibres and the multiplicative logarithmic cells are not the same mathematical object. Yet both proofs win by choosing coordinates in which local constructions can be made strong and then finding an exact invariant that prevents those constructions from interfering. In the stability problem, the same philosophy says to quotient the neutral directions first. Once private-prime lifts are removed, the true obstruction becomes visible: a weighted cover-free family on a tiny prime hub. That is the right object for the next theorem.
P. Chojecki, The Second Term for Strongly 2-Primitive Sets, user-supplied five-page manuscript, July 2026; no public identifier located at the time of writing.
Let $\left(M^n, g\right)$ be a closed Riemannian manifold with $C^1$-smooth $g_{i j}$. The spectrum of the Laplace operator on $M$ is discrete. There is a sequence of eigenvalues
$$ 0=\lambda_1<\lambda_2 \leq \lambda_3 \ldots $$
that tend to $\infty$ and a sequence of (real) eigenfunctions $u_k$ such that
$$ \Delta_g u_k+\lambda_k u_k=0 . $$
Our enumeration of eigenvalues is non-standard. We start with $\lambda_1=0$ and $u_1=1$ on $M$. The nodal domains of $u_k$ are the connected components of $M \backslash Z_{u_k}$, where $Z_{u_k}$ is the zero set of $u_k\left(Z_{u_k}\right.$ is called the nodal set of $\left.u_k\right)$. The Courant nodal domain theorem states that the $k$-th eigenfunction $u_k$ has at most $k$ nodal domains. If the multiplicity of an eigenvalue is more than 1 , one may enumerate the eigenfunctions corresponding to this eigenvalue in any order. Our main result is the local version of Courant’s theorem.
In many mathematical and engineering problems, we are interested in finding solutions that satisfy certain constraints. A powerful modern paradigm is to train a neural network as a generator that proposes candidate solutions, and use a differentiable validator (i.e., loss function) to evaluate how well they satisfy those constraints. This feedback is then used to update the network via gradient descent.
In this note, we illustrate this approach by using a neural network to generate vectors that nearly achieve equality in the Cauchy-Schwarz inequality.
1. Problem Setup: Making Cauchy-Schwarz Nearly Tight
Equality holds if and only if x\mathbf{x} and y\mathbf{y} are linearly dependent: y=kx\mathbf{y} = k \mathbf{x} for some scalar kk.
Objective:
Given an input vector x\mathbf{x}, train a neural network NN to output a vector y=N(x)\mathbf{y} = N(\mathbf{x}) such that x,y\mathbf{x}, \mathbf{y} are as close to colinear as possible.
2. Generator: Neural Network Design
Let N(⋅)N(\cdot) be a feedforward neural network (e.g. MLP) with:
Structure: Simple MLP with 1–2 hidden layers (ReLU), and a linear output layer (no activation)
3. Validator: Loss Function Design
To measure how close x,y\mathbf{x}, \mathbf{y} are to colinearity, use cosine similarity: cos(θ)=⟨x,y⟩∥x∥⋅∥y∥+ε\cos(\theta) = \frac{\langle \mathbf{x}, \mathbf{y} \rangle}{\|\mathbf{x}\| \cdot \|\mathbf{y}\| + \varepsilon}
We define the loss as: L(x,y)=1−∣⟨x,y⟩∥x∥⋅∥y∥+ε∣L(\mathbf{x}, \mathbf{y}) = 1 – \left|\frac{\langle \mathbf{x}, \mathbf{y} \rangle}{\|\mathbf{x}\| \cdot \|\mathbf{y}\| + \varepsilon}\right|
L=0L = 0 when x\mathbf{x} and y\mathbf{y} are perfectly aligned or anti-aligned
ε≪1\varepsilon \ll 1 is a small constant for numerical stability
This validator provides a differentiable measure of alignment quality.
4. Training Procedure
Data generation: Sample random input vectors x\mathbf{x} from e.g. N(0,I)\mathcal{N}(0, I)
Loss computation: Evaluate L(x,y)L(\mathbf{x}, \mathbf{y})
Backpropagation: Compute ∇θL\nabla_\theta L and update θ\theta using an optimizer (e.g. Adam)
Repeat until convergence
At the end of training, the network learns to generate vectors y\mathbf{y} nearly colinear with x\mathbf{x}, thus making the Cauchy-Schwarz inequality nearly tight.
5. General Framework: Generator + Validator
This method exemplifies a general and powerful pattern in deep learning:
Component
Role
Description
Neural Network NN
Generator / Solver
Maps input (or noise) to a candidate solution
Validator VV
Loss / Constraint Function
Evaluates how well the candidate satisfies the constraints (must be differentiable)
Optimizer
Learning Engine
Uses gradients to update NN so that the solutions improve over time
Functional problems: e.g. finding extremals in calculus of variations
Neural symbolic systems: e.g. generating logic-constrained expressions
Inverse design: input-to-output mappings constrained by physical or mathematical laws
Conclusion
Training a neural network to minimize a differentiable validator is a powerful method to learn constrained solutions. The Cauchy-Schwarz example shows how even classical inequalities can be embedded into a modern optimization loop, potentially aiding in automated reasoning, symbolic learning, or mathematical discovery.
Would you like this exported to PDF with rendered math? Or should I write a minimal PyTorch implementation to match?
We begin with the Weierstrass form of elliptic equation, i.e. look it as an embedding cubic curve in .
Definition 1 (Weierstrass form), In general the form is given by,
If , then, we have a much more simper form,
Remark 1
Where .
We have two way to classify the elliptic curve living in a fix field . \paragraph{j-invariant} The first one is by the isomorphism in . i.e. we say two elliptic curves is equivalent iff
is a isomorphism such that .
Definition 2 (j-invariant) For a elliptic curve , we have a j-invariant of , given by,
Why j-invariant is important, because j-invariant is the invariant depend the equivalent class of under the classify of isomorphism induce by . But in one equivalent class, there also exist a structure, called twist.
Definition 3 (Twist) For a elliptic curve , all elliptic curve twist with is given by,
So the twist of a given elliptic curve is given by:
Remark 2 Of course a elliptic curve is the same as , induce by the map .
But this moduli space induce by the isomorphism of is not good, morally speaking is because of the abandon of universal property. see \cite{zhang}. \paragraph{Level structure} We need a extension of the elliptic curve , this is given by the integral model.
Definition 4 (Integral model), . is regular and minimal, the construction of is by the following way, we first construct and then blow up. is given by the Weierstrass equation with coefficent in .
Remark 3 The existence of integral model need Zorn’s lemma.
Definition 5 (Semistable) the singularity of the minimal model of are ordinary double point.
Remark 4 Semistable is a crucial property, related to Szpiro’s conjecture.
Definition 6 (Level structure)
The weil pairing of is given by a unit in cycomotic fields, i.e.
What happen if ? In this case we have a analytic isomorphism:
Given by,
Where , and the Weierstrass equation is given by . The full n tructure of it is given by and the value of , i.e.
Where is induce by
The key point is following:
Theorem 7, the moduli of elliptic curves with full level n-structure is identified with
Now we discuss the Mordell-Weil theorem.
Theorem 8 (Mordell-Weil theorem)
The proof of the theorem divide into two part:
Weak Mordell-Weil theorem, i.e. , is finite.
There is a quadratic function,
, is finite.
Remark 5 The proof is following the ideal of infinity descent first found by Fermat. The height is called Faltings height, introduce by Falting. On the other hand, I point out, for elliptic curve , there is a naive height come from the coefficient of Weierstrass representation, i.e. .
While the torsion part have a very clear understanding, thanks to the work of Mazur. The rank part of is still very unclear, we have the BSD conjecture, which is far from a fully understanding until now.
But to understanding the meaning of the conjecture, we need first constructing the zeta function of elliptic curve, .
\paragraph{Local points} We consider a local field , and a locally value map , then we have the short exact sequences,
Topologically, we know are union of disc indexed by ,
. Define , then we have Hasse principle:
Theorem 9 (Hasse principle)
Remark 6 I need to point out, the Hasse principle, in my opinion, is just a uncertain principle type of result, there should be a partial differential equation underlying mystery.
So count the points in reduce to count points in , reduce to count the Selmer group . We have a short exact sequences to explain the issue.
I mention the Goldfold-Szipiro conjecture here. , there such that:
\paragraph{L-series} Now I focus on the construction of , there are two different way to construct the L-series, one approach is the Euler product.
Where or when has bad reduction on .
The second approach is the Galois presentation, one of the advantage is avoid the integral model. Given is a fixed prime, we can consider the Tate module:
Then by the transform of different embedding of , we know , decompose it into a lots of orbits, so we can define , the decomposition group of (extension of to ). We define is the inertia group of .
Then is generated by some Frobenius elements
So we can define
And then .
Faltings have proved is the invariant depending the isogenous class in the follwing meaning:
Theorem 10 (Faltings) is an isogenous ivariant, i.e. isogenous to iff , .
Where come from an automorphic representation for . Now we give the statement of BSD onjecture. is the regulator of , i.e. the volume of fine part of with respect to the Neron-Tate height pairing. be the volume of Then we have,
.
.
Here is an explictly positive integer depending only on for dividing .
补充说明
以下是新整理的中文说明;上方旧博客原文保持不变。
Birch and Swinnerton-Dyer 猜想把椭圆曲线的有理点群和它的 $L$ 函数在 $s=1$ 处的零点阶联系起来。它是数论中最核心的桥之一:一边是 Diophantine 方程的解,另一边是解析函数的特殊值。
Remark 1 Why it is but not 1? if it is 1, then the action distribute is not trasitive on , i.e. every element in unite group present a connected component
Now we consider the subgroup .
We are most interested in the case . So how to investigate ? We can look at the action of it on something, for particular, we look at the action of it on Riemann sphere, i.e. given by fraction linear map:
Remark 2 What is fraction linear map? This action carry much more information than the action on vector, thanks for the exist of multiplication in and the algebraic primitive theorem. Due to I always looks the fraction linear map as something induce by the permutation of the roots of polynomial of degree 2, this is true at least for fix points, and could natural extension. So how about the higher dimension generate? consider the transform of tuples induce by polynomial with degree ?
Remark 3
, then action faithful on , i.e. except identity, every action is nontrivial. This is easy to be proved, observed,
Up half plane is invariant under the action of , i.e. , . The proof is following,
Now we focus on or the same,. All the argument for make sense for
Fix , define,
Then is the kernel of map , i.e. we have short exact sequences,
Remark 4 The relationship of is just like .
Definition 1 (Congruence group) A subgroup of is called a congruence group iff , .
Example 1 We give two examples of congruence subgroups here.
Definition 2 (Fundamental domain)
Now here is a theorem charistization the fundamental domain.
Theorem 3 This domain is a fundamental domain of
Proof: Ths key point is have two generators,
.
.
Thanks to this two generator exactly divide the action of on into a lots of scales, then is a fundamental domain is a easy corollary.
Remark 5 This is not rigorous, need be replace by , but this is very natural to get a modification to a right one.
Remark 6 are equivalent iff and or if on the unit circle and
Remark 7 If , then expect in the following three case:
if .
if .
if .
Where .
Remark 8 The group is generated by the two elements , . In other word, any fraction linear transform is a “word” induce by . But not free group, we have relationship .
The natural function space on is the memorphic function, under the map: , it has a -expension,
And there are only finite many negative such that .