博客

  • 从支撑接触到 Hausdorff 分层:凸 k-Hessian 解的四步证明机制

    这篇 note 介绍论文第3–6章的主证明。 目标是说明:若 $u$ 是开凸域 $\Omega$ 上满足

    $$\sigma_k(D^2u)=1,\qquad 2\le k\le n,$$

    的凸黏性解,则其非局部 $C^2$ 集满足

    $$\mathcal H^{n-1}_{\mathrm{loc}}\bigl(\operatorname{Sing}(u)\bigr)=0.$$

    证明把支撑接触集的维数、John 椭球的短轴、严格 section 的体积衰减和 Mooney 覆盖定理连接成一条闭合链。临界接触维数是 $q=n-k+1$。

    1. 接触分层与主证明链

    对 $x\in\Omega$ 和 $p\in\partial u(x)$,令 $\ell_{x,p}$ 为相应的支撑仿射函数,并用 $d_u(x,p)$ 表示局部接触集 $\{u=\ell_{x,p}\}$ 在小尺度稳定后的仿射维数。定义

    $$C_m(u)=\left\{x:d_u(x,p)\ge m\ \text{for every }p\in\partial u(x)\right\}.$$

    “对每一个支撑斜率”是必要量词,因为第6章的 covering theorem 要求同一点的每一个 section 都满足体积衰减。

    1. 第3章:若某个支撑斜率的接触维数不超过 $n-k$,则 $u$ 在该点附近光滑。
    2. 第4章:正的 $k$-Hessian 密度控制小次水平集 John 椭球最短 $k$ 根轴的乘积。
    3. 第5章:$m$ 维接触与二次局部化共同产生 $|S_h|\lesssim h^{(m+k)/2}$。
    4. 第6章:Mooney 覆盖把 section 体积衰减转化为 Hausdorff 分层。

    $$\operatorname{Sing}(u)\subset C_q(u),\qquad q=n-k+1,\qquad \mathcal H^{2n-m-k}_{\mathrm{loc}}\bigl(C_m(u)\bigr)=0.$$

    最后取 $m=q$,便有 $2n-q-k=n-1$。

    2. 第一个引擎:小接触集推出局部光滑

    这一节中,标准逼近、内点二阶估计和 Evans–Krylov 都是熟悉的步骤;真正决定论证能否启动的是一个马鞍面。减去一个支撑仿射函数后,归一化为 $u(0)=0$、$u\ge0$。若 $\{u=0\}$ 在小球内落入至多 $n-k$ 维的子空间 $Z$,取正交分解 $\mathbb R^n=Y\oplus Z$。此时 $\dim Y\ge k$;只需在 $Y$ 中选出 $k$ 个方向,并构造

    $$Q(y,z)=A|y|^2-|z|^2+\alpha r^2.$$

    这个马鞍面的两组方向承担相反而互补的任务:$Y$ 中至少 $k$ 个强正曲率方向保证 $D^2Q\in\Gamma_k$;沿接触空间 $Z$ 的负曲率则把 $Q$ 在球面边界压到 $u$ 的下方。选择 $A$、$\alpha$ 并适当缩放后,可以同时做到“边界上 $Q<u$”与“中心附近 $Q>u$”。这正是单纯凸抛物面无法提供的几何。

    对 CNS 给出的光滑 $k$-admissible 逼近 $u_j$,选择正则值 $t_j$,并取包含原点的连通分支

    $$D_j=\{Q-u_j>t_j\}.$$

    边界与中心的符号差保证 $D_j$ 是困在小球内部的移动区域,而正则值保证其边界光滑。于是可以在 $D_j$ 上应用 Chou–Wang 局部二阶估计,得到与 $j$ 无关的 Hessian 界。随后对凹算子 $F=\sigma_k^{1/k}$ 使用 Evans–Krylov 与 Schauder,才得到极限 $u$ 的局部光滑性。

    一致收敛 $u_j\to u$ 本身不产生 $C^2$ 极限;真正的正则化步骤是马鞍面切出的变化域和 Chou–Wang Hessian 界。

    图1。 马鞍面的两个方向:接触方向中的负曲率负责封住边界,至少 $k$ 个横向正曲率负责保持 $k$-admissibility。图面可用鼠标拖动旋转,并用滚轮缩放。

    3. 第二个引擎:John 椭球与内部黏性测试

    严格 section $K_h=\{u\le h\}\cap\overline B_R$ 虽然是凸集,却可能有复杂边界、随尺度旋转,并且各方向尺度相差很大;证明并没有假定它本身是椭球。John 定理的作用,是用一个内接椭球把这种高复杂度压缩成 $n$ 个半轴和一组主方向:

    $$E_h=z_h+A_hB_1,\qquad E_h\subset K_h\subset z_h+nA_hB_1,$$

    把 $A_h$ 的半轴按 $a_1\ge\cdots\ge a_n$ 排列。于是 section 的全部凸几何复杂度,在固定维数常数 $n$ 的损失内,变成了这 $n$ 根有序轴的账本。由 $E_h$ 构造测试函数

    $$P_h(x)=h\left|A_h^{-1}(x-z_h)\right|^2.$$

    令 $\mu_h=\min(P_h-u)$。平移后的 $P_h-\mu_h$ 在某个辅助内点 $y_h$ 从上方接触 $u$。因为 $D^2P_h=2hA_h^{-2}>0$,黏性不等式给出

    $$a_q(h)\cdots a_n(h)\lesssim \lambda^{-1/2}h^{k/2},\qquad q=n-k+1.$$

    这是最短 $k$ 根 John 轴乘积的上界。方程控制的是乘积,而不是每一根轴。

    图2。 第4章内部黏性测试的几何结构。抬高平面显示 graph cap;$K_h$ 与 $E_h$ 位于自变量空间,测试点 $y_h$ 由最小化 $P_h-u$ 自动产生。

    4. 第三个引擎:接触维数与 section 体积

    这一节的核心不是把 $K_h$ 想象成某个规则形状,而是把上一节得到的有序 John 轴分成三段。记 $q=n-k+1$。测试函数与黏性方程给出最短 $k$ 根轴的乘积上界

    $$a_q\cdots a_n\lesssim \lambda^{-1/2}h^{k/2}.$$

    另一方面,若支撑接触集含有半径 $\rho$ 的 $m$ 维相对球,John 上包含和宽度的极小极大表征给出 $a_m\gtrsim\rho$。由于 $a_1\ge\cdots\ge a_n$,这实际意味着

    $$a_1,\ldots,a_m\gtrsim\rho.$$

    当 $m\ge q$ 时,两段索引在中间确实重合:$a_q,\ldots,a_m$ 同时属于“接触给下界”的区间和“PDE 控制乘积”的区间。把这些已有下界的重合因子从短轴乘积中除去,便得到

    $$a_{m+1}\cdots a_n\lesssim \lambda^{-1/2}\rho^{-(m-q+1)}h^{k/2}.$$

    这一步迫使重合区间之后的最短 $n-m$ 根轴总体快速衰减,但还留下一个漏洞:前 $m$ 根轴虽然有下界,却完全可能非常长;所以仅靠上式还不能控制整个 section 的体积。修补这个漏洞的技巧是令

    $$v(x)=u(x)+\frac{|x|^2}{2}.$$

    加不加这个光滑二次函数,奇异集完全相同:$u$ 在一点附近属于 $C^2$ 当且仅当 $v$ 属于 $C^2$。这里也不对 $v$ 使用 $k$-Hessian 方程;所有 PDE 轴乘积估计仍来自 $u$。二次项只承担几何局部化。对相应严格 section,$v(x)\le h$ 立即蕴含 $|x|^2/2\le h$,因此

    $$S_h^v\subset K_h\cap B_{\sqrt{2h}}.$$

    这只是一项光滑抬升,却把可能任意长的前 $m$ 个接触方向全部截到 $O(\sqrt h)$,贡献 $h^{m/2}$;重合消去后剩余短轴的乘积贡献 $h^{k/2}$。两部分合在一起得到

    $$|S_h^v|\lesssim \lambda^{-1/2}\rho^{-(m-q+1)}h^{(m+k)/2}.$$

    图3。 半轴按 $a_1\ge\cdots\ge a_n$ 排列。上括号覆盖 $q\le i\le n$,下括号覆盖 $1\le i\le m$,所以 $q\le i\le m$ 是两种控制的重合区间;着色的 $m+1\le i\le n$ 正是消去重合后要处理的短轴。
    图4。 Proposition 5.1 的几何顺序:接触片、$K_h$、John 轴下界以及 $S_h^v\subset K_h\cap B_{\sqrt{2h}}$。

    5. 第四个引擎:从体积衰减到 Hausdorff 分层

    对 $x\in C_m(u)$ 以及 $v$ 的每一个支撑斜率,前一节给出

    $$|S_h^v(x)|\le C_{x,p}h^{(m+k)/2}.$$

    令 $s=m+k-n$,则 $(m+k)/2=(n+s)/2$。Mooney 的 section-covering theorem 因而推出

    $$\mathcal H^{n-s}\bigl(C_m(u)\bigr)=\mathcal H^{2n-m-k}\bigl(C_m(u)\bigr)=0.$$

    覆盖定理允许常数与高度阈值依赖 $(x,p)$,所以接触球半径 $\rho$ 没有统一下界并不妨碍论证。另一方面,满维接触会产生曲率任意小的上方二次测试,与 $\sigma_k\ge\lambda$ 矛盾,因此 $C_n(u)=\varnothing$。

    第3章给出 $\operatorname{Sing}(u)\subset C_q(u)$。在 $m=q=n-k+1$ 处应用分层结论,便得到 $\mathcal H^{n-1}_{\mathrm{loc}}(\operatorname{Sing}(u))=0$。

    6. 模型核算:$n=4$、$k=3$

    此时 $q=2$。对临界层 $m=2$,接触几何给出 $a_2\gtrsim\rho$,而 PDE 给出 $a_2a_3a_4\lesssim h^{3/2}$。加入二次项后,

    $$|S_h|\lesssim h^{5/2},\qquad \mathcal H^3(C_2)=0.$$

    对更深的 $m=3$ 层,

    $$|S_h|\lesssim h^3,\qquad \mathcal H^2(C_3)=0.$$

    图5。 $n=4$、$k=3$ 时,$m=2$ 与 $m=3$ 两层的 John 轴乘积、section 体积指数和 Hausdorff 指数。

    7. 端点与限制

    • $k=n$:$q=1$,结论恢复 Monge–Ampère 情形的 $\mathcal H^{n-m}(C_m)=0$。
    • $k=2$:本方法给出临界接触层的余维一零测;全光滑还需要 strict $2$-convexity 的额外输入。
    • 轴的控制:方程只控制短轴乘积,主轴可以随尺度旋转并重新分配,所以当前结论是 Hausdorff nullity,而不是统一 Minkowski packing。

    关于本文。 这是一篇介绍 Xiyu Hu, Sharp Hausdorff Bounds for the Interior Singular Set of Convex k-Hessian Solutions 第3–6章的证明导读 note。论文第8章的 mean-value 公式提供概念动机,但不作为上述主证明链的输入。


  • From Sum-Free Sets to Strongly 2-Primitive Sets: Localization, Patching, and the Search for Stability

    Two extremal problems can share a proof architecture without sharing a proof. Benjamin Bedert’s 2025 breakthrough on large sum-free subsets is additive and Fourier-analytic. A July 2026 manuscript on strongly $2$-primitive sets is multiplicative and hypergraph-theoretic. Their common language is not a transferable sieve, but a four-step design: localize, build a strong object in each local block, prevent leakage between blocks, and add the gains. That comparison points toward the right stability question – and also reveals why the most naive version of stability is false.

    Status and scope. Bedert’s theorem is the February 2025 arXiv preprint cited below. The $27/2$ theorem is contained in a five-page manuscript by Przemek Chojecki supplied for this post in July 2026; no public preprint or peer-reviewed version was found, and the Erdős Problems page for #793 still labels the problem open at the time of writing. The written proof was checked here for internal coherence, but that is not independent peer review. The factor-graph and normalized-stability results later in this article are the additional analysis developed in the accompanying long note. Conditional and open statements are labelled explicitly.

    Here is the headline before the details. The supplied manuscript claims

    $$
    F(n)=\pi(n)+\left(\frac{27}{2}+o(1)\right)
    \frac{n^{2/3}}{(\log n)^2},
    $$

    where $F(n)$ is the largest size of a strongly $2$-primitive subset of $[1,n]$. One might expect every near-extremizer to be close to the construction behind this formula: almost all primes, plus products of three primes near $n^{1/3}$ arranged as a nearly saturated linear $3$-uniform hypergraph. That literal statement is false. There is, however, an exact and useful replacement: after contracting harmless private-prime lifts, every near-extremizer has a canonical prime layer and a small exceptional hub carrying the entire second-order excess. If the residual hub is mostly made of squarefree triples, then pair saturation and $n^{1/3}$-scale stability follow. Proving – or refuting – that triple-core hypothesis is the remaining inverse problem.

    1. Two problems that look similar only from far away

    1.1 The sum-free extraction problem

    A set $B\subset\mathbb Z$ is sum-free if it contains no $x,y,z$, not necessarily distinct, with

    $$x+y=z.$$

    For a finite set $A$ of integers, write

    $$
    S(A)=\max\{|B|:B\subseteq A,\ B\text{ is sum-free}\},
    $$

    and, for $N$-element sets of positive integers,

    $$
    S(N)=\min_{|A|=N}S(A).
    $$

    This is an extraction problem. An adversary gives us an arbitrary host set $A$, and we must find a large structured subset inside it. Erdős’s middle-third argument gives $S(A)\ge |A|/3$. Alon and Kleitman improved this to $(N+1)/3$, and Jean Bourgain proved $S(N)\ge (N+2)/3$ in 1997. The long-standing qualitative question was whether the additive improvement can tend to infinity:

    $$S(N)\ge \frac N3+\omega(N),\qquad \omega(N)\longrightarrow\infty.$$

    Bedert answered yes, proving that an absolute $c\gt0$ exists such that

    $$
    S(A)\ge \frac{|A|}{3}+c\log\log |A|.
    $$

    1.2 The strongly $2$-primitive packing problem

    A set $A\subseteq[1,n]$ is strongly $2$-primitive when

    $$
    a\nmid bc
    \qquad
    (a,b,c\in A,\ a\notin\{b,c\}),
    $$

    where $b=c$ is allowed. The word strongly matters. Under a more recent convention, a $2$-primitive set only forbids witnesses $b,c$ that are distinct. For example, $\{4,5,6\}$ passes that weaker test but fails the strong one because $4\mid6^2$.

    Now define

    $$
    F(n)=\max\{|A|:A\subseteq[1,n],\ A\text{ is strongly }2\text{-primitive}\}.
    $$

    This is a packing problem. The host interval is fixed, and we directly construct the largest possible forbidden-divisibility family. The all-primes set gives the leading term $\pi(n)$. In 1938 Erdős proved upper and lower bounds of the form

    $$
    \pi(n)+c_1\frac{n^{2/3}}{(\log n)^2}
    \le F(n)\le
    \pi(n)+c_2\frac{n^{2/3}}{(\log n)^2},
    $$

    and later asked whether the second-order term has an asymptotic constant. This is the modern Erdős Problem #793. The supplied 2026 manuscript proposes that the constant is $27/2$; it is important not to say that Erdős himself conjectured this numerical value.

    Comparison of the sum-free extraction problem and the strongly 2-primitive packing problem
    Figure 1. The quantifiers already separate the two questions. Sum-free theory extracts a subset from an arbitrary host; the multiplicative problem packs a forbidden-divisibility family into a fixed interval.

    1.3 Equality versus order

    The relation $a\nmid bc$ is not ordinary product-freeness. The set $\{6,10,15\}$, for instance, has no internal equality $xy=z$, yet $6\mid10\cdot15$. Prime valuations expose the real geometry:

    $$
    a\mid bc
    \quad\Longleftrightarrow\quad
    v_p(a)\le v_p(b)+v_p(c)
    \quad\text{for every prime }p.
    $$

    Thus the forbidden relation is a coordinatewise domination inequality in the divisor lattice. For squarefree integers, if $E_a=\{p:p\mid a\}$, it becomes

    $$E_a\subseteq E_b\cup E_c.$$

    So the natural combinatorial object is a $2$-cover-free family, not the solution set of a linear equation. Fourier characters are superb at detecting equations such as $x+y=z$. They do not come with an evident contractive projection that detects the partial order $\nu(a)\le\nu(b)+\nu(c)$. This is the first reason Bedert’s proof cannot simply be copied into the multiplicative setting.

    2. The $27/2$ upper bound: every element needs a private factor

    Set

    $$
    y=n^{1/3},\qquad
    M=\frac{y}{\log n},\qquad
    \Sigma_n=M^2=\frac{n^{2/3}}{(\log n)^2}.
    $$

    The upper bound begins with a small lemma that contains more information than the inequality it proves.

    Private-factor lemma. Let $\mathcal B$ be a set of positive integers, and choose for every $a\in A$ a factorization $a=u_av_a$ with $u_a,v_a\in\mathcal B$. If $A$ is strongly $2$-primitive, then $|A|\le|\mathcal B|$.

    Think of $E_a=\{u_a,v_a\}$ as a two-element multiset. If $u_a\ne v_a$ and neither coordinate is private to $a$, another chosen pair contains $u_a$ and another contains $v_a$; the corresponding two elements have product divisible by $a$. If $u_a=v_a=x$, failure of privacy would give another pair containing $x$ twice, hence another element equal to $x^2=a$. Therefore every $a$ has a private coordinate, and these coordinates are distinct. That is the injection $A\hookrightarrow\mathcal B$.

    The manuscript chooses the multiplicative $2$-basis

    $$
    \begin{aligned}
    \mathcal B_0&=[1,n^{3/5}],\\
    \mathcal B_1&=\{p\text{ prime}:n^{3/5}\lt p\le n\},\\
    \mathcal B_2&=\{pq:p,q\le y\text{ prime}\},\\
    \mathcal B_3&=\{qr:y\lt q\le n^{2/5},\ r\le n/q^2,\ q,r\text{ prime}\}.
    \end{aligned}
    $$

    Every $m\le n$ can be factored into two elements of $\mathcal B=\mathcal B_0\cup\mathcal B_1\cup\mathcal B_2\cup\mathcal B_3$. The case split is elementary but carefully tuned. Small $m$ can be balanced into two factors below $n^{3/5}$, a prime factor above $n^{2/5}$ can be split off, and the remaining difficult case groups two of the three largest prime factors into an element of $\mathcal B_2$ or $\mathcal B_3$.

    The main prime layer comes from $\mathcal B_1$. The genuinely second-order counts are

    $$
    |\mathcal B_2|
    =\binom{\pi(y)+1}{2}
    =\left(\frac92+o(1)\right)\Sigma_n,
    $$

    because $\pi(n^{1/3})\sim3n^{1/3}/\log n=3M$, and

    $$
    |\mathcal B_3|
    =\sum_{y\lt q\le n^{2/5}}\pi\!\left(\frac n{q^2}\right)
    =(9+o(1))\Sigma_n.
    $$

    The contribution of $\mathcal B_0$ and all relevant overlaps is $o(\Sigma_n)$. Hence

    $$
    |\mathcal B|
    =\pi(n)+\left(\frac92+9+o(1)\right)\Sigma_n
    =\pi(n)+\left(\frac{27}{2}+o(1)\right)\Sigma_n.
    $$

    The private-factor injection then gives the upper bound. A caution that becomes crucial for stability: $\mathcal B_2$ and $\mathcal B_3$ are labels in an upper-bound certificate. Near-saturation of these labels does not immediately say that the original elements of $A$ are themselves products of two or three primes.

    3. The lower bound: turn properly coloured edges into prime triples

    The lower bound lives in a linear $3$-uniform hypergraph. Let $\mathcal H$ be a family of triples of distinct primes such that any two triples share at most one prime and every edge product is at most $n$. Define

    $$
    A_{\mathcal H}
    =\{p\le n:p\text{ prime and }p\notin V(\mathcal H)\}
    \cup
    \left\{\prod_{p\in E}p:E\in\mathcal H\right\}.
    $$

    Then

    $$|A_{\mathcal H}|=\pi(n)-|V(\mathcal H)|+|\mathcal H|.$$

    Why is this strongly $2$-primitive? A target triple product has three distinct prime coordinates. Each other hyperedge supplies at most one of them, so two other elements supply at most two. The same argument still works when the two witnesses coincide. Primes outside the vertex set remain singleton elements and cannot divide any other chosen element.

    3.1 Logarithmic prime bins

    Fix a small mesh $h\gt0$ and divide primes near $y=n^{1/3}$ into bins

    $$
    P_i=\{p\text{ prime}:ye^{ih}\lt p\le ye^{(i+1)h}\},
    \qquad
    \Delta_i=e^{(i+1)h}-e^{ih}.
    $$

    For a cell satisfying

    $$i\le j,\qquad i+2j\le-4,$$

    put $k=-i-j-3$. Then $i\le j\lt k$, and any $p\in P_i$, $q\in P_j$, $r\in P_k$ obeys $pqr\le n$.

    If $i\lt j$, properly edge-colour the complete bipartite graph between $P_i$ and $P_j$, using primes of $P_k$ as colours. When $i=j$, do the same with the complete graph on $P_i$. The lower pair $\{p,q\}$, coloured by $r$, becomes the hyperedge $\{p,q,r\}$.

    A complete bipartite graph between two prime bins properly edge-coloured by a third prime bin
    Figure 2. Within a cell, every colour class is a matching, so two produced triples never share a lower pair. Across cells, the sorted signature $(i,j,k)$ has fixed sum $-3$; two shared bin indices force the third and hence force the same cell.

    The proper colouring is the local no-interference mechanism. The constant-sum signature is the global one. Together they make the union over all cells a linear hypergraph.

    3.2 The cell weight

    For any fixed finite collection of bins, the prime number theorem gives

    $$|P_i|=(3+o(1))M\Delta_i.$$

    Off the diagonal, a cell contributes approximately $9M^2\Delta_i\Delta_j$ edges. A diagonal cell contributes half as many unordered pairs. The exact geometric-series identity is

    $$
    \sum_{\substack{i\lt j\\i+2j\le-4}}\Delta_i\Delta_j
    +\frac12\sum_{i\le-2}\Delta_i^2
    =e^{-h}+\frac12e^{-2h}.
    $$

    Letting $h\to0$, the weight tends to $3/2$, so

    $$
    |\mathcal H|
    =\left(9\cdot\frac32-o(1)\right)\Sigma_n
    =\left(\frac{27}{2}-o(1)\right)\Sigma_n.
    $$

    The number of vertices used is only $o(\Sigma_n)$. Replacing those primes by the triple products therefore yields the matching lower bound.

    The logarithmic feasible pair region split into the 9 over 2 and 9 counting regimes
    Figure 3. The same constant is visible in the canonical feasible pair space. The labels $9/2$ and $9$ are prime-counting contributions, not Euclidean areas of the drawing.

    4. From Erdős’s middle third to Bedert’s $c\log\log N$

    We now return to the additive problem. Let $\mathbb T=\mathbb R/\mathbb Z$, and let $\phi=\mathbf1_{(1/3,2/3)}$. The middle third of the circle is sum-free: two points in it cannot add, modulo $1$, to another point in it. Therefore, for every $x\in\mathbb T$,

    $$A_x=\{a\in A:ax\pmod1\in(1/3,2/3)\}$$

    is sum-free. Averaging $|A_x|$ over $x$ gives $|A|/3$. The problem is to force a positive fluctuation above that mean.

    After a harmless normalization, the relevant Fourier series has the form

    $$
    F_A(x)=\sum_{a\in A}\sum_{m\ge1}
    \frac{\chi(m)}m\cos(2\pi max),
    $$

    where $\chi$ is the non-principal real character modulo $3$. A lower bound for $\|F_A\|_1$ gives a one-sided large value and hence a large sum-free subset.

    4.1 Where Bourgain’s sieve loses a logarithm

    The historical correction is worth making explicit. The relevant predecessor is Jean Bourgain’s 1997 paper, not “John Bogan, 1993.” Also, the Littlewood $L^1$ conjecture had already been proved in 1981, independently by McGehee-Pigno-Smith and Konyagin. Bourgain’s argument combines Fourier analysis with a Möbius operation that filters harmonic indices.

    Schematically, if $R_Q$ denotes the positive $Q$-rough integers, then

    $$
    \sum_{d\mid\prod_{p\le Q}p}
    \frac{\mu(d)\chi(d)}d F_A(dx)
    =\sum_{a\in A}\cos(2\pi ax)+\mathcal R_Q(x).
    $$

    The left side costs

    $$
    \prod_{p\le Q}\left(1+\frac1p\right)\asymp\log Q
    $$

    under the triangle inequality. If $Q$ is large enough to make the rough remainder negligible directly, this factor consumes the logarithm delivered by a Littlewood-type lower bound. This explains why merely “using the Littlewood theorem harder” does not produce the desired unbounded gain.

    4.2 Bedert’s split: medium primes are sieved, small primes are projected

    Bedert chooses

    $$Q_1=(\log N)^{1/2},\qquad Q=(\log N)^{20}.$$

    He Möbius-sieves only the medium primes $Q_1\le p\le Q$. Their reciprocal sum is $O(1)$, so the norm loss is a constant rather than a logarithm. The small primes $p\le Q_1$ are handled by a different operation.

    For every small prime, record the exact valuation $\nu_p(a)$ and the unit residue after removing that prime power. Chinese remaindering packages all this data into a joint fibre $A(r,\nu)$. Fourier projection to a residue class,

    $$
    \operatorname{Proj}(H;\rho\bmod q)(x)
    =\sum_{m\equiv\rho\ (q)}\widehat H(m)e(mx),
    $$

    is an $L^1$-contraction:

    $$
    \|\operatorname{Proj}(H;\rho\bmod q)\|_1\le\|H\|_1.
    $$

    At the main valuation level, the projection removes small-prime harmonic contamination without paying the full Möbius product. After truncation, one sees a large exponential sum plus an $L^2$-small error, to which a robust McGehee-Pigno-Smith test function applies.

    At a general valuation level, lower fibres can still leak into the projected channel. Bedert proves an isolation-or-descent alternative: either the desired $L^1$ lower bound is already present, or a large fibre has a strictly lower fibre losing at most a controlled polylogarithmic factor. Iterating and reversing this descent produces a chain

    $$
    \nu^{(1)}\prec\nu^{(2)}\prec\cdots\prec\nu^{(J)},
    \qquad
    J\gg\frac{\log N}{\log\log N},
    $$

    whose fibre sizes grow geometrically.

    4.3 The non-Archimedean no-leakage lemma

    Associate to these levels the nested moduli

    $$
    q_i=\prod_{p\le Q_1}p^{\nu_p^{(i)}+1},
    \qquad q_1\mid q_2\mid\cdots\mid q_J.
    $$

    Let $g_i$ be the normalized indicator of the $i$-th residue block and put $Q_i(x)=\exp(-|\widehat g_i(x)|)$. The non-Archimedean MPS test function is assembled from

    $$
    \Phi_J=
    \widehat g_J+\widehat g_{J-1}Q_J+\cdots+
    \widehat g_1Q_2\cdots Q_J.
    $$

    The decisive observation is that $|\widehat g_i|$ is $1/q_i$-periodic, hence $\widehat Q_i$ is supported on $q_i\mathbb Z$. Multiplying an earlier block by a later $Q_k$ shifts its frequencies only by multiples of $q_k$. Since $q_i\mid q_k$, the earlier block cannot escape its residue class modulo $q_i$. Main inner products contribute one controlled unit per block, while geometric fibre growth makes the cross terms summable.

    That is Bedert’s genuine no-leakage mechanism. It is more precise than the slogan “use many congruence classes and patch them.”

    One further qualification prevents a common misstatement. Bedert obtains a suitable $F_4$-isomorphic model $B$ of the original set and proves the required $L^1$ bound there. The norm $\|F_A\|_1$ is not claimed to be invariant under an $F_4$-isomorphism. What is preserved is the four-term additive information needed for sum-freeness, so $S(A)=S(B)$.

    Bedert’s inverse output is also stronger than the numerical lower bound. If $S(A)\le N/3+C$, his results force low additive dimension, a dense $F_4$-model, large additive energy in every substantial subset, and a “99% Structure Theorem” decomposing all but a small exceptional set into large small-doubling pieces. This is a genuine inverse theorem for host sets resisting sum-free extraction, but not an edit-distance classification by one canonical extremizer.

    5. The real bridge: localize, build, prevent leakage, sum

    Parallel flowcharts for Bedert's p-adic Fourier proof and the logarithmic hypergraph construction
    Figure 4. The analogy is architectural. Bedert preserves Fourier residue lanes through nested moduli; the multiplicative construction preserves low codegrees through matching colours and unique scale signatures.

    The dictionary is now clean:

    Role Bedert’s sum-free proof Strongly $2$-primitive proof
    Underlying relation Linear equation in an abelian group Coordinatewise domination of prime valuations
    Local coordinates Exact small-prime valuations and unit residues Logarithmic sizes of prime factors
    Local block Joint $p$-adic fibre and MPS block Scale cell and properly coloured graph
    No interference Nested moduli preserve Fourier support Colour matchings and unique cell signatures preserve linearity
    Accumulation One controlled inner product per fibre One triple per admissible lower pair
    Inverse output Additive dimension, energy, small-doubling pieces Private factors and a cover-free exponent core

    This is a substantial connection, but it is a connection of proof design. Bedert’s Fourier projection has no direct analogue for the divisor partial order. The most plausible transfer is therefore not a line-by-line proof but a research program: find a multiplicative localization, an isolation-or-descent alternative, and a no-leakage invariant that survives repeated prime powers and larger supports.

    6. The first stability guess is false

    The lower construction suggests a tempting statement: every set with

    $$
    |A|\ge
    \pi(n)+\left(\frac{27}{2}-o(1)\right)\Sigma_n
    $$

    should differ in only $o(\Sigma_n)$ places from almost all primes plus a near-optimal linear family of triple products. Two examples show why this is too rigid.

    6.1 A private-prime lift

    Start with a near-optimal linear triple family $\mathcal H_n$ using odd primes below $n/2$. Keep primes above $n/2$, replace every unused prime $p\le n/2$ by $2p$, and retain the triple products. Each $2p$ has the private prime divisor $p$; no other chosen element contains that prime. The triple products are still protected by linearity. The new set remains strongly $2$-primitive and has the same second-order asymptotic size.

    But essentially every prime below $n/2$ has been replaced by a composite. The symmetric-difference distance from the all-primes model is

    $$
    \Theta(\pi(n/2))=\Theta\!\left(\frac n{\log n}\right),
    $$

    and

    $$
    \frac{\pi(n/2)}{\Sigma_n}
    \asymp n^{1/3}\log n\longrightarrow\infty.
    $$

    So literal edit-distance stability fails by much more than the scale of the second-order term.

    6.2 A nonlinear sunflower at almost no cost

    Even if we insist on squarefree triple products, linearity need not hold in the original representation. Add a sunflower

    $$\{\{2,3,r\}:r\in R\}$$

    for many private petals $r$. Remove the singleton primes $2,3,r$ and add the products $6r$. The net cardinality loss is only $2$, while the support family contains $\asymp n/\log n$ edges sharing the pair $\{2,3\}$. It is $2$-cover-free because every target has its own private petal, but any linear subfamily contains at most one sunflower edge.

    The two examples have the same cause: degree-one prime coordinates carry a huge amount of neutral decoration. Stability can only become true after quotienting out this freedom.

    7. Near equality in the upper bound is an exact star forest

    Fix the chosen factorizations $a=u_av_a$ from the multiplicative basis proof. Make a graph $G_A$ with vertex set $\mathcal B$, one edge $\{u_a,v_a\}$ for each $a\in A$, and allow loops.

    Exact factor-graph theorem. Every nonloop edge has a degree-one endpoint, and every loop is an isolated component. Thus $G_A$ is a disjoint union of stars and isolated loops. If $U$ is the number of unused vertices and $c(G_A)$ is the number of nonloop star components, then

    $$|\mathcal B|-|A|=U+c(G_A).$$

    The proof is the private-factor lemma read without discarding information. If a nonloop edge $\{u,v\}$ had both endpoints in other edges, the corresponding two elements would have product divisible by $uv=a$. If a loop at $u$ met another edge, then $u^2$ would divide the square of the other element.

    A factor graph decomposed into stars, an isolated loop and unused vertices
    Figure 5. The upper-bound defect is counted exactly: an unused basis coordinate costs one, and each nonloop star component costs one.

    More quantitatively,

    $$
    \bigl|\{b\in\mathcal B\setminus\mathcal B_0:
    \deg_{G_A}(b)\ne1\}\bigr|
    \le |\mathcal B|-|A|.
    $$

    If

    $$
    |A|\ge
    \pi(n)+\left(\frac{27}{2}-\delta_n\right)\Sigma_n,
    $$

    then the defect $D_n=|\mathcal B|-|A|$ is at most $(\delta_n+o(1))\Sigma_n$. Hence almost every secondary coordinate in both $\mathcal B_2$ and $\mathcal B_3$ occurs in exactly one chosen factor pair. This is a rigorous stability theorem for the upper-bound certificate. It is not yet a classification of the original integers.

    8. Private-prime compression gives the correct normal form

    Suppose a prime $p$ divides $a\in A$ and divides no other element of $A$. Replacing $a$ by $p$ preserves the cardinality and strong $2$-primitivity. The new target $p$ cannot divide a product of two other elements because neither contains $p$; for any unchanged target, replacing $a$ by a divisor only decreases the product available to cover it.

    For primes $p\gt n^{3/5}$, the basis factorization can be chosen to expose $p$ in every multiple. Thus a degree-one large-prime coordinate is genuinely a globally private prime and can be contracted. Contract all such coordinates simultaneously.

    Normalized stability modulo private-prime lifts. From every strongly $2$-primitive $A\subseteq[1,n]$ satisfying the near-extremal bound above, private-prime contractions produce a strongly $2$-primitive set $\widetilde A$ of the same size with

    $$\widetilde A=(\mathbb P\cap[1,n]\setminus V)\cup C,\qquad \operatorname{supp}(c)\subseteq V\quad(c\in C).$$

    where

    $$|V|\le(\delta_n+o(1))\Sigma_n,\qquad |C|-|V|=|A|-\pi(n).$$

    In particular, if $\delta_n=o(1)$, then

    $$
    |V|=o(\Sigma_n),
    \qquad
    |C|=\left(\frac{27}{2}+o(1)\right)\Sigma_n.
    $$

    The geometry is now transparent. Before compression, the leading $\pi(n)$ coordinates may be represented by arbitrary private multiples. After compression, almost every prime appears literally. Every residual composite is supported entirely on the small exceptional prime hub $V$, and its excess over the missing primes is exactly the second-order gain.

    Private-prime lifts contracted to a normal form with singleton primes and a small exceptional hub
    Figure 6. Literal stability fails on the left. The quotient by private-prime contractions produces the exact normal form on the right.

    There is one more unconditional consequence. At most $|V|$ members of $C$ have support of size at most two. Indeed, any one- or two-prime residual element must have a valuation coordinate in which it is a strict global maximum; assign the element to such a prime. Two elements cannot receive the same prime. Therefore, in a normalized near-extremizer, all but $o(\Sigma_n)$ residual composites have at least three distinct prime divisors.

    This is close to the hoped-for triple picture, but it does not prove that the typical element has exactly three prime factors, that it is squarefree, or that its prime factors lie near $n^{1/3}$.

    9. Conditional stability when the residual core is made of triples

    Now impose an additional hypothesis:

    Squarefree-triple hypothesis. All but $o(\Sigma_n)$ elements of $C$ are products of three distinct primes.

    Let $\mathcal H$ be the support family of those triples. Strong $2$-primitivity says precisely that it is $2$-cover-free:

    $$
    E\nsubseteq F\cup G
    \qquad(E,F,G\in\mathcal H,\ E\notin\{F,G\}),
    $$

    with $F=G$ allowed.

    9.1 Pruning to a linear core

    If two triples $E,F$ share a pair and $x$ is the third vertex of $E$, then $x$ has degree one. Otherwise another edge $G$ containing $x$, together with $F$, would cover $E$. Delete such a petal edge whenever a repeated pair occurs. Every deletion removes at least one vertex as well as one edge, so the excess $|\mathcal H|-|V(\mathcal H)|$ does not decrease.

    The result is a linear subfamily $\mathcal L$ with

    $$
    |\mathcal L|-|V(\mathcal L)|
    \ge |\mathcal H|-|V(\mathcal H)|,
    $$

    and only $o(\Sigma_n)$ edges are lost in the normalized near-extremal setting. Notice the nuance: linearity is recovered after pruning private petals; it is not asserted for every cover-free representation.

    9.2 Saturating the feasible pair space

    Define

    $$
    \mathcal D_n=
    \{(p,q):p\lt q\text{ prime and }pq^2\le n\}.
    $$

    Sort an edge of $\mathcal L$ as $p\lt q\lt r$ and map it to $(p,q)$. Linearity makes this map injective. Since $r\gt q$ and $pqr\le n$, the image lies in $\mathcal D_n$. Direct counting gives

    $$
    |\mathcal D_n|
    =\left(\frac{27}{2}+o(1)\right)\Sigma_n.
    $$

    The range $q\le n^{1/3}$ contributes $(9/2+o(1))\Sigma_n$; the range $n^{1/3}\lt q\le n^{2/5}$ contributes $(9+o(1))\Sigma_n$; the tail is negligible. Since the linear core already has $(27/2-o(1))\Sigma_n$ edges, its lower-pair map misses only $o(\Sigma_n)$ feasible pairs.

    Moreover, for every fixed $\varepsilon\gt0$, all but $o(\Sigma_n)$ edges satisfy

    $$
    n^{1/3-\varepsilon}
    \le p,q,r\le
    n^{1/3+\varepsilon}.
    $$

    This is the desired scale stability. It does not imply uniqueness. Different one-factorizations or different proper edge-colourings can change $\Theta(\Sigma_n)$ triples while preserving the same pair occupancy and extremal count. The stable object is the saturated pair space, not an individual colouring.

    10. The remaining inverse problem is weighted and cover-free

    After normalization, write every $c\in C$ as an exponent vector

    $$
    \nu(c)=(v_p(c))_{p\in V}\in\mathbb Z_{\ge0}^{V}.
    $$

    Strong $2$-primitivity becomes the weighted cover-free condition

    $$
    \nu(c)\nleq\nu(c_1)+\nu(c_2)
    \quad\text{coordinatewise}
    $$

    whenever $c\notin\{c_1,c_2\}$. We know that

    $$
    |C|=\left(\frac{27}{2}+o(1)\right)\Sigma_n,
    \qquad |V|=o(\Sigma_n),
    $$

    and that almost every vector has support at least three. What remains unknown is the sharper assertion

    $$
    \#\{c\in C:c\text{ is not a squarefree product of three primes}\}
    =o(\Sigma_n).
    $$

    This gap may conceal genuine alternative near-extremizers. If a triple $pqr$ has product slack, one can contemplate replacing it by $2pqr$ while deleting the singleton prime $2$. For a linear support family, the underlying three private coordinates still prevent coverage. It is not known whether a positive density of such bounded-hub multiplier layers can coexist globally near the $27/2$ threshold, or whether cross-layer collisions force them to be negligible.

    Open normalized stability problem. Classify weighted $2$-cover-free exponent families under the product constraint $c\le n$ whose excess is $(27/2-o(1))\Sigma_n$. Decide whether the residual core is predominantly squarefree of support three, or whether bounded-hub multiplier layers produce genuinely different normalized near-extremizers.

    10.1 What a Bedert-style descent would need

    The additive proof suggests three design requirements, not three ready-made lemmas.

    1. A local order projection. One needs to isolate a valuation or support layer while controlling contamination from lower exponent vectors. A divisor-lattice zeta or Möbius transform is a possible language, but no analogue of Bedert’s $L^1$-contractive residue projection is presently available.
    2. An isolation-or-descent alternative. If the secondary basis slots do not have predominantly prime cofactors, the argument should descend to a structured hub of repeated factors with a quantitative gain or a summable loss.
    3. A no-leakage invariant. Low codegree and scale signatures work for squarefree triples. A general invariant must survive repeated exponents and supports of size at least four.

    The star-forest theorem is already a first inverse statement of this kind: near equality forces almost every basis coordinate to be a leaf attached to a small collection of hubs. The hard step is to turn that certificate-level hub structure into a classification of the original weighted exponent vectors.

    11. What is proved, conditional, and open

    • Published/preprint additive result: Bedert proves $S(A)\ge |A|/3+c\log\log|A|$, together with inverse information involving additive dimension, dense $F_4$-models, energy and a $99\%$ decomposition into large small-doubling pieces.
    • Supplied 2026 manuscript: the claimed $27/2$ asymptotic follows from the multiplicative-basis upper bound and the logarithmic-cell hypergraph construction described above. The status caveat at the beginning remains in force.
    • Unconditional stability developed in the accompanying note: literal edit stability is false; the upper-bound factor graph is an exact star forest; private-prime compression gives the normal form $(\mathbb P\setminus V)\cup C$; and almost every residual composite has at least three distinct prime divisors.
    • Conditional stability: if almost all residual composites are squarefree triples, pruning gives a near-maximal linear core whose lower pairs saturate $\mathcal D_n$, and almost every prime factor lies at scale $n^{1/3+o(1)}$.
    • Open: prove that the normalized residual core is predominantly squarefree of support three, or construct a different near-extremal weighted cover-free core.

    The conceptual moral is simple but useful. Bedert’s $p$-adic fibres and the multiplicative logarithmic cells are not the same mathematical object. Yet both proofs win by choosing coordinates in which local constructions can be made strong and then finding an exact invariant that prevents those constructions from interfering. In the stability problem, the same philosophy says to quotient the neutral directions first. Once private-prime lifts are removed, the true obstruction becomes visible: a weighted cover-free family on a tiny prime hub. That is the right object for the next theorem.

    References and source trail

    1. P. Erdős, On sequences of integers no one of which divides the product of two others and on some related problems, 1938.
    2. P. Erdős, On some applications of graph theory to number-theoretic problems, 1969; see also Erdős Problem #793.
    3. O. Carruth McGehee, L. Pigno and B. Smith, Hardy’s inequality and the $L^1$-norm of exponential sums, Annals of Mathematics 113 (1981), 613-618.
    4. J. Bourgain, Estimates related to sumfree subsets of sets of integers, Israel Journal of Mathematics 97 (1997), 71-92.
    5. B. Bedert, Large sum-free subsets of sets of integers via $L^1$-estimates for trigonometric series, arXiv:2502.08624v1, 12 February 2025.
    6. P. Chojecki, The Second Term for Strongly 2-Primitive Sets, user-supplied five-page manuscript, July 2026; no public identifier located at the time of writing.
  • [后端同步测试] 富媒体、代码与数学公式渲染

    这是一个由你的后端服务器 AI 自动生成的测试文章。用于验证从后端发布富媒体内容到你的前端博客页面的最终渲染效果。

    1. 文本与代码高亮测试

    这里包含粗体斜体,以及一段 Python 模型代码测试:

    import torch
    import torch.nn as nn
    
    class SimpleModel(nn.Module):
        def __init__(self):
            super().__init__()
            self.linear = nn.Linear(128, 10)
    
        def forward(self, x):
            return self.linear(x)

    2. 数学公式测试 (LaTeX 格式)

    很多 AI 算法项目都需要展示数学公式。请检查前端是否成功配置了 KaTeX 或 MathJax 插件来渲染这些公式:

    行内公式测试:著名的质能方程是 $E = mc^2$,激活函数 Sigmoid 定义为 $\sigma(x) = \frac{1}{1 + e^{-x}}$。

    块级独立公式测试(如损失函数):

    $$ L = -\frac{1}{N} \sum_{i=1}^{N} [y_i \log(\hat{y}_i) + (1 – y_i) \log(1 – \hat{y}_i)] $$

    3. 外部图片插入测试

    下面是一张通过外部图床 URL 插入的网络测试图片(未来你可以把图片传到图床或者随代码发给我):

    测试风景图
  • 1

    Let $\left(M^n, g\right)$ be a closed Riemannian manifold with $C^1$-smooth $g_{i j}$. The spectrum of the Laplace operator on $M$ is discrete. There is a sequence of eigenvalues

    $$
    0=\lambda_1<\lambda_2 \leq \lambda_3 \ldots
    $$

    that tend to $\infty$ and a sequence of (real) eigenfunctions $u_k$ such that

    $$
    \Delta_g u_k+\lambda_k u_k=0 .
    $$

    Our enumeration of eigenvalues is non-standard. We start with $\lambda_1=0$ and $u_1=1$ on $M$. The nodal domains of $u_k$ are the connected components of $M \backslash Z_{u_k}$, where $Z_{u_k}$ is the zero set of $u_k\left(Z_{u_k}\right.$ is called the nodal set of $\left.u_k\right)$. The Courant nodal domain theorem states that the $k$-th eigenfunction $u_k$ has at most $k$ nodal domains. If the multiplicity of an eigenvalue is more than 1 , one may enumerate the eigenfunctions corresponding to this eigenvalue in any order. Our main result is the local version of Courant’s theorem.

  • Using Neural Networks to Optimize the Cauchy-Schwarz Inequality: A Generator-Validator Framework


    Introduction

    In many mathematical and engineering problems, we are interested in finding solutions that satisfy certain constraints. A powerful modern paradigm is to train a neural network as a generator that proposes candidate solutions, and use a differentiable validator (i.e., loss function) to evaluate how well they satisfy those constraints. This feedback is then used to update the network via gradient descent.

    In this note, we illustrate this approach by using a neural network to generate vectors that nearly achieve equality in the Cauchy-Schwarz inequality.


    1. Problem Setup: Making Cauchy-Schwarz Nearly Tight

    Recall the Cauchy-Schwarz inequality: $∣⟨x,y⟩∣≤∥x∥⋅∥y∥|\langle \mathbf{x}, \mathbf{y} \rangle| \leq \|\mathbf{x}\| \cdot \|\mathbf{y}\|$

    Equality holds if and only if x\mathbf{x} and y\mathbf{y} are linearly dependent: y=kx\mathbf{y} = k \mathbf{x} for some scalar kk.

    Objective:

    Given an input vector x\mathbf{x}, train a neural network NN to output a vector y=N(x)\mathbf{y} = N(\mathbf{x}) such that x,y\mathbf{x}, \mathbf{y} are as close to colinear as possible.


    2. Generator: Neural Network Design

    Let N(⋅)N(\cdot) be a feedforward neural network (e.g. MLP) with:

    • Input: nn-dimensional vector x∈Rn\mathbf{x} \in \mathbb{R}^n
    • Output: nn-dimensional vector y=N(x;θ)\mathbf{y} = N(\mathbf{x}; \theta)
    • Structure: Simple MLP with 1–2 hidden layers (ReLU), and a linear output layer (no activation)

    3. Validator: Loss Function Design

    To measure how close x,y\mathbf{x}, \mathbf{y} are to colinearity, use cosine similarity: cos⁡(θ)=⟨x,y⟩∥x∥⋅∥y∥+ε\cos(\theta) = \frac{\langle \mathbf{x}, \mathbf{y} \rangle}{\|\mathbf{x}\| \cdot \|\mathbf{y}\| + \varepsilon}

    We define the loss as: L(x,y)=1−∣⟨x,y⟩∥x∥⋅∥y∥+ε∣L(\mathbf{x}, \mathbf{y}) = 1 – \left|\frac{\langle \mathbf{x}, \mathbf{y} \rangle}{\|\mathbf{x}\| \cdot \|\mathbf{y}\| + \varepsilon}\right|

    • L=0L = 0 when x\mathbf{x} and y\mathbf{y} are perfectly aligned or anti-aligned
    • ε≪1\varepsilon \ll 1 is a small constant for numerical stability

    This validator provides a differentiable measure of alignment quality.


    4. Training Procedure

    1. Data generation: Sample random input vectors x\mathbf{x} from e.g. N(0,I)\mathcal{N}(0, I)
    2. Forward pass: Compute y=N(x)\mathbf{y} = N(\mathbf{x})
    3. Loss computation: Evaluate L(x,y)L(\mathbf{x}, \mathbf{y})
    4. Backpropagation: Compute ∇θL\nabla_\theta L and update θ\theta using an optimizer (e.g. Adam)
    5. Repeat until convergence

    At the end of training, the network learns to generate vectors y\mathbf{y} nearly colinear with x\mathbf{x}, thus making the Cauchy-Schwarz inequality nearly tight.


    5. General Framework: Generator + Validator

    This method exemplifies a general and powerful pattern in deep learning:

    ComponentRoleDescription
    Neural Network NNGenerator / SolverMaps input (or noise) to a candidate solution
    Validator VVLoss / Constraint FunctionEvaluates how well the candidate satisfies the constraints (must be differentiable)
    OptimizerLearning EngineUses gradients to update NN so that the solutions improve over time

    6. Applications and Extensions

    This framework generalizes to many domains:

    • Inequality tightness: AM-GM, Hölder, Jensen inequalities
    • Constraint solving: linear/quadratic programming, geometric constraints
    • Functional problems: e.g. finding extremals in calculus of variations
    • Neural symbolic systems: e.g. generating logic-constrained expressions
    • Inverse design: input-to-output mappings constrained by physical or mathematical laws

    Conclusion

    Training a neural network to minimize a differentiable validator is a powerful method to learn constrained solutions. The Cauchy-Schwarz example shows how even classical inequalities can be embedded into a modern optimization loop, potentially aiding in automated reasoning, symbolic learning, or mathematical discovery.

    Would you like this exported to PDF with rendered math? Or should I write a minimal PyTorch implementation to match?

  • 世界,您好!

    欢迎使用 WordPress。这是您的第一篇文章。编辑或删除它,然后开始写作吧!

  • Ramanujan-Nagell theorem:平方数与 2^n 相差 7 的有限性

    旧博客原文

    原题:The Ramanujan-Nagell Theorem: Understanding the Proof

    The Ramanujan-Nagell Theorem: Understanding the Proof


    补充说明

    以下是新整理的中文说明;上方旧博客原文保持不变。

    Ramanujan-Nagell theorem 研究的是一个看起来非常小的指数丢番图方程:

    $$x^2+7=2^n.$$

    它的结论是整数解只有有限个,而且正整数解恰好为

    $$(x,n)=(1,3),(3,4),(5,5),(11,7),(181,15).$$

    Ramanujan-Nagell theorem:平方数与 2^n 相差 7 的有限性
    Ramanujan-Nagell 方程把初等同余、二次域分解和 Lucas 序列约束压缩在同一个指数丢番图问题中。

    1. 初等筛选

    先看奇偶性。若 $x$ 为偶数,则左边 $x^2+7$ 为奇数,不可能等于 $2^n$;所以 $x$ 必为奇数。再看模 $8$,奇数平方恒为 $1$,所以 $x^2+7\equiv0\pmod 8$,这只说明 $n\ge3$,但已经排除了很多无意义情形。

    2. 代数数论中的分解

    真正有力的观察是把方程写成

    $$(x+\sqrt{-7})(x-\sqrt{-7})=2^n.$$

    在 $\mathbb Q(\sqrt{-7})$ 的整数环中,$2$ 可以分解成两个共轭因子。由于这个二次域的类数很小,理想层面的分解可以被提升为元素层面的约束,于是 $x+\sqrt{-7}$ 必须接近某个基本元素的 $n$ 次幂。

    3. Lucas 序列的出现

    把共轭相减,得到的不是任意等式,而是一个 Lucas 型序列项必须等于很小的数:

    $$\frac{\alpha^n-\bar\alpha^n}{\alpha-\bar\alpha}=\pm 1\quad\text{or}\quad \pm 7.$$

    这种递推序列增长很快,同时在模意义下有强限制。少数小 $n$ 需要直接检查,大 $n$ 则被递推结构和同余条件排除。

    4. 为什么这条定理有代表性

    这类问题的典型形状是:初等同余给出粗过滤,二次域分解给出结构,最后用 Lucas 序列或线性形式估计把无限可能压成有限检查。Ramanujan-Nagell 方程的漂亮之处在于,所有这些工具都集中在一个非常短的公式里。

  • BSD 猜想入门:椭圆曲线、L 函数与 Mordell-Weil rank

    旧博客原文

    原题:A crash introduction to BSD conjecture

    The pdf version is A crash introduction to BSD conjecture .

    We begin with the Weierstrass form of elliptic equation, i.e. look it as an embedding cubic curve in {\mathop{\mathbb P}^2}.

    Definition 1 (Weierstrass form) {E \hookrightarrow \mathop{\mathbb P}^2 }, In general the form is given by,

    \displaystyle E: y^2+a_1xy+a_3y=x^3+a_2x^2+a_4x+a_6 \ \ \ \ \ (1)

    If {char F \neq 2,3}, then, we have a much more simper form,

    \displaystyle y^2=x^3+ax+b, \Delta:=4a^3+27b^2\neq 0. \ \ \ \ \ (2)

    Remark 1

    \displaystyle \Delta(E)=\prod_{1\leq i,\neq j\leq 3}(z_i-z_j)

    Where {z_i^3+az_i+b=0, \forall 1\leq i\leq 3}.

    We have two way to classify the elliptic curve {E} living in a fix field {F}. \paragraph{j-invariant} The first one is by the isomorphism in {\bar F}. i.e. we say two elliptic curves {E_1,E_2} is equivalent iff

    \displaystyle \exists \rho:\bar F\rightarrow \bar F

    is a isomorphism such that {\rho(E_1)=E_2}.

    Definition 2 (j-invariant) For a elliptic curve {E}, we have a j-invariant of {E}, given by,

    \displaystyle j(E)=1728\frac{4a^3}{4a^3+27b^2} \ \ \ \ \ (3)

    Why j-invariant is important, because j-invariant is the invariant depend the equivalent class of {E} under the classify of isomorphism induce by {\bar F}. But in one equivalent class, there also exist a structure, called twist.

    Definition 3 (Twist) For a elliptic curve {E:y^2=x^3+ax+b}, all elliptic curve twist with {E} is given by,

    \displaystyle E^{(d)}:y^2=x^3+ad^2x+bd^3 \ \ \ \ \ (4)

    So the twist of a given elliptic curve {E} is given by:

    \displaystyle H^1(Gal(\bar F/ F), Aut(E_{\bar F})) \ \ \ \ \ (5)

    Remark 2 Of course a elliptic curve {E:y^2=x^3+ax+b} is the same as {E:y^2=x^3+ad^2x+bd^4}, induce by the map {\mathop{\mathbb P}^1\rightarrow \mathop{\mathbb P}^1, (x,y,1)\rightarrow (x,dy,1)}.

    But this moduli space induce by the isomorphism of {F} is not good, morally speaking is because of the abandon of universal property. see \cite{zhang}. \paragraph{Level {n} structure} We need a extension of the elliptic curve {E}, this is given by the integral model.

    Definition 4 (Integral model) {s:=Spec(\mathcal{O}_F)}, {E\rightarrow E_s}. {E_s} is regular and minimal, the construction of {E_s} is by the following way, we first construct {\widetilde{E_s} } and then blow up. {\widetilde E_s} is given by the Weierstrass equation with coefficent in {\mathcal{O}_F}.

    Remark 3 The existence of integral model need Zorn’s lemma.

    Definition 5 (Semistable) the singularity of the minimal model of {E} are ordinary double point.

    Remark 4 Semistable is a crucial property, related to Szpiro’s conjecture.

    Definition 6 (Level {n} structure)

    \displaystyle \phi: ({\mathbb Z}/n{\mathbb Z})_s^2\longrightarrow E[N] \ \ \ \ \ (6)

    {P=\phi(1,0), Q=\phi(o,1)} The weil pairing of {P,Q} is given by a unit in cycomotic fields, i.e. {<P,Q>=\zeta_N\in \mu_{N}(s)}

    What happen if {k={\mathbb C}}? In this case we have a analytic isomorphism:

    \displaystyle E({\mathbb C})\simeq {\mathbb C}/\Lambda \ \ \ \ \ (7)

    Given by,

    \displaystyle {\mathbb C}/\Lambda \longrightarrow \mathop{\mathbb P}^2 \ \ \ \ \ (8)

    \displaystyle z\longrightarrow (\mathfrak{P}(z), \mathfrak{P}'(z), 1 ) \ \ \ \ \ (9)

    Where {\mathfrak{P(z)}=\frac{1}{z^2}+\sum_{\lambda\in \Lambda,\lambda\neq 0}(\frac{1}{(z-\lambda)^2}-\frac{1}{\lambda^2})}, and the Weierstrass equation {E} is given by {y^2=4x^3-60G_4(\Lambda)x-140G_6(\Lambda)}. The full n tructure of it is given by {{\mathbb Z}+{\mathbb Z}\lambda} and the value of {P,Q}, i.e.

    \displaystyle P=\frac{1}{N}, Q=\frac{\tau}{N} \ \ \ \ \ (10)

    Where {\tau} is induce by

    \displaystyle \Gamma(N):=ker(SL_2({\mathbb Z})\rightarrow SL_2({\mathbb Z}/n{\mathbb Z})) \ \ \ \ \ (11)

    The key point is following:

    Theorem 7 {k={\mathbb C}}, the moduli of elliptic curves with full level n-structure is identified with

    \displaystyle \mu_N^*\times H/\Gamma(N) \ \ \ \ \ (12)

    Now we discuss the Mordell-Weil theorem.

    Theorem 8 (Mordell-Weil theorem)

    \displaystyle E(F)\simeq {\mathbb Z}^r\oplus E(F)_{tor}

    The proof of the theorem divide into two part:

    1. Weak Mordell-Weil theorem, i.e. {\forall m\in {\mathbb N}}, {E(F)/mE(F)} is finite.
    2. There is a quadratic function,

      \displaystyle \|\cdot\|: E(F)\longrightarrow {\mathbb R} \ \ \ \ \ (13)

      {\forall c\in {\mathbb R}}, {E(F)_c=\{P\in E(F), \|P\|<c\}} is finite.

    Remark 5 The proof is following the ideal of infinity descent first found by Fermat. The height is called Faltings height, introduce by Falting. On the other hand, I point out, for elliptic curve {E}, there is a naive height come from the coefficient of Weierstrass representation, i.e. {\max\{|4a^3|,|27b^2|\}}.

    While the torsion part have a very clear understanding, thanks to the work of Mazur. The rank part of {E({\mathbb Q})} is still very unclear, we have the BSD conjecture, which is far from a fully understanding until now.

    But to understanding the meaning of the conjecture, we need first constructing the zeta function of elliptic curve, {L(s,E)}.

    \paragraph{Local points} We consider a local field {F_v}, and a locally value map {F\rightarrow F_{\nu}}, then we have the short exact sequences,

    \displaystyle 0\longrightarrow E^0(F_{\nu})\longrightarrow E(F_{\nu})=E_s(\mathcal{O}_F)\longrightarrow E_s(K_0)\longrightarrow 0 \ \ \ \ \ (14)

    Topologically, we know {E(F_{\nu})} are union of disc indexed by {E_s(k_{\nu})},

    \displaystyle |E_s(k_{\nu})| \sim q_{\nu}+1=\# \mathop{\mathbb P}^1(k_{\nu})

    . Define {a_{\nu}=\# \mathop{\mathbb P}^1(k_{\nu})-|E_s(k_{\nu})|}, then we have Hasse principle:

    Theorem 9 (Hasse principle)

    \displaystyle |a_{\nu}|\leq 2\sqrt{q_{\nu}} \ \ \ \ \ (15)

    Remark 6 I need to point out, the Hasse principle, in my opinion, is just a uncertain principle type of result, there should be a partial differential equation underlying mystery.

    So count the points in {E(F)} reduce to count points in {H^1(F_{\nu},E(m))}, reduce to count the Selmer group {S(E)[m]}. We have a short exact sequences to explain the issue.

    \displaystyle 0\longrightarrow E(F)/mE(F) \longrightarrow Sha(E)[m] \longrightarrow E(F)/mE(F)\longrightarrow 0 \ \ \ \ \ (16)

    I mention the Goldfold-Szipiro conjecture here. {\forall \epsilon>0}, there {\exists C_{\epsilon}(E)} such that:

    \displaystyle \# (E)\leq c_{\epsilon}(E)N_{E/{\mathbb Q}}(N)^{\frac{1}{2}+\epsilon} \ \ \ \ \ (17)

    \paragraph{L-series} Now I focus on the construction of {L(s,E)}, there are two different way to construct the L-series, one approach is the Euler product.

    \displaystyle L(s,E)=\prod_{\nu: bad}(1-a_{\nu}q_{\nu}^{-s})^{-1}\cdot \prod_{\nu:good}(1-a_{\nu}q_{\nu}^{-s}+q_{\nu}^{1-2s})^{-1} \ \ \ \ \ (18)

     

    Where {a_{\nu}=0,1} or {-1} when {E_s} has bad reduction on {\nu}.

    The second approach is the Galois presentation, one of the advantage is avoid the integral model. Given {l} is a fixed prime, we can consider the Tate module:

    \displaystyle T_l(E):=\varprojlim_{l^n} E[l^n] \ \ \ \ \ (19)

    Then by the transform of different embedding of {F\hookrightarrow \bar F}, we know { T_{l}(E)/Gal(\bar F/F)}, decompose it into a lots of orbits, so we can define {D_{\nu}}, the decomposition group of {w}(extension of {\nu} to {\bar F}). We define {I_{\nu}} is the inertia group of {D_{\nu}}.

    Then {D_{\nu}/I_{\nu}} is generated by some Frobenius elements

    \displaystyle Frob{\nu}x\equiv x^{q_{\nu}} (mod w),\forall x\in \mathcal{O}_{\bar Q} \ \ \ \ \ (20)

    So we can define

    \displaystyle L_{\nu}(s,E)=(1-q_{\nu}^{-s}Frob_{\nu}|T_{l}(E)^{I_{\nu}})^{-1} \ \ \ \ \ (21)

    And then {L(s,E)=\prod_{\nu}L_{\nu}(s,E)}.

    Faltings have proved {L_{\nu}(s,E)} is the invariant depending the isogenous class in the follwing meaning:

    Theorem 10 (Faltings) {L_{\nu}(s,E)} is an isogenous ivariant, i.e. {E_1} isogenous to {E_2} iff {\forall a.e. \nu}, {L_{\nu}(s,E_1)=L_{\nu}(s,E_2)}.

    \displaystyle L(s,E)=L(s-\frac{1}{2},\pi ) \ \ \ \ \ (22)

    Where {\pi} come from an automorphic representation for {GL_2(A_F)}. Now we give the statement of BSD onjecture. {R} is the regulator of {E}, i.e. the volume of fine part of {E(F)} with respect to the Neron-Tate height pairing. {\Omega} be the volume of {\prod_{v|\infty}F(F_v)} Then we have,

    1. {ord_{s=1}L(s,E)=rank E(F)}.
    2. {|Sha(E)|<\infty}.
    3. {\lim_{s\rightarrow 0}L(s,E)(s-1)^{-rank(E)}=c\cdot \Omega(E)\cdot R(E)\cdot |Sha(E)|\cdot |E(F)_{tor}|^{-2}}

    Here {c} is an explictly positive integer depending only on {E_{\nu}} for {\nu} dividing {N}.

     


    补充说明

    以下是新整理的中文说明;上方旧博客原文保持不变。

    Birch and Swinnerton-Dyer 猜想把椭圆曲线的有理点群和它的 $L$ 函数在 $s=1$ 处的零点阶联系起来。它是数论中最核心的桥之一:一边是 Diophantine 方程的解,另一边是解析函数的特殊值。

    BSD 猜想入门:椭圆曲线、L 函数与 Mordell-Weil rank
    BSD 猜想连接椭圆曲线有理点群的 rank 与 $L$ 函数在 $s=1$ 的零点阶。

    1. Weierstrass 形式

    在特征不是 $2,3$ 的域上,椭圆曲线常写成

    $$E:y^2=x^3+ax+b,$$

    并要求判别式

    $$\Delta=-16(4a^3+27b^2)\ne0.$$

    非零判别式保证曲线光滑。椭圆曲线不仅是代数曲线,它的点还带有 Abelian group 结构。

    2. $j$-invariant 与 twist

    $j$-invariant 分类椭圆曲线在代数闭包上的同构类:

    $$j(E)=1728\frac{4a^3}{4a^3+27b^2}.$$

    但在固定域 $K$ 上,同一个 $j$-invariant 可能对应不同的 twists。twist 说明“几何上同构”和“在基域上同构”之间有差别。

    3. Mordell-Weil theorem

    若 $E$ 定义在 $\mathbb Q$ 上,则 Mordell-Weil theorem 说

    $$E(\mathbb Q)\cong E(\mathbb Q)_{\operatorname{tors}}\oplus\mathbb Z^r.$$

    整数 $r$ 称为 Mordell-Weil rank。求 $r$ 是理解有理点结构的核心问题。

    4. Hasse-Weil L 函数

    对每个好素数 $p$,令

    $$a_p=p+1-\#E(\mathbb F_p).$$

    椭圆曲线的 $L$ 函数由 Euler product 组成:

    $$L(E,s)=\prod_p(1-a_pp^{-s}+p^{1-2s})^{-1}$$

    再在坏素数处加入修正因子。模性定理保证这个 $L$ 函数有解析延拓和函数方程。

    5. BSD 猜想

    BSD 猜想最核心的断言是

    $$\operatorname{rank}E(\mathbb Q)=\operatorname{ord}_{s=1}L(E,s).$$

    更精细的版本还给出 $L(E,s)$ 在 $s=1$ 处首项系数,涉及 regulator、Tate-Shafarevich group、Tamagawa numbers 和 torsion subgroup。

    这个猜想的意义在于,它把有理点这个离散、代数的问题,转化成了 $L$ 函数特殊值这个解析问题。许多数论现代方法都在这座桥上来回移动。

  • $SL_2(\mathbb Z)$ 与 congruence subgroups:分式线性作用和 modular curves

    旧博客原文

    原题:SL_2(Z) and its congruence subgroups

    The pdf version is SL_2(Z) and its congruence subgroups.

    We know we can always do the following thing:

    \displaystyle R\ commutative\ ring \longrightarrow \ "general\ linear\ group" \ GL_2(R) \ \ \ \ \ (1)

    Where

    \displaystyle GL_2(R):=\{\begin{pmatrix} a & b \\ c & d \end{pmatrix}: det \begin{pmatrix} a & b \\ c & d \end{pmatrix}=R^*, a,b,c,d\in R\} \ \ \ \ \ (2)

     

    Remark 1 Why it is {R^*} but not 1? if it is 1, then the action {R/GL_2(R)} distribute is not trasitive on {R}, i.e. every element in unite group present a connected component

    Now we consider the subgroup {SL_2(R)\subset GL_2(R)}.

    \displaystyle SL_2(R)=\{\begin{pmatrix} a & b\\ c & d \end{pmatrix}: det \begin{pmatrix} a & b\\ c & d \end{pmatrix}=1, a,b,c,d\in R \} \ \ \ \ \ (3)

    We are most interested in the case {R={\mathbb Z}, {\mathbb Z}/n{\mathbb Z}}. So how to investigate {SL_2({\mathbb R})}? We can look at the action of it on something, for particular, we look at the action of it on Riemann sphere, i.e. { \hat {\mathbb C}/({\mathbb R})} given by fraction linear map:

    \displaystyle g(z):=\frac{az+b}{cz+d}, g(\infty)=\frac{a}{c} \ \ \ \ \ (4)

    Remark 2 What is fraction linear map? This action carry much more information than the action on vector, thanks for the exist of multiplication in {{\mathbb C}} and the algebraic primitive theorem. Due to I always looks the fraction linear map as something induce by the permutation of the roots of polynomial of degree 2, this is true at least for fix points, and could natural extension. So how about the higher dimension generate? consider the transform of {k-1} tuples induce by polynomial with degree {k}?

    Remark 3

    1. {SL_2({\mathbb R})/\pm I:=PSL_2({\mathbb R})}, then {PSL_2({\mathbb R}) } action faithful on {\hat C}, i.e. except identity, every action is nontrivial. This is easy to be proved, observed,

      \displaystyle \frac{az+b}{cz+d}=z,\forall z\in \hat{\mathbb C}\Longrightarrow \begin{pmatrix} a & b\\ c & d \end{pmatrix}=\begin{pmatrix} 1 & 0\\ 0&1 \end{pmatrix} or \begin{pmatrix} -1&0\\ 0&-1 \end{pmatrix} \ \ \ \ \ (5)

    2. Up half plane {H} is invariant under the action of {PSL_2({\mathbb R})}, i.e. {\forall g\in PSL_2({\mathbb R})}, {gH=H}. The proof is following,
    3. \displaystyle \begin{array}{rcl} Im(\frac{az+b}{cz+d}) & = & Im(\frac{(az+b)(c\bar z+d)}{|cz+d|^2})\\ & = & Im(\frac{ac|z|^2+bc\bar z+adz+bd}{|cz+d|^2})\\ & > & 0,\ due\ to\ ad=bc+1. \end{array}

     

    Now we focus on {SL_2({\mathbb Z})} or the same,{PSL({\mathbb Z})}. All the argument for {SL_2(R)} make sense for

    \displaystyle \Gamma:= SL_2({\mathbb Z}), \bar \Gamma:=SL_2({\mathbb Z})/\pm I \ \ \ \ \ (6)

    Fix {N\in {\mathbb N}}, define,

    \displaystyle \Gamma(N):=\{\begin{pmatrix} a &b\\ c&d \end{pmatrix}, a,d \equiv 1(mod N), b,c\equiv 0(mod N)\} \ \ \ \ \ (7)

    Then {\Gamma(N)} is the kernel of map {SL_2({\mathbb Z})\rightarrow SL_2({\mathbb Z}/n{\mathbb Z})}, i.e. we have short exact sequences,

    \displaystyle 0\longrightarrow \Gamma(N)\longrightarrow \Gamma\longrightarrow SL_2({\mathbb Z}/n{\mathbb Z})\longrightarrow 0 \ \ \ \ \ (8)

    Remark 4 The relationship of {\Gamma(N)\subset \Gamma} is just like {N{\mathbb Z}+1\subset {\mathbb Z}}.

    Definition 1 (Congruence group) A subgroup of {\Gamma} is called a congruence group iff {\exists n\in {\mathbb N}}, {\Gamma(N)\subset G}.

    Example 1 We give two examples of congruence subgroups here.

    1. \displaystyle \Gamma_1(N)=\{\begin{pmatrix} 1& *\\ 0 &1 \end{pmatrix} mod N\} \ \ \ \ \ (9)

    2. \displaystyle \Gamma_0(N)=\{\begin{pmatrix} * & *\\ 0 & * \end{pmatrix}mod N\} \ \ \ \ \ (10)

     

    Definition 2 (Fundamental domain)

    \displaystyle F=\{z\in H:-\frac{1}{2}\leq Re(z)\leq \frac{1}{2}\ and |z|\geq 1\} \ \ \ \ \ (11)

    Now here is a theorem charistization the fundamental domain.

    Theorem 3 This domain {F} is a fundamental domain of {\hat {\mathbb H}/({\mathbb R})}

    Proof: Ths key point is {SL_2({\mathbb Z})} have two generators,

    1. {\tau_a: z\rightarrow z+a, \forall a\in {\mathbb Z}}.
    2. {s: z\rightarrow \frac{1}{z}}.

    Thanks to this two generator exactly divide the action of {\Gamma} on {H} into a lots of scales, then {\Omega} is a fundamental domain is a easy corollary. \Box

    Remark 5 This is not rigorous, {H} need be replace by {\hat H}, but this is very natural to get a modification to a right one.

    Remark 6 {z_1,z_2\in \partial F} are {\Gamma} equivalent iff {Re(z)=\pm \frac{1}{2}} and {z_2=z_1\pm 1} or if {z_1} on the unit circle and {z_2=-\frac{1}{z_1}}

    Remark 7 If {z\in F}, then {\Gamma_z=\pm I} expect in the following three case:

    1. {\Gamma=\pm \{\tau,s\}} if {z=i}.
    2. {\Gamma=\pm\{ I,s\tau, (s\tau)^2\}} if {z=w=-\frac{1}{2}+\frac{\sqrt{-3}}{2}}.
    3. {\Gamma=\pm\{I,\tau s, (\tau s)^2\}} if {z=-\bar w=\frac{1}{2}+\frac{\sqrt{-3}}{2}}.

    Where {\tau=\tau_1}.

    Remark 8 The group {\bar \Gamma=SL_2({\mathbb Z})/\pm I} is generated by the two elements {s}, {\tau}. In other word, any fraction linear transform is a “word” induce by {s,\tau,s^{-1}.\tau^{-1}}. But not free group, we have relationship {s^2=-I,(s\tau)^3=-I}.

    The natural function space on {F} is the memorphic function, under the map: {H\rightarrow D-\{0\}}, it has a {q}-expension,

    \displaystyle f(q)=\sum_{k\in {\mathbb Z}}a_kq^k \ \ \ \ \ (12)

    And there are only finite many negative {k} such that {a_k\neq 0}.


    补充说明

    以下是新整理的中文说明;上方旧博客原文保持不变。

    $SL_2(\mathbb Z)$ 是 modular forms 和 modular curves 的基本群。它通过分式线性变换作用在上半平面上,而 congruence subgroups 则把模 $N$ 的算术信息放进几何商空间。

    $SL_2(\mathbb Z)$ 与 congruence subgroups:分式线性作用和 modular curves
    $SL_2(\mathbb Z)$ 通过分式线性变换作用在上半平面,congruence subgroups 对应 modular curves 的 level structure。

    1. 分式线性作用

    $$\gamma=\begin{pmatrix}a&b\\ c&d\end{pmatrix}\in SL_2(\mathbb Z),$$

    定义

    $$\gamma z=\frac{az+b}{cz+d}.$$

    若 $\operatorname{Im}z>0$,则 $\operatorname{Im}(\gamma z)>0$,所以上半平面 $\mathbb H$ 在作用下不变。

    2. 为什么是 $PSL_2$

    $I$ 和 $-I$ 给出同一个分式线性变换。因此真正忠实作用的是

    $$PSL_2(\mathbb Z)=SL_2(\mathbb Z)/\{\pm I\}.$$

    3. Congruence subgroups

    主同余子群定义为

    $$\Gamma(N)=\ker\bigl(SL_2(\mathbb Z)\to SL_2(\mathbb Z/N\mathbb Z)\bigr).$$

    更常见的还有 $\Gamma_0(N)$ 和 $\Gamma_1(N)$,它们分别对矩阵的某些项加模 $N$ 条件。

    4. Modular curves

    商空间

    $$Y(\Gamma)=\Gamma\backslash\mathbb H$$

    是 modular curve 的开部分。补上 cusps 后得到紧化 $X(\Gamma)$。不同 congruence subgroup 对应不同 level structure。

    5. 算术与几何

    $SL_2(\mathbb Z)$ 的作用把矩阵、分式线性变换、椭圆曲线的 level structure 和 modular forms 连接起来。研究 congruence subgroups,就是研究模 $N$ 算术信息如何改变上半平面商的几何。