分类: Harmonic analysis

  • 离散调和函数:格点 Laplacian、坏点集合与 Liouville 型问题

    旧博客原文

    原题:Discrete harmonic function in Z^n

    There is some gap, in fact I can improve half of the argument of Discrete harmonic function , the pdf version is Discrete harmonic function in Z^n, but I still have some gap to deal with the residue half…

     

    1. The statement of result

    First of all, we give the definition of discrete harmonic function.

    Definition 1 (Discrete harmonic function) We say a function {f: {\mathbb Z}^n \rightarrow {\mathbb R}} is a discrete harmonic function on {{\mathbb Z}^n} if and only if for any {(x_1,...,x_n)\in {\mathbb Z}^n}, we have:

    \displaystyle f(x_1,...,x_n)=\frac{1}{2^n}\sum_{(\delta_1,...,\delta_n )\in \{-1,1\}^n}f(x_1+\delta_1,...,x_n+\delta_n ) \ \ \ \ \ (1)

     

    In dimension 2, the definition reduce to:

    Definition 2 (Discrete harmonic function in {{\mathbb R}^2}) We say a function {f: {\mathbb Z}^2 \rightarrow {\mathbb R}} is a discrete harmonic function on {{\mathbb Z}^2} if and only if for any {(x_1,x_2)\in {\mathbb Z}^2}, we have:

    \displaystyle f(x_1,x_2)=\frac{1}{4}\sum_{(\delta_1,\delta_2)\in \{-1,1\}^2}f(x_1+\delta_1,x_2+\delta_2 ) \ \ \ \ \ (2)

     

    The result establish in \cite{paper} is following:

    Theorem 3 (Liouville theorem for discrete harmonic functions in {{\mathbb R}^2}) Given {c>0}. There exists a constant {\epsilon>0} related to {c} such that, given a discrete harmonic function {f} in {{\mathbb Z}^2} satisfied for any ball {B_R(x_0)} with radius {R>R_0}, there is {1-\epsilon} portion of points {x\in B_R(x_0)} satisfied {|f(x)|<c}. then {f} is a constant function in {{\mathbb Z}^2}.

    Remark 1 This type of result contradict to the intuition, at least there is no such result in {{\mathbb C}}. For example. the existence of poisson kernel and the example given in \cite{paper} explain the issue.

    Remark 2 There are reasons to explain why there could not have a result in {{\mathbb C}} but in {{\mathbb Z}^2},

    1. The first reason is due to every radius {R} there is only {O(R^2)} lattices in {B_R(x)} in {{\mathbb Z}^2} so the mass could not concentrate very much in this setting.
    2. The second one is due to there do not have infinite scale in {{\mathbb Z}^2} but in {{\mathbb C}}.
    3. The third one is the function in {{\mathbb Z}^2} is automatically locally integrable.

     

    The generation is following:

    Theorem 4 (Liouville theorem for discrete harmonic functions in {{\mathbb R}^n}) Given {c>0,n\in {\mathbb N}}. There exists a constant {\epsilon>0} related to {n,c} such that, given a discrete harmonic function {f} in {{\mathbb Z}^n} satisfied for any ball {B_R(x_0)} with radius {R>R_0}, there is {1-\epsilon} portion of points {x\in B_R(x_0)} satisfied {|f(x)|<c}. then {f} is a constant function in {{\mathbb Z}^n}.

    In this note, I give a proof of 4, and explicit calculate a constant {\epsilon_n>0} satisfied the condition in 3, this way could also calculate a constant {\epsilon_n} satisfied 4. and point the constant calculate in this way is not optimal both in high dimension and 2 dimension.

    2. some element properties with discrete harmonic function

    We warm up with some naive property with discrete harmonic function. The behaviour of bad points could be controlled, just by isoperimetric inequality and maximum principle we have following result.

    Definition 5 (Bad points) We divide points of {{\mathbb Z}^n} into good part and bad part, good part {I} is combine by all point {x} such that {|f(x)|<c}, and {J} is the residue one. So {A\amalg B={\mathbb Z}^n}.

    For all {B_R(0)}, we define {J_R:=J\cap B_R(0), I_R=I\cap B_R(0)} for convenient.

    Theorem 6 (The distribution of bad points) For all bad points {J_R} in {B_R(0)}, they will divide into several connected part, i.e.

    \displaystyle J_R=\amalg_{i\in S_R}A_i \ \ \ \ \ (3)

    and every part {A_i} satisfied {A_i\cap \partial B_R(0)\neq \emptyset}.

    Remark 3 We say {A} is connected in {{\mathbb Z}^n} iff there is a path in {A} connected {x\rightarrow y, \forall x,y\in A}.

    Remark 4 the meaning that every point So the behaviour of bad points are just like a tree structure given in the gragh.

    Proof: A very naive observation is that for all {\Omega\subset {\mathbb Z}^n} is a connected compact domain, then there is a function

    \displaystyle \lambda_{\Omega}: \partial \longrightarrow {\mathbb R} \ \ \ \ \ (4)

    such that {\lambda_{\Omega}(x,y)\geq 0, \forall (x,y)\in \Omega \times \mathring{\Omega}}. And we have:

    \displaystyle f(x)=\sum_{y\in\partial \Omega}\lambda_{\Omega}(x,y)(y) \ \ \ \ \ (5)

     

    This could be proved by induction on the diameter if {\Omega}. Then, if there is a connected component of {\Omega} such that contradict to theorem 6 for simplify assume the connected component is just {\Omega}, then use the formula 5we know

    \displaystyle \begin{array}{rcl} \sup_{x\in \Omega}|f(x)| & = & \sup_{x\in \Omega}\sum_{y\in\partial \Omega}\lambda_{\Omega}(x,y)(y) \\ & \leq & \sup_{\partial \Omega}|f(x)| \\ & \leq & c \end{array}

    The last line is due to consider around {\partial \Omega}. But this lead to: {\forall x\in \Omega, |f(x)|<c} which is contradict to the definition of {\Omega}. So we get the proof. \Box

    Now we begin another observation, that is the freedom of extension of discrete harmonic function in {{\mathbb Z}^n} is limited.

    Theorem 7 we can say something about the structure of harmonic function space of {Z^n}, the cube, you will see, if add one value, then you get every value, i.e. we know the generation space of {Z^n}

    Proof: For two dimension case, the proof is directly induce by the graph. The case of {n} dimensional is similar. \Box

    Remark 5 The generation space is well controlled. In fact is just like n orthogonal direction line in n dimensional case.

    3. sktech of the proof for \ref

    }

    The proof is following, by looking at the following two different lemmas establish by two different ways, and get a contradiction.

    \paragraph{First lemma}

    Lemma 8 (Discrete poisson kernel) the poisson kernel in {{\mathbb Z}^n}. We point out there is a discrete poisson kernel in {{\mathbb Z}^n}, this is given by:

    \displaystyle f(x)=\sum_{y\in \partial B_R(z)}\lambda_{B_R(z)}(x,y)(y) \ \ \ \ \ (6)

    And the following properties is true:

    1. {\lambda_{B_R(z+h)}(x+h,y+h)=\lambda_{B_R(z)}(x,y)} , {\forall x\in \Omega, h\in {\mathbb Z}^n}.
    2. \displaystyle \lambda_{B_R(z)}(x,y)\rightarrow \rho_R(x,y) \ \ \ \ \ (7)

    Remark 6 The proof could establish by central limit theorem, brown motion, see the material in the book of Stein \cite{stein}. The key point why this lemma 8 will be useful for the proof is due to this identity always true {\forall x\in B_R(0)}, So we will gain a lots of identity, These identity carry information which is contract by another argument.

    \paragraph{Second lemma} The exponent decrease of mass.

    Lemma 9 The mass decrease at least for exponent rate.

    Remark 7 the proof reduce to a random walk result and a careful look at level set, reduce to the worst case by brunn-minkowski inequality or isoperimetry inequality.

    \paragraph{Final argument} By looking at lemma 1 and lemma 2, we will get a contradiction by following way, first the value of {f} on {\partial B_R(0)} increasing too fast, exponent increasing by lemma2, but on the other hand, it lie in the integral expresion involve with poisson kernel, but the pertubation of poisson kernel is slow, polynomial rate in fact…

    \newpage

    {99} \bibitem{paper} A DISCRETE HARMONIC FUNCTION BOUNDED ON A LARGE PORTION OF Z2 IS CONSTANT

    \bibitem{stein} Functional analysis

     


    补充说明

    以下是新整理的中文说明;上方旧博客原文保持不变。

    离散调和函数是连续调和函数在格点 $\mathbb Z^n$ 上的版本。它保留了平均值性质、最大值原理和随机游走解释,但也带来新的组合几何问题:如果一个离散调和函数在很多点上都不大,是否能推出它必须是常数?

    离散调和函数:格点 Laplacian、坏点集合与 Liouville 型问题
    离散调和函数的平均值性质可以用随机游走解释,坏点集合的几何受到最大值原理控制。

    1. 定义

    函数 $u:\mathbb Z^n\to\mathbb R$ 称为离散调和,如果对每个 $x\in\mathbb Z^n$,

    $$u(x)=\frac1{2n}\sum_{y\sim x}u(y),$$

    其中 $y\sim x$ 表示 $y$ 与 $x$ 相邻。等价地,离散 Laplacian

    $$\Delta_d u(x)=\sum_{y\sim x}(u(y)-u(x))$$

    满足 $\Delta_d u=0$。

    在二维中,这就是

    $$u(i,j)=\frac14\bigl(u(i+1,j)+u(i-1,j)+u(i,j+1)+u(i,j-1)\bigr).$$

    2. 最大值原理

    离散调和函数满足最大值原理:如果 $u$ 在有限连通区域内部调和,那么最大值和最小值出现在边界上。原因很简单:一个点的值是邻点平均,若内部点达到严格最大值,则所有邻点也必须取同样的值,连通性迫使整个区域常数。

    这使得坏点集合的形状受到限制。设

    $$G=\{x:|u(x)|\le 1\},\qquad B=\mathbb Z^n\setminus G.$$

    如果某个坏点连通块完全被好点包围,那么最大值原理会迫使它不能真正坏。因此坏点必须以某种方式连接到边界或无穷远。

    3. Poisson kernel 与随机游走

    在有限区域 $\Omega\subset\mathbb Z^n$ 上,离散调和函数由边界值决定:

    $$u(x)=\sum_{z\in\partial\Omega}P_\Omega(x,z)u(z).$$

    这里 $P_\Omega(x,z)$ 是从 $x$ 出发的简单随机游走第一次离开 $\Omega$ 时落在 $z$ 的概率。这就是离散 Poisson kernel。

    这个表示把分析问题转成概率问题:若边界上大值点所占比例很小,那么内部点看到大值的概率也会受到控制。

    4. Liouville 型命题

    经典 Liouville theorem 说,有界的整调和函数必须是常数。离散版本也有类似结论。更细的问题是:如果 $u$ 不假设处处有界,但在每个大球里都有固定比例的点满足 $|u|\le 1$,是否仍能推出 $u$ 是常数?

    这类命题的证明通常要比较两个方向。第一,离散调和性和 Poisson 表示迫使内部值由边界平均控制。第二,若存在越来越大的坏点连通块,那么等周不等式或随机游走逃逸概率会给出大值传播。二者冲突时,就只能得到常数解。

    5. 为什么维数和尺度重要

    在 $\mathbb Z^n$ 中,每个半径球只有有限个格点,边界体积和内部体积之间有明确关系。坏点集合若想在所有尺度上保持稀疏,就很难同时支撑一个非平凡调和函数的增长。

    这和连续情形的差别在于,离散空间把局部传播路径变成了组合对象。一个值要从边界影响内部,必须沿随机游走路径进入;而路径数量、逃逸概率和等周结构都会参与估计。

    6. 证明图像

    可以把证明想成两条 lemma 的冲突。一个 lemma 来自 Poisson kernel:内部值是边界值的随机平均,因此不能随意增长。另一个 lemma 来自坏点集合的几何:如果坏点在每个尺度都存在足够结构,它的质量会以某种速率传播。若假设“好点比例”在所有尺度上都足够大,这两种趋势最终矛盾。

    这个问题的有趣之处在于,它把调和分析、随机游走和离散等周不等式绑在一起。离散调和函数不是连续理论的机械翻译,而是一种真正带有格点几何味道的分析对象。

  • 不确定性原理:局部化、Bernstein 估计与 Gaussian

    旧博客原文

    原题:Uncertainty principle

    The pdf version is Uncertainty principle. The nice note of terrence tao seems given a nice answer for the problem below.

    1. Introduction

    Is there a Brunn-Minkowski inequality approach to the phenomenon charged by uncertainty principle? More precisely, is it possible to say some thing about the Gaussian distribution

    \displaystyle G(x)=e^{-|x|^2} \ \ \ \ \ (1)

     

    to be the best choice that {\|\hat G-G\|_2} arrive minimum?

    Remark 1 Or some other suitable distance space on reasonable function (may be some gromov hausdorff distance? Any way, to say the guassian distribution is the best function to defect the influence of uncertain principle.

    I do not know the answer of the problem 1, but this is a phenomenon of a universal phylosphy, aid, uncertainty principle, heuristic:

    It is not possible for both function {f} and its Foriour transform {\hat f} to be localized on small set.

    Now let me give some approach by intuition to explain why the phenomenon of “uncertainty principle” could happen.

    The approach is based on:

    1. level set decomposition.
    2. area formula (or coarea formula), anyway, some kind of change variable formula.
    3. integral by part.
    4. Basic understanding on exponential sum.

    Let our function {f\in S} the Shwarz space, we begin with a intuition (not very rigorous) calculate:

    \displaystyle \begin{array}{rcl} \hat f(\xi) & = & \int e^{2\pi i<\xi, x>}f(x)dx\\ & \overset{integral \ by \ part}= &\int \frac{1}{-2\pi i\xi}e^{-2\pi i<x,\xi>}\cdot \nabla f(x)\\ & \overset{Fubini}= & \int_{inf |f|}^{max |f|}\int_{level set A(t)} \frac{-e^{2\pi i<\xi,x>}}{-2\pi i\xi}\nabla f(x)dH^{n-1}(A)dt \end{array}

    Now we try to understanding the result of the calculate, it is,

    \displaystyle \hat f(\xi) =\int_{inf |f|}^{max |f|}\int_{level set A(t)} \frac{-e^{2\pi i<\xi,x>}}{-2\pi i\xi}\nabla f(x)dH^{n-1}(A)dt \ \ \ \ \ (2)

     

    \displaystyle \pounds(A(t),\xi)=\int_{level set A(t)} \frac{-e^{2\pi i<\xi,x>}}{-2\pi i\xi}\nabla f(x)dH^{n-1}(A) \ \ \ \ \ (3)

     

    The calculate is wrong, but not very far from the thing that is true, the key point is now the exponential sum involve. We could use the pole coordinate in the frequence space and get some very rough intuition of why the the uncertainty principle could occur.

    Remark 2 Why we consider the level set decomposition, due to the integral is a combination of linear sum of the integral on every level set, so shape of level set is the key point.

    The part of {\frac{-e^{2\pi i<\xi,x>}}{-2\pi i\xi}} in 2 is a rotation on the level set, a wave correlation of it and the christization function {\chi_{A_t}} of level set {A_t} in the whole space, this is of course a exponential sum.

    Now we can begin the final intuition explain of the phenomenon of uncertainty principle. If the density of function {f} is very focus on some small part of the physics space, then it is the case for level sets of {f}, but we could say some thing for the exponential sum {\pounds(A(t),\xi)} 3 related to the level set, just by very simply argument with hardy litterwood circle method or Persaval identity? Any way, something similar to this argument will make sense, due to if the diameter of level set focus ois small, then we can not get a decay estimate for {\pounds(A(t),\xi)} when {\xi\rightarrow \infty} along one direction in frequency space, in fact we could say the inverse, i.e. it could not decay very fast.

    2. Bernstein’s bound and Heisenberg uncertainty principle

     

     

    2.1. Motivation and Bernstein’s bound

    There is two different Bernstein’s bound, we discuss the first with the motivation, and proof the second rigorously. \paragraph{Form 1} {A} is a invertible affine map, then for a ball {B}, {A(B)=\epsilon} is a ellipsoid.

    \displaystyle \epsilon=\{x\in {\mathbb R}^d|\sum_{j=1}^{d}r_j^{-2}(x_j-y_j)^2\leq 1\} \ \ \ \ \ (4)

     

    By a orthogonal transform we could make {A} to be a diagonal matrix, i.e. {A=diag(r_1,...,r_d)}. It is said, for {\forall f\in S} or {f} is a smooth bump function, {f_A=f\circ A^{-1}}, so we have,

    \displaystyle \hat f_A(\xi)=\int e^{2\pi i<x,\xi>}\cdot f\circ A^{-1}(x)dx \ \ \ \ \ (5)

    We define dual of {\epsilon}, {\epsilon^*:=\{\xi\in {\mathbb R}^d| \sum_{j=1}^d\xi_j^2r_j^2\leq 1\}}.

    Remark 3 Why there we use the metric {\xi_j^2r_j^2\leq 1} but not the standard inner product {<\xi,x>}? How to understand the choice?

    Proposition 1 We have the following property:

    1. {f_A\in L^{\infty}\Longrightarrow \|\hat f_A\|_{1}\leq +\infty}.
    2. {|\hat f_A(\xi)|\leq c_N|\xi|(1+|\xi|^2_{\epsilon^*})^{-N}}

     

    Remark 4

    \displaystyle |\xi|^2_{\epsilon^*}=\sum_{i=1}^d\xi_j^2r_j^2

    This is a norm of {{\mathbb R}^d} related to {\epsilon^*}.

    Proof: Suffice to proof 2.

    \displaystyle \begin{array}{rcl} |\hat f_A(\xi)| & = & |\int_{{\mathbb R}^d}e^{2\pi i<x,\xi>}f_A(x)dx|\\ & \overset{integral \ by \ part}\sim & \frac{1}{(1+|\xi|)^N}\int |e^{2\pi i<x,\xi>}\partial^N f_A(x)dx|\\ & = & c_N |\xi|(1+|\xi|_{\epsilon^*}^2)^{-N} \end{array}

    \Box

    More quantitative we have rigorous one: \paragraph{Form 2} If {f\in L^2({\mathbb R}^d)}, {supp f\in B(r,0)}, then it is not possible for {\hat f} to be concentrate on a scale much less than {R^{-1}}.

    Proposition 2 (Bernstein’s bound) Suppose {f\in L^2({\mathbb R}^d)}, {supp f\subset B_R(0)}. Then,

    \displaystyle \|\partial^{\alpha}\hat f\|_2\leq (2\pi r)^{|\alpha|}\|f\|_2, \forall \alpha. \ \ \ \ \ (6)

    Proof: {\alpha=0} case is trivial by Paserval identity, which said on {L^2({\mathbb R}^d)}, fourier transform is a isometry, {\|f\|_2=\|\hat f\|_2}. For general case, integral by part, and use trivial estimate,

    \displaystyle \begin{array}{rcl} \|\partial^{\alpha}\hat f\|_2 & \overset{integral \ by \ part}= & \|x^{\alpha} f\|_2\\ & \leq & (2\pi r)^{\alpha}\|f\|_2 \end{array}

    \Box

    2.2. Heisenberg inequality

    Theorem 3 (Heisenberg uncertain principle) {f\in L^2({\mathbb R}^d)}, so {\hat f\in L^2({\mathbb R}^d)}, {\|f\|_2=\|\hat f\|_2}. then for any {x_0,\xi_0\in {\mathbb R}^d}, every direction, we have

    \displaystyle \|f\|_2^2=\|f\|_2\|\hat f\|_2\leq \|(x-x_0)f\|_2\|(\xi-\xi_0)\hat f\|_2 \ \ \ \ \ (7)

     

    Remark 5 We could understand the inequality by the following way. suffice to prove it with {f\in S} and then by approximation argument. {f\otimes \hat f\in S({\mathbb R}^d\times {\mathbb R}^d)}, define {\|f\otimes \hat f\|_2:= \|f\|_{L^2({\mathbb R}^d)}\cdot \|\hat f\|_{L^2({\mathbb R}^d)}}. then we have the following:

    \displaystyle \|f\otimes \hat f\|_{L^2}\leq 4\pi \|xf\otimes \hat{xf}\|_{L^2} \ \ \ \ \ (8)

    Remark 6 The inequality is shape, the extremizers being precisely given by the modulated Gaussians: arbitrary

    \displaystyle f(x)= c e^{2\pi i\xi_0x}e^{-\pi \delta(x-x_0)^2} \ \ \ \ \ (9)

     

    There are two proof strategies I have tried, I try them for several hour but not work out with a satisfied answer, the method more involve, I explain what happen in section 1, I have not tried, I will try it later. Both this two strategies i face some difficulties, I explain why I can not work out them with a proof: \paragraph{Strategy 1} The first one is, we could work with {f\in S} of course, by approximation, then we find, by Paserval, {\|f\|_2=\|\hat f\|_2, \|\partial_x f\|_2=\|\xi \hat f\|2} and are both true. then we use our favourite way to use Cauchy-Schwarz, the difficulty is we can not use a integral by part argument directly, even after restrict ourselves with monotonic radical symmetry inequality and by a rearrangement inequality argument, it seems reasonable due to rearrangement decreasing the kinetic energy as said in Lieb’s book. But even work with monotonic one, then one involve with some complicated form, try to use Fubini theorem to rechange the order of integral try to say something, it is possible to work out by this way but I do not know how to do. There is some calculate under this way,

    \displaystyle \begin{array}{rcl} \|xf\|_2\|\xi \hat f\|_2 & \overset{Cauchy-Schwarz}\geq & \int xf\cdot \partial_x f\\ & \sim & \int f^2 \end{array}

    but you know, at a point we have {\partial_x(Xf)=f+x\partial_x f\neq f}, the reasonable calculate is following,

    \displaystyle \partial_x(xf)=f+x\partial_x f \ \ \ \ \ (10)

    We want {\partial_x P(x,f) =f}, Then

    \displaystyle \begin{array}{rcl} \partial_x P(x,f) & = & f\\ & = & \partial_x(xf)-x\partial_x f\\ & = & \partial_x(xf)-\partial_x(\frac{1}{2}x^2\partial_x f)+\frac{1}{2}x^2\partial_{x^2}f\\ & = & \partial_x(\frac{1}{6}x^3\partial_{x^2}f)-\frac{1}{6}x^3\partial_{x^3}f\\ ...\\ & = &\partial_x(\sum_{i=1}^{\infty}(-1)^{i+1}x^i\frac{1}{i!}\partial_{x^i}f)+(-1)^{i+1}x^i\frac{1}{i!}\partial_{x^{i+1}}f \end{array}

    Seems to be {f=\partial_x(ln(f))}… I do not know.

    \paragraph{Strategy 2} The second strategy is, in the quantity {\|xf\|_2\|\xi \hat f\|_x} we lose two cone very near {x_0,\xi}, we need use the extra thing to make up them. May be effective argument come from some geometric inequality.

    3. The Amerein-Berthier theorem

    Next we investigate following problem, the problem is following: if {E,F\subset {\mathbb R}^d} are of finite measure, can there be a nonzero {f\in L^2({\mathbb R}^d)} with {supp (f)\subset E} and {supp(\hat f)\subset F}? Some argument is folowing: Observe that:

    \displaystyle \chi_{F}\hat f=\hat f \Longrightarrow \chi_{E}(\chi_F \hat f)^{\vee}=f. \ \ \ \ \ (11)

    Assume that: {Tf:=\chi_{E}(\chi_F \hat f)^{\vee}} then {Tf=f}. So we have, at least {\|T\|_{2-2}\geq 1}. Some dirty calculate show:

    \displaystyle \begin{array}{rcl} (Tf)(x) & = & \int e^{2\pi i\xi x}\chi_F\hat f(\xi)\chi_E(x)d\xi\\ & = & \int \int e^{2\pi i\xi(x-y)}f(y)\chi_F(\xi)\chi_E(x)dy d\xi\\ & \overset{Fubini}= &\int_{{\mathbb R}^d}\chi_E(x)\chi_F^{\vee}(x-y)f(y)fy \end{array}

    So we can define kernel of {T},

    \displaystyle K(x,y)=\chi_E(x)\chi_F(x-y)^{\vee} \ \ \ \ \ (12)

     

    By Fubini, we calculate the Hilbert-Schmidt norm:

    \displaystyle \int_{{\mathbb R}^{2d}}|K(x,y)|^2dxdy=|E||F|=\sigma^2<+\infty \ \ \ \ \ (13)

     

    So {T} is a compact operator and its {L^2} operator norm satisfied {\|T\|=\min(\sigma,1)}. So if {\sigma<1} then we can canculate we can not have {f\neq 0} in the original question.

    The story is in fact more interesting, the answer of the question is no even for {\sigma\geq 1}, so in all case. We have the following quatitative theorem:

    Theorem 4 {E,F} finite measure in {{\mathbb R}^d}, then

    \displaystyle \|f\|_{L^2({\mathbb R}^d)}\leq c(\|f\|_{L^2(E^c)}+\|\hat f\|_{L^2(F^c)}) \ \ \ \ \ (14)

    for some constant {c=c(E,F,d)}.

    Remark 7 There is a naive approach for this theorem: Area formula trick, the shape of level set. Obvioudly we have:

    \displaystyle \|f\|_{L^2({\mathbb R}^d)}\leq \|f\|_{L^2(E)}+\|f\|_{L^2(E^c)} \ \ \ \ \ (15)

     

    Key point is proof:

    \displaystyle \|f\|_{L^2(E)}\leq c(E,F,d)\|\hat f\|_{L^2(F^c)} \ \ \ \ \ (16)

     

    Let us do some useless further calculate:

    \displaystyle \begin{array}{rcl} |f\|_{L^2(E)} & = & \|\chi_E \cdot f\|_{L^2({\mathbb R}^d)}\\ & = & \|\widehat {\chi_E\cdot f}\|_{L^2({\mathbb R}^d)}\\ & = & \|\hat \chi_E \cdot \hat{f^{\vee}}\|_{L^2({\mathbb R}^d)} \\ & = & \|\chi_E^{\vee} * f^{\vee}\|_{L^2({\mathbb R}^d)} \end{array}

    So suffice to have:

    \displaystyle \|\chi_E^{\vee} * f^{\vee}\|_{L^2({\mathbb R}^d)}\leq c(E,F,d)\|\hat f\|_{L^2(F^c)} \ \ \ \ \ (17)

    But there is connter example given by modified scaling Gaussian distribution… The point is form 15 to 16 is too loose.

    Following I given a right approach, following by my sprite on level set and area formula argument and discritization.

    Proof: The story is the same for a discretization one. We need point out, change the space {{\mathbb R}^d} to {{\mathbb Z}^d}, then every thing become a discretization one, and the change could been argue as a approximation way. What happen then, we have a naive picture in mind which is:

    \displaystyle \delta \rightarrow wave , \ wave \rightarrow \delta

    What is the case with {L^2} norm, it become the standard nner product on {{\mathbb Z}^d}, and the scale involve, i.e. we have the following basic estimate:

    \displaystyle \|\chi_E f\|_1^2\leq \|\chi_E f\|_2 \cdot|E| \ \ \ \ \ (18)

     

    Now image if the density of {f} concentrate in a very small area, then by a cut off argument we consider the supp of {f}, {supp f=E} is very small, then use the argument 18, we could conclute the density of {\hat f} could not very concentrate in the fraquence space. The constant {c(E,F,d)} could be given presicely by this way, but I do not care about it. \Box

     

    4. Logvinenko-Sereda theorem

    Next we formulate some result that provide further evidence of the non-concentration property of functions with Fourier support on {B_1}.

    4.1. A toy model

    Theorem 5 Let {\alpha>1} an suppose that {S\subset {\mathbb R}^d} satisfies,

    \displaystyle |S\cap B|<\alpha |B|, \ for \ all \ balls \ B \ of \ radius\ 1. \ \ \ \ \ (19)

    If {f\in L^2({\mathbb R}^d)} satisfies {supp(\hat f)\subset B(0,1)} then

    \displaystyle \|f\|_{L^2(S)}\leq \delta(\alpha)\|f\|_2 \ \ \ \ \ (20)

    Where {\delta(\alpha)\rightarrow 0} as {\alpha \rightarrow 0}.

    Proof: This is a easy corollary of the argument I give in the proof of Amerein-Berthier theorem 4. \Box

    4.2. A refine version

    Theorem 6 Suppose that a measurable set {E\subset {\mathbb R}^d} satisfies the following “thinkness” condition: there exists {\gamma\in (0,1)} such that

    \displaystyle |E\cap B|>\gamma |B| \ for \ all \ balls \ B \ of \ radius\ R^{-1}. \ \ \ \ \ (21)

    where {R>0} is arbitrary but fixed. Assume that {supp(\hat f)\subset B(0,R)}. Then

    \displaystyle \|f\|_{L^2({\mathbb R}^d)}\leq C\|f\|_{L^2(E)}. \ \ \ \ \ (22)

    where the constant {C} depends only on {d} and {\gamma}.

    Remark 8 This proof need some very good estimate come from several complex variables.

    5. The Malgrange-Ehrenpreis theorem

    Theorem 7 Let {\Omega} be a bounded domain in {{\mathbb R}^d} and let {p\neq 0} be a polynomial, Then, for all {g\in L^3(\Omega)}, there exists {f\in L^2(\Omega)} such that {p(D)f=g} in a distribution sence.

     


    补充说明

    以下是新整理的中文说明;上方旧博客原文保持不变。

    不确定性原理的核心直觉是:一个函数和它的 Fourier transform 不可能同时高度集中。这里的“集中”可以有许多精确版本:支撑大小、方差、指数衰减、带限函数的空间尺度,或者通过 level sets 来衡量的几何集中。

    不确定性原理:局部化、Bernstein 估计与 Gaussian
    不确定性原理的基本图像:物理空间中的强集中会迫使频率空间展开,反之亦然。

    1. 基本直觉

    Fourier transform 把函数分解成平面波:

    $$\widehat f(\xi)=\int_{\mathbb R^n}f(x)e^{-2\pi i x\cdot \xi}\,dx.$$

    如果 $f$ 集中在很小的物理区域里,那么积分里的相位 $x\cdot \xi$ 在许多方向上变化不够快,$\widehat f$ 就不可能在频率空间里迅速衰减。反过来,如果 $\widehat f$ 被限制在很小的频率球内,$f$ 就必须在物理空间里足够平滑、足够展开。

    从 level set 的角度看,$f$ 的每一层都贡献一个振荡积分。level set 的几何形状越小、越刚性,相位平均带来的抵消就越弱。

    2. Bernstein 估计

    一个非常实用的形式是 Bernstein inequality。若 $\widehat f$ 支撑在半径 $R$ 的球内,那么

    $$\|\nabla f\|_{L^p}\lesssim R\|f\|_{L^p}.$$

    这说明带限函数不能在尺度远小于 $R^{-1}$ 的区域内剧烈变化。频率支撑越小,物理空间越平滑;频率支撑越大,函数才有能力产生更细尺度的振荡。

    更一般地,如果频率支撑在一个椭球中,那么物理空间的自然尺度由对偶椭球控制。这个“对偶尺度”正是不确定性原理的几何版本。

    3. Heisenberg 不等式

    最经典的方差形式是

    $$\left(\int_{\mathbb R^n}|x|^2|f(x)|^2\,dx\right)
    \left(\int_{\mathbb R^n}|\xi|^2|\widehat f(\xi)|^2\,d\xi\right)
    \ge C_n\|f\|_2^4.$$

    一维情形可以通过分部积分和 Cauchy-Schwarz 证明。核心计算是把 $\|f\|_2^2$ 写成 $-\int x\partial_x(|f|^2)$ 的形式,再把 $x f$ 和 $\partial_x f$ 配对。Plancherel 把 $\partial_x f$ 转成频率侧的 $\xi\widehat f$。

    4. Gaussian 为什么是极值

    Gaussian 同时在物理空间和频率空间保持 Gaussian 形状:

    $$f(x)=e^{-\pi |x|^2},\qquad \widehat f(\xi)=e^{-\pi |\xi|^2}.$$

    它正好平衡了两个方向的集中,因此在 Heisenberg 型不等式中给出等号情形。这个现象提示一个更几何的问题:是否可以用类似 Brunn-Minkowski 或 entropy convexity 的语言解释 Gaussian 的最优性?

    从现代观点看,答案常常与热流、log-Sobolev 不等式和最优输运有关。热流让任意函数逐渐 Gaussian 化;entropy 的凸性则记录了这个过程中的不确定性增加。

    5. 一个可继续追问的问题

    如果把“集中程度”换成某种几何距离,例如 level sets 的体积增长、分布之间的 transport distance,或者函数图像诱导的度量结构,那么 Gaussian 是否仍然是最自然的最优对象?这个问题不一定有唯一答案,但它把 Fourier 分析里的不确定性和凸几何、概率、几何测度论连在了一起。

  • 伪微分算子与奇异积分:从 Fourier multiplier 到 symbol calculus

    旧博客原文

    原题:Pesudo differential opertor and singular integral

    I already understand this material 3days ago but it is a little difficult for me to type the latex…

     

    1. Introduction

    There is two space to understand a function’s behaviour, the physics space and the frequency space (Why thing going like this? Why there is such a duality?). Namely, we have:

    \displaystyle \hat f(\xi)=\int_{{\mathbb R}^d}e^{2\pi i\xi x}f(x)dx \ \ \ \ \ (1)

     

    The key point is, waves is a parameter group of scaling of definition of a constant fraquence wave, so it connected the multiplication and addition. Basically due to it can be look as the correlation of a function and the scaling of wave with carry all the information about {f}. A generation of this obeservation is the wavelet theory.

    So as we well know, the key ingredient of Fourier transform is to image function as a sum of series waves. A famous theorem of Mikhlion said that a translation-invariant operator {T} on {R^n} could be represented by a multiplication operator on the Fourier transform side. translation is the meaning, {h\circ T=T\circ h, \forall h} is a translation.

    In a formal level, consider it as distribution (compact distribution or temperature distribution is both OK). We have:

    \displaystyle T(e^{2\pi ix\xi})=a(\xi)e^{2\pi ix\xi}, \forall \xi \in {\mathbb R}^n \ \ \ \ \ (2)

     

    the meaning is if we consider {T} is a operator on distribution space, {T:S'\rightarrow S'}, then {\forall f\in S},

    \displaystyle \int T(e^{2\pi ix\xi})f=\int a(\xi)e^{2\pi ix\xi}f

    due to the linear combination of {e^{2\pi ix\xi}} will consititue a dense set in {S}. So this could extend to the whole distribution space by dual and give the definition of {T}, i.e.

    \displaystyle (Tf)(x)=\int_{{\mathbb R}^n}a(x,\xi)e^{2\pi ix\xi}\hat f(\xi)d\xi \ \ \ \ \ (3)

     

    Remark 1 {T} is bounded on {L^2({\mathbb R}^n)} when {a} is a bounded function, thanks to Parevel theorem. When {a} is a bounded function, the composition of two such operator could be defined, and the symbol of composition operator corresponding to the composite of their symbol, i.e.

    \displaystyle T_a\circ T_b(e^{2\pi ix\xi})=b(\xi)a(\xi)e^{2\pi ix\xi} \ \ \ \ \ (4)

     

    Remark 2 For parenval theorem, i.e. {\|\hat f\|_2=\|f\|_2}, there is two approach, heat kernel approximation approach and discretization.

    We wish to investigate the operator given by multiplier, i.e.

    \displaystyle (Tf)(x)=\int_{{\mathbb R}^n}a(x,\xi)e^{2\pi ix\xi}\hat f(\xi)d\xi \ \ \ \ \ (5)

    When it is satisfied {\|T\|_{p-p}<\infty}?

    Intuition, the following calculate is only morally true, not rigorous.

    \displaystyle \begin{array}{rcl} \|Tf\|^p_p & = & \int_{{\mathbb R}^d }|\int_{{\mathbb R}^d}a(x,\xi) e^{2\pi ix\xi} \hat f(\xi)d\xi|^pdx \\ & \overset{\exists f}\sim & \int_{{\mathbb R}^d }\int_{{\mathbb R}^d}|a(x,\xi) e^{2\pi ix\xi} \hat f(\xi)|^p d\xi dx \\ & \overset{Fubini}\sim & \int_{{\mathbb R}^d}\int_{{\mathbb R}^d}|\widehat{a(x,\xi)}\hat f(\xi) |^pdxd\xi \\ & \sim & \int_{{\mathbb R}^n}\int_{{\mathbb R}^n}|{a(x,\xi)}^{\vee}*f(\xi)|^pdxd\xi \end{array}

     

    So we need some restriction on { {a(x,\xi)}^{\vee}}, namely {\widehat{a(x,\xi)}}, so we need some decay condition on {|\partial_x^{\alpha}\partial_{\xi}^{\beta}a(x,\xi)|}, why this, just consider integral by part for {a(x,\xi)\in S}. The rigorozaton of this intuition inspirit us to the definition of symbol calss.

    Definition 1 we say {a(x,\xi)} is in symbol class {S_m} iff,

    \displaystyle |\partial_x^{\beta}\partial_{\xi}^{\alpha}a(x,\xi)|\leq A_{\alpha,\beta}(1+|\xi|)^{m-|\alpha|} \ \ \ \ \ (6)

    for all {\alpha,\beta} is multi-indece.

    Remark 3

    1. we note that all partial differential operator, whose coefficient, together with all their derivatives are bounded belong to this class, In this particular circumstance, the symbol is a polynomial in {\xi}, essentially the “characteristic polynomial” of the operator.
    2. The general operator of this class have a parallel description in terms of their kernels. That is, in a suitable sense,

      \displaystyle (Tf)(x)=\int_{{\mathbb R}_x}K(x,y)f(y)dy \ \ \ \ \ (7)

      besides enjoying a cancellation property, {K} is here characterized by differential inequalities “dual” to those for {a(x,\xi)}. In the key case where the order {m=0}, this kernel representation makes {T} a singular integral operator.

    3. The crucial {L^2} estimate, when {m=0}, is atelatively simple consequences of Plancherel’s theorem for the Fourier transform. With this, the {L^p} theory introduce in previous note is therefore applicable.
    4. The product identity that holds in the translation-invariant case generalized to the situation treated here as a symbolic calculus for the composition of operators. That is, there is an asymptotic formula for the composition of two such operators, whose main term is the point-wise product of their symbols.
    5. The succeeding terms of the formula are of decreasing orders. These orders measure not only the size of the symbols, but determine also the increasing smoothing properties of the corresponding operators. The smoothing properties are most neatly expressed in terms of the Sobolev space {W_k^p} and the Lipschitz space {\Lambda_{\alpha}}.

     

    2. Pseudo-differential operator

    “Freezing principle”: from variable coefficient differential equation to constant coefficient differential equation by approximation. divide into 2 steps:

    1. divide space into small cubes.
    2. take average of the coefficient of differential equation in every cubes.

    Suppose we are interested in study the solution of the classical elliptic second order equation.

    \displaystyle (Lu)(x)=\sum a_{ij}(x)\frac{\partial^2 u(x)}{\partial x_i\partial x_j}=f(x) \ \ \ \ \ (8)

     

    Where the coefficient matrix {\{a_{ij}(x)\}} is assume to be real, symmetric, positive definite and smooth in {X}. Understanding {P}, such that,

    \displaystyle PL=I \ \ \ \ \ (9)

     

    Looking for a {P}. Such that {PL=I+E}. {E} is a error term which have good control. To do this, fix an arbituary point {x_0}, freeze the operator {L} at {x_0}:

    \displaystyle L_{x_0}=\sum a_{ij}(x_0)\frac{\partial^2}{\partial x_i\partial x_j} \ \ \ \ \ (10)

     

    In Fourier sense ({L^2} sence).

    \displaystyle \begin{array}{rcl} L_{x_0}f(x) & = & \int e^{2\pi ix\xi}(\widehat{ \sum a_{ij}(x_0)\frac{\partial^2 f(\xi)}{\partial x_i\partial x_j}}) d\xi\\ & = & \int e^{2\pi ix\xi}\int e^{-2\pi i \xi y}\sum a_{ij}(x_0)\frac{\partial^2}{\partial x_i\partial x_j}f(y)dyd\xi\\ & = & \int e^{2\pi ix\xi}(-4 \pi^2)\sum_{i,j}a_{ij}(x_0)\xi_i\xi_j \end{array}

    Remark 4 The remark is, morally speaking, for application of fourier transform in PDE. morally we could only solve the problem with linear differential equation (although we could consider the hyperbolic type). The main obstacle for Fourier transform application into PDE:

    1. it only make sense with Schwarz class or its dual, this is not main obstacle, in principle could be solved by rescaling.
    2. the main obstacle is it only compatible with linear differential equation.

     

    Cut-off function: {\eta} vanish near the origin,

    \displaystyle (P_{x_0}f)=\int_{{\mathbb R}^n}(-4\pi^2\sum_{i,j}a_{ij}(x_0)\xi_i\xi_j)^{-1}\eta(\xi)\hat {f(\xi)}e^{2\pi ix\xi}d\xi. \ \ \ \ \ (11)

     

    then:

    \displaystyle L_{x_0}P_{x_0}=I+E_{x_0}.

    {E_{x_0}} is actually a smoothing operator, because it is given by convolution with a fixed test function. It should be seasonable when {x} near {x_0}, {(Pf)(x)} is well approximated by {(P_{x_0}f)(x)}, it is actually the case, define {((Pf)(x):=(P_xf)(x)}, i.e.

    \displaystyle (Pf)(x)=\int_{{\mathbb R}^n}(-4\pi^2\sum_{i,j}a_{ij}(x)\xi_i\xi_j)^{-1}\eta(\xi)\hat f(\xi)e^{2\pi ix\xi}d\xi \ \ \ \ \ (12)

     

    The operator {P} so given is a propotype of a pesudo-differential operator. Moreover, one has {LP=I+E}, where the error operator {E} is “smoothing of order 1”. That this is indeed the case is the main part of the symbolic calculus described.

    Definition 2 (symbol class) A function {a(x,\xi)} belong to {S^m} and is said to be of order {m} of {a(x,\xi)} is a {C^{\infty}} function of {(x,\xi)\in {\mathbb R}^n\times {\mathbb R}^n} and satisfies the differential inequality:

    \displaystyle |\partial_x^{\beta}\partial_{\xi}^{\alpha}a(x,\xi)|\leq A_{\alpha,\beta}(1+|\xi|)^{m-|\alpha|} \ \ \ \ \ (13)

     

    For all {\alpha,\beta} are multi-indece.

    Now we trun to the exact meaning of pesudo-differential operator, i.e. how them action on functions. Under some suffice given regularity condition, for {a\in S^m}, {T_a:S\rightarrow S}.

    \displaystyle (Tf)(x)=\int_{{\mathbb R}^n}a(x,\xi)\hat f(\xi)e^{2\pi ix\xi}d\xi \ \ \ \ \ (14)

     

    Remark 5 {T_a:S\rightarrow S} is continuous and for {a_k\rightarrow a} pointwise, {a_k\in S,\forall k\in {\mathbb N}^*}, {T_{a_k}(f)\rightarrow T_a(f)} in {S}.

    then expense it, we get:

    \displaystyle (T_af)(x)=\int\int a(x,\xi)e^{2\pi i\xi(x-y)}f(y)dyd\xi \ \ \ \ \ (15)

     

    This could be diverge, even when {f\in S}. The key point is we do not have control with the second integral, morally speaking, this phenomenon is the weakness of Lesbegue integral which would not happen in Riemann integral, so sometime we need the idea from Riemann integral, this phnomenon is settle by multi a cut off function {\eta_{\epsilon}} and take {\epsilon\rightarrow \infty}, the same deal also occur as the introduced of P.V. integral in Hilbert transform. The precise method to deal with the obstacle is following: {a_{\epsilon}(x,\xi)=a(x,\xi)\gamma(\epsilon x,\epsilon \xi)}, if {a\in S^m}, {a_{\epsilon}\in S^m}. {T_{a_{\epsilon}}\rightarrow T_a} in the sense:

    {\forall f\in S}, {T_{a_{\epsilon}}(f)\rightarrow T_a(f)},

    \displaystyle (T_af)(x)=\lim_{\epsilon\rightarrow 0}\int\int a_{\epsilon}(x,\xi)e^{2\pi i\xi(x-y)}f(y)dyd\xi \ \ \ \ \ (16)

    We also have:

    \displaystyle <T_af,g>=<f,T_a^*g>, \forall f,g\in S. \ \ \ \ \ (17)

     

    Then we have:

    \displaystyle (T^*_ag)(y)=\lim_{\epsilon\rightarrow 0}\int\int \bar a_{\epsilon}(x,\xi)e^{2\pi i\xi(y-x)}g(x)dxd\xi \ \ \ \ \ (18)

    and {<f,g>} denotes {\int_{{\mathbb R}^n}f(x)\bar g(x)dx}. Thus the pesudo-differential operator {T_a} initially defined as a mapping from {S} to {S}, extend via the identity 17 to a mapping from the space of temperatured distribution {S'} to itself {S'}. Notice also that {T_a} is automatically continuous in this space. \newpage

    3. {L^p} bounded theorem

    We first introduce a powerful tools, called dyadic decomposition,

    Lemma 3 (dyadic decomposition) In eculid space {{\mathbb R}^n} there exists a function {\phi\in C^{\infty}({\mathbb R}^n)} such that,

    \displaystyle \sum_{i\in {\mathbb Z}}\phi(2^{-i}x)=1 \ \ \ \ \ (19)

     

    and {\forall x\in {\mathbb R}^n}, there is only two of {i\in {\mathbb Z}} such that {\phi(2^{-i}x)\neq 0}, and we can choose {\phi} to be radical and {\phi(x) \geq 0,\forall x\in {\mathbb R}^n}.

    Remark 6

    So for a given mutiplier {a(x,\xi)}, we will have {a(x,\xi)=\sum_{i\in {\mathbb Z}}a_i(x,\xi)=\sum_{i\in Z}\phi(2^{-ix})a(x)}.

    Proof: The proof is easy, after rescaling we just need observed there is a bump function satisfied whole condition. \Box

    Theorem 4 Suppose {a} is a symbol of order 0, i.e. that {a\in S^0} Then the operator {T_a}, initially defined on {S}, extends to a bounded operator from {L^2({\mathbb R}^n)} to itself.

    Remark 7 Suffice to show {\|T_a(f)\|_{L^2}\leq A\|f\|_{L^2}, \forall f\in S} and by dual.

    In fact we can directly proof a more general theorem:

    Theorem 5 Let {m:{\mathbb R}^d-\{0\}\rightarrow {\mathbb C}} satisfy, for any multi-index {\gamma} of length {|\gamma|\leq d+2},

    \displaystyle |\partial^{\gamma}m(\xi)|\leq B|\xi|^{-|\gamma|}

    For all {\xi\neq 0}. Then, for any {0<p<\infty}, there is a constant {C=C(d,p)} such that,

    \displaystyle \| (m\hat f)^{\vee}\|_p\leq C(p,d)\|f\|_p \ \ \ \ \ (20)

     

    for all {f\in S}.

    Proof: {a\in S^0}, so we have:

    \displaystyle |\partial_x^{\beta}\partial_{\xi}^{\alpha}a(x,\xi)\leq a_{\alpha,\beta}(1+|\beta|)^{-|\alpha|} \ \ \ \ \ (21)

    {\forall \alpha,\beta} are multi indeces. Then we consider dyadic decomposition, the is a function {\phi} satisfied the condition in 19, define {a_i(x,\xi)=\phi(2^{-i}x)a(x)}. then {supp a_i(x,\xi)} cpt, {\|a_i\|<\infty}. So {a_i\in L^p({\mathbb R}^n)}, we have,

    \displaystyle \begin{array}{rcl} \|T_{a_i}f\|_p^p & = & \int_{{\mathbb R}^d}|K_i*f(x)|^pdx \\ & = & \int_{{\mathbb R}^d}|\int_{{\mathbb R}^d}K_i(x-y)f(y)dy|^pdydx\\ & \leq &\int_{{\mathbb R}^d}\int_{{\mathbb R}^d}|K_i(x-y)f(y)|^pdydx\\ & = & \|K_i\|_p^p\|f\|_p^p \end{array}

    {\|K_i\|_p} have good decay estimate, thanks to {u_i=\phi(2^{-i}x)a(x)\in S^0}, this estimate is deduce morally along the same ingredient of “station phase”, it is come from a argument combine “counting point” argument and a rescaling argument. So,

    \displaystyle \begin{array}{rcl} \|Tf\|_p & = & \|\sum T_if\|_p\\ & \leq & (\sum \|K_i\|_p^p)\|f\|_p \end{array}

    But we have {\sum\|K_i\|_p^p\leq \infty}, ending the proof. \Box

    Remark 8 this method also make sense of restrict the condition to be:

    \displaystyle |\partial_x^{\beta}\partial_{\xi}^{\alpha}a(x,\xi)|\leq A_{\alpha,\beta}(1+|\xi|)^{-|\alpha|}, \forall |\alpha|\leq d+2. \ \ \ \ \ (22)

    Where {d} is the dimension of the space, and we could change {2} to {1+\epsilon}.

    Remark 9

    \displaystyle \widehat{\frac{\partial^2 u}{\partial x_i\partial x_j}}|\xi|=\frac{\xi_i\xi_j}{|\xi|^2}\widehat \Delta u(\xi), m(\xi)=\frac{\xi_i\xi_j}{|\xi|^2} \ \ \ \ \ (23)

    is a counter example for {p=1,\infty}.

    Remark 10 The key point is the estimate

    \displaystyle \int|\int e^{2\pi i\xi x}\phi(2^{-j}\xi)m(\xi)d\xi |^pdx \ \ \ \ \ (24)

    Correlation of taylor expension and wavelet expension. This is also crutial for the theory of station phase.

    4. Calculus of symbols

    This calculus of symbols would imply there is some structure on this set.

    Theorem 6 Suppose {a,b} are symbols belonging to {S^{m_1}} and {S^{m_2}} respectively. Then there is a symbol {c} in {S^{m_1+m_2}} so that:

    \displaystyle Tc=T_a\circ T_b

    Moreover,

    \displaystyle c\sim \sum_{\alpha}\frac{(2\pi i)^{-|\alpha|}}{\alpha}(\partial_{\xi}^{\alpha}a)(\partial_x^{\alpha}b). \ \ \ \ \ (25)

    in the sense that,

    \displaystyle c-\sum_{|\alpha|<N}\frac{(2\pi i)^{-|\alpha|}}{\alpha !}\partial_{\xi}^{\alpha}\partial_x^{\alpha}b\in S^{m_1+m_2-N} \ \ \ \ \ (26)

    For all {N>0}.

    The following “proof” is not rigorous, we just calculate it formally, we could believe it is true rigorously, by some approximation process. Proof: We assume {a,b} have compact support so that our manipulations are justified. We use the alternate formula 15 to write,

    \displaystyle (T_af)(y)=\int b(y,\xi)e^{2\pi i\xi(y-z)}f(z)dzd\xi \ \ \ \ \ (27)

    Then we apply {T_a}, again in the form 15, but here with the variable {\eta} replacing in the integration. The result is,

    \displaystyle T_a(T_bf)(x)=\int a(x,\eta)b(y,\xi)e^{2\pi i\eta(x-y)}e^{2\pi i\xi(y-z)}f(z)dzd\xi dyd\eta. \ \ \ \ \ (28)

    This calculate is easy to derive, but the following is more tricky. Now {e^{2\pi i\eta(x-y)}\cdot e^{2\pi i\xi(y-z)}=e^{2\pi i(x-y)(\eta-\xi)}\cdot e^{2\pi i(x-z)\xi}}, so

    \displaystyle T_a(T_bf)(x)=\int c(x,\xi)e^{2\pi i(x-z)\xi}f(z)dz d\xi \ \ \ \ \ (29)

    with

    \displaystyle c(x,\xi)=\int a(x,\eta)b(y,\xi)e^{2\pi i(x-y)(\eta-\xi)}dyd\eta \ \ \ \ \ (30)

     

    we can also carry out the integration in the y-variable. This leads to the corresponding Fourier transform of {b} in that variable, and allows us to rewrite 30 as,

    \displaystyle c(x,\xi)=\int a(x,\xi+\eta)\hat b(\eta,\xi)e^{2\pi i x\eta}d\eta. \ \ \ \ \ (31)

    With this form in hand, use taylor expense to the symbol {a(x,\xi+\eta)}, i.e.

    \displaystyle a(x,\xi+\eta)=\sum_{|\alpha|<N}\partial_{\xi}^{\alpha}a(x,\xi)\eta^{\alpha}+R_N(x,\xi,\eta) \ \ \ \ \ (32)

    with a suitable error term {R_N}, due to

    \displaystyle \frac{1}{\alpha !}\int \partial_{\xi}^{\alpha}a(x,\xi)\hat\eta(\eta,\xi)e^{2\pi ix\eta}d\eta=\frac{(2\pi i)^{|\alpha|}}{\alpha !}(\partial_{\xi}^{\alpha}a(x,\xi))(\partial_x^{\alpha}b(x,\xi)). \ \ \ \ \ (33)

    we only need to proof {R_N\in S^{m_1+m_2-N}} and it is definitely the case, we get the theorem. \Box

    Remark 11 We need replace {a,b} with {a_{\epsilon},b_{\epsilon}}, where

    \displaystyle a_{\epsilon}(x,\xi)=a(x,\xi)\cdot \gamma(\epsilon ,\epsilon \xi), b_{\epsilon}(x,\xi)=b(x,\xi)\cdot \gamma(\epsilon ,\epsilon \xi). \ \ \ \ \ (34)

    we note that {a_{\epsilon},b_{\epsilon}} satisfy the same differential inequalities that {a} and {b} do, uniformly in {\epsilon, 0<\epsilon\leq 1} .passage to the limit as {\epsilon\rightarrow 0} will then give us our desired result.

    5. Estimate in {L^p} , Sobolev, and Lipchitz space

    We now take up the regularity properties of our pesudo-differential operator as expressed in terms of the standard function spaces, we begin with the {L^p} boundedness of an operator of order {0}.

    5.1. {L^p} estimate

    Suppose {a} belongs to the symbol class {S^0}. Then, we can express {T=T_a} as

    \displaystyle (Tf)(x)=\int K(x,y)f(y)dy=\int K(x,x-y)f(y)dy \ \ \ \ \ (35)

    due to {a\in S^0}, we know, with some approximation argument and first do it with a cutoff symbol of {a}, i.e. {a_{\epsilon}}, that,

    \displaystyle |K(x,y)|\leq A|x-y|^{-n} \ \ \ \ \ (36)

    So that the integral coverage whenever {f\in S} and {x} is away from the support of {f}. Since we know that {T} is bounded on {L^2({\mathbb R})}, this representation extends to all {f\in L^2({\mathbb R})} for almost every {x\notin supp f}. More generally, we have,

    \displaystyle |\partial_{x}^{\alpha}\partial_{y}^{\beta}K(x,y)|\leq A_{\alpha,\beta}|x-y|^{-n-|\alpha|-|\beta|} \ \ \ \ \ (37)

    hence {K} satisfies,

    \displaystyle \int_{|x-y|\geq 2\delta}|K(x,y)-K(x,\bar y)|dx\leq A, \ if\ |y-\bar y|\leq \delta, all \ \ \delta>0. \ \ \ \ \ (38)

    Use the general singular integral theory we get the following {L^p} estimate.

    Theorem 7 Suppose {T_a} is the pseudo-differential operator corresponding to a symbol {a} in {S^0}, then {T_a} extends to a bounded operator on {L^p({\mathbb R}^n)} to itself, for {1<p<\infty}.

    5.2. Sobolev spaces

    We first recall the definition of the Sobolev spaces {W_k^p}, where {k} is a positive integer. A function {f} belongs to {W_k^p({\mathbb R}^n)} if {f\in L^p({\mathbb R}^n)} and the partial derivatives {\partial_x^{\alpha}f}, taken in the sense of distribution, belong to {L^p({\mathbb R}^n)}, whenever {0\leq |\alpha|\leq k}. The norm in {W_k^p} is given by,

    \displaystyle \|f\|_{W_k^p}=\sum_{|\alpha|\leq k}\|\partial_x^{\alpha}f\|_{L^p} \ \ \ \ \ (39)

    the following result is the directly corollary of 7.

    Theorem 8 Suppose {T_a} is a pseudo-differential operator whose symbol {a} belongs to {S^m}. If {m} is an integer and {k\geq m}, then {T_a} is a bounded mapping from {W_k^p} to {W_{k-m}^p}, whenever {1<p<\infty}.

    Remark 12 This theorem remain valid for arbitrary real {k,m}.

    5.3. Lipschitz spaces

    Theorem 9 Suppose {a} is a symbol in {S^m}. Then the operator {T_a} is a bounded mapping from {\Lambda_{\gamma}} to {\Lambda_{\gamma-m}}, whenever {\gamma>m}.

    Lemma 10 Suppose the symbol {a} belongs to {S^m}, and define {T_{a_j}=T_a\Delta_j}. Then, as operator from {\L^{\infty}({\mathbb R}^n)} to itself, the {T_{a_j}} have norms that satisfy

    \displaystyle \|T_{a_j}\|\leq A2^{jm} \ \ \ \ \ (40)

    We shall now point out a very simple but useful alternative characterization of {\Lambda_{\gamma}}. This is in terms of approximation by smooth functions; it is also closely connected with the definition of {\Lambda_{\alpha}} space as intermediate spaces, using the “real” method of interpolation.

    Corollary 11 A function {f} belongs to {\Lambda_{\gamma}} if and only if there is a decomposition,

    \displaystyle f=\sum_{j=0}^{\infty}f_j \ \ \ \ \ (41)

    with {\|\partial_x^{\alpha}f_j\|_{L^infty}\leq A2^{-j\gamma}\cdot 2^{j|\alpha|}}, for all {0\leq |\alpha|\leq l}, where {l} is the smallest integer {>\gamma}.

    When {f\in \Lambda_{/\gamma}}, the argument prove 10, with {T_a=I}, {f_j=F_j=\Delta_j(f)}, gives the required estimate for the {f_j}.

    A second consequence of 9 is the following:

    Corollary 12 The operator {(I-\Delta)^{\frac{m}{2}}} gives an isomorphism from {\Lambda_{\gamma}} to {\Lambda_{\gamma-m}}, whenever {\gamma>m}.

    This is clear because {(I-\Delta)^{\frac{m}{2}}} is continuous from {\Lambda_{\gamma}} to {\Lambda_{\gamma-m}}, and its inverse, {(I-\Delta)^{\frac{-m}{2}}}, is continuous from {\Lambda_{\gamma-m}} to {\Lambda_{\gamma}}.


    补充说明

    以下是新整理的中文说明;上方旧博客原文保持不变。

    理解伪微分算子的一个自然入口,是先理解 Fourier multiplier。平移不变算子在频率侧通常只是乘以一个函数;一旦系数依赖空间位置,乘子就必须升级为同时依赖 $x$ 和 $\xi$ 的 symbol。这篇笔记沿着这个想法,从 multiplier 走到 kernel,再走到 Calderon-Zygmund 型奇异积分。

    伪微分算子与奇异积分:从 Fourier multiplier 到 symbol calculus
    伪微分算子把位置变量 $x$ 和频率变量 $\xi$ 同时放进 symbol,kernel 的奇异性由 symbol 的阶数控制。

    1. Fourier multiplier

    若 $T$ 与平移可交换,那么在 Fourier 侧它应当满足

    $$\widehat{Tf}(\xi)=m(\xi)\widehat f(\xi).$$

    这里 $m$ 是 multiplier。Plancherel 定理直接给出一个基本结论:如果 $m\in L^\infty$,那么 $T$ 在 $L^2$ 上有界,并且

    $$\|Tf\|_2\le \|m\|_\infty\|f\|_2.$$

    这解释了为什么频率空间是处理常系数线性方程的自然语言:微分在频率侧变成乘法,算子代数变成函数代数。

    2. 从 multiplier 到 symbol

    变系数方程的困难在于,算子不再平移不变。伪微分算子的基本形式是

    $$Af(x)=\int_{\mathbb R^n}e^{ix\cdot\xi}a(x,\xi)\widehat f(\xi)\,d\xi.$$

    函数 $a(x,\xi)$ 称为 symbol。它描述算子在点 $x$ 附近、频率 $\xi$ 方向上的局部行为。常系数情形是 $a$ 不依赖 $x$ 的特例。

    标准 symbol class $S^m_{1,0}$ 要求对所有多重指标 $\alpha,\beta$,有

    $$|\partial_x^\alpha\partial_\xi^\beta a(x,\xi)|\le C_{\alpha\beta}(1+|\xi|)^{m-|\beta|}.$$

    对 $\xi$ 求导会降低阶数,这正对应物理空间 kernel 的衰减和光滑性。

    3. Freezing principle

    变系数椭圆方程可以先在一个点附近“冻结系数”。例如二阶椭圆算子

    $$Lu=-\partial_i(a^{ij}(x)\partial_j u)$$

    在点 $x_0$ 附近可以用常系数算子

    $$L_{x_0}u=-a^{ij}(x_0)\partial_i\partial_j u$$

    作为第一近似。Fourier transform 可以求解冻结后的模型;误差则由系数的变化控制。伪微分算子把这个“每个点冻结一次”的过程系统化了。

    4. Kernel 与奇异积分

    伪微分算子也可以写成 kernel 形式

    $$Af(x)=\int K(x,y)f(y)\,dy.$$

    当 symbol 的阶数 $m=0$ 时,kernel 在 $x=y$ 附近有奇异性,但又满足 cancellation 和光滑估计。这就把问题接到 Calderon-Zygmund 理论:$L^2$ 有界性来自 Plancherel,弱 $(1,1)$ 或 $L^p$ 有界性则依赖实变量分解。

    因此,伪微分算子不是孤立的一套语言,而是 Fourier 分析和实变量奇异积分之间的桥。

    5. Symbolic calculus

    平移不变情形下,两个 multiplier 的复合对应乘子相乘。伪微分算子也有类似结论,只是要加入低阶修正。形式上,如果 $A$ 的 symbol 是 $a$,$B$ 的 symbol 是 $b$,那么 $AB$ 的 symbol 有渐近展开

    $$a\# b(x,\xi)\sim \sum_\alpha \frac1{\alpha!}\partial_\xi^\alpha a(x,\xi)D_x^\alpha b(x,\xi).$$

    主项是 $ab$,后面的项阶数越来越低,表示越来越强的 smoothing。这个公式是椭圆正则性、参数构造和微局部分析的基本计算规则。

  • 奇异积分一般理论一瞥:近似恒等、Fourier multiplier 与 square function

    旧博客原文

    原题:A glimpse to the general theory

    1. Introduction

    We have talked about a very basic result in singular integral, i.e. if we have an additional condition, i.e. {q-q} bounded condition, then by interpolation theorem we only need to establish the weak {1-1} bound then we establish the {p-p} bound of {T}, {\forall 1< p< q }.

    The category of of singular integral is very general, in fact the singular integral we interested in always equipped more special structure. We discuss following 3 types result which world be the central role in this further series note.

    1. Approximation of the identity.
    2. Singular integral with {L^2} bounded translation invariant operator.
    3. Maximal function, singular integral, and square functions.

    The underlying object we consider in both the three case is some special singular integral, in the first case, it looks like a {T=sup_{t>0} \Phi_{t}*f}, this, among the other thing, has a close relationship with the maximal operator {Mf}. This is discussed in 2. For the singular integral with {L^2} bound, the Fourier transform or its discretization version, Fourier series is natural involved. And there is a “representation theorem” similar to the sprite of Reisz representation theorem, said, roughly speaking, if we consider the {L^2} bound operator adding the condition of transform invariant, then it is really coinside with the case of our image, the operator must behaviour as a Fourier multiple. This is the contant of famous Mikhlin multiplier theorem, and we discuss some technique difficulty in the process of establishing such a theorem, this is the contact of 3. At last we discuss some deep relationship between three basic underlying intution and objects in harmonica analysis, the Maximal function, singular integral, and square functions. They could all be understanding as tools to understanding the variant complicated emerging in singular integral. But there is definitely some common points. This is the theme of 4. Of course there are some further topic which are also interesting, but I do not want to discuss them here, maybe somewhere else.

    2. Approximation of the identity

    First topic, we discuss the approximation of the identity, this play a central role in understanding solution of PDE, why, I think a key point is this tools carry a lots of information about the scaling of the space, as it well known, analysis could roughly divide into two parts, “hard analysis” and “soft analysis”, approximation of the identity supply a way to transform a result form “hard analysis” side to “soft analysis” side and reverse. And when it shows its whole power always along with the involving of following Dominate convergence theorem:

    Theorem 1 (DCT)

    Let {\{f_n\}_{n=1}^{\infty}} be a series of function on measure space {(X,\Sigma,\mu)}, and {f_n \rightarrow f, a.e. x\in X}, and {\{f_n\}_{n=1}^{\infty}} satisfied a controlling condition, i.e. we can find a integrable function {g\in L^{1}(X)}, such that {|f_n(x)|\leq |g(x)|, a.e. x\in X, \forall n\in {\mathbb N}^*}, then we know,

    \displaystyle \lim_{n\rightarrow \infty} \int_{X}f_n(x)d\mu\rightarrow \int f(x)d\mu \ \ \ \ \ (1)

     

    In fact we have even stronger,

    \displaystyle \lim_{n\rightarrow \infty} \int_{X}|f_n(x)-f(x)|d\mu=0 \ \ \ \ \ (2)

     

    This is a standard theorem in real analysis, we give the proof.

    Proof: {f} is the point-wise limit of {f_n} so we know f is measurable and also dominate by {g}, so by triangle inequality we have:

    \displaystyle |f-f_n|\leq 2|g|

    Then the 1 is trivially true, due to a diagonal taking subsequences trick. For more subtle result 2, we need use reverse Fatou theorem to show it is true, roughly speaking we have,

    \displaystyle \limsup_{n\rightarrow \infty}\int_{X}|f_n-f|\leq \int_{X}\limsup_{n\rightarrow \infty}|f_n-f|=0

    The key point is the first inequality above used the reverse Fatou theorem. \Box

    Now we discuss of the main result of the approximation identity. So first we need to define what is a approximation identity. a key ingredient is scaling. i.e. we given a function {\Phi} and consider {\Phi_t=t^{-n}\Phi(\frac{x}{t})}, and we wish,

    \displaystyle \lim_{t\rightarrow 0}(f*\Phi_t(x))=f(x), for a.e. x\ \in {\mathbb R}^n \ \ \ \ \ (3)

     

    Whenever {f\in L^p, 1\leq p\leq \infty}, but there need some technique assume to make this intuition to be tight, this lead the following definition.

    Definition 2 (Approximation of the identity) Suppose {\Phi} is a fixed function on {{\mathbb R}^n} that is appropriated small at infinity (have good enough decay rate), for example, take,

    \displaystyle |\Phi(x)|\leq A(1+|x|)^{-n-\epsilon} \ \ \ \ \ (4)

     

    Then we define {\{\Phi_t:\Phi_t(x)=t^{-n}\Phi(\frac{x}{t})\}} to be an approximation of the identity.

    The key theorem is the following, related the approximation of the indentity with the maximal operator.

    Theorem 3

    \displaystyle \sup_{t>0}|(\Phi_t*f)(x)|\leq c_{\Phi}Mf(x) \ \ \ \ \ (5)

     

    For heat kernel, the thing is more subtle.

    Theorem 4 [Heat kernel estimate]

    \displaystyle \|f-e^{t\Delta}f\|_2\leq \|\nabla f\|_2\sqrt{t} \ \ \ \ \ (6)

     

    Remark 1 I know this theorem from Lieb’s book. The power of 4 combine with Plancherel theorem could use to establish the Sobolev inequality, at least for the index {p=2}.

    There are 3 ingredients which cold be useful.

    1. the power of Rearrangement inequality involve in the Approximation of indentity operator. we could consider the relationship between {f*\Phi_t} and {f*\overline \Phi_t}, where {\overline \Phi} is constructed by take the average of {\Phi} on the level set but the foliation of scaling. Intuition seems some monotonic property natural emerge.
    2. There is a discretization model, i.e. the toy model on gragh, or we think it as correlation between particles, the key point is the rescaling deformation could be instead by semi group or renormalization property.
    3. We consider the more general case, now there is not only one {\Phi} but a group of them, i.e. {\Phi_k, k\in A}, this will involve some amenable theory I think.

    We give two of the original and most important examples, First, if

    \displaystyle \Phi(x)=c_n(1+|x|^2)^{\frac{-(n+1)}{2}}

    where

    \displaystyle c_n=\frac{\Gamma(\frac{n+1}{2})}{\pi^{\frac{n+1}{2}}}

    then {\Phi_t(x)} is the possion kernel, and,

    \displaystyle u(x,t)=(f*\Phi_t)(x)

    Gives the solution of the Dirichlet problem for the upper half space,

    \displaystyle {\mathbb R}^{n+1}_{+}=\{(x,t):x\in {\mathbb R}^n,t>0\}

    Namely

    \displaystyle \Delta u=(\frac{\partial^2}{\partial t^2}+ \sum_{j=1}^n\frac{\partial^2}{\partial x_j^2})u(x,t)=0,\ u(x,0)\equiv f(x) \ \ \ \ \ (7)

     

    The second example is the Gaussian kernel,

    \displaystyle \Phi(x)=(4\pi)^{-\frac{n}{2}}e^{-\frac{|x|^2}{4}}.

    This time, if {u(x,t)=(f*\Phi_{t^{\frac{1}{2}}})(x)}, then {u} is a solution of the heat equation,

    \displaystyle (\frac{\partial}{\partial t}- \sum_{j=1}^n\frac{\partial^2}{\partial x_j^2})u(x,t)=0,\ u(x,0)\equiv f(x) \ \ \ \ \ (8)

     

     

    3. Singular integral with {L^2} bounded translation invariant operator

    The main result proved in last note about singular integral is a conditional one, guaranteeing the boundedness on {L^p} for a range {1<p\leq q}, on the assupution that the boundedness on {L^q} is already known; the most important instance of this occurs when {q=2}. In keeping with this, we consider bounded linear transformation {T} from {L^2({\mathbb R}^n)} to itself that commute with translation. As is well known, such operator are characterized by the existence of a bounded function {m} on {{\mathbb R}^n} (the “multiper”), so that {T} can be realized as,

    \displaystyle \widehat{Tf(\xi)}=m(\xi)\widehat f(\xi) \ \ \ \ \ (9)

    Where {\widehat{}} denotes the Fourier transform. Alternatively, at least on test function {f\in S}, {T} can be realized in terms of convolution with a kernel {K},

    \displaystyle Tf=f*K \ \ \ \ \ (10)

     

    Where {K} is the distribution given by {\hat K=m}. We shall now examine how the theorem with condition on singular integral weill lead to some result of this type of operator. Roughly speaking, it is due to now we know the boundedness on {L^2}, for technique condition, we need to assume the distribution {K} agree away from the origin with a function that is locally integrable away from the origin with a function that is locally integrable away from the origin; in this case we define the function by {K(x)}. Then 10 implies that,

    \displaystyle Tf(x)=\int K(x-y)f(y)dy,\ for \ a.e. x\notin supp f. \ \ \ \ \ (11)

    Whenever {f} is in {L^2} and {f} has campact support. Tis is the representation of singular integral in the present context. Next, the crucial hormander condition is then equivalent with,

    \displaystyle \int_{|x|\geq c|y|}|K(x-y)-K(x)|dx\leq A \ \ \ \ \ (12)

     

    for all {y\neq 0}, where {c>1}. In this case, the condition 12 have a further understanding, in fact,

    Lemma 5

    \displaystyle |(\frac{\partial}{\partial x}^{\alpha}K(x))|\leq A_{\alpha}|x|^{-n-|\alpha|},\ for\ all \ \ \alpha \ \ \ \ \ (13)

     

    or its weaker form, (here {\gamma>0} is fixed )

    \displaystyle |K(x-y)-K()|\leq A\frac{|y|^{\gamma}}{|x|^{n+\gamma}}, whenever \ |x|\geq c|y|. \ \ \ \ \ (14)

    imply the hormander condition 12

    Proof: Integral by part. \Box

    So, now the key point is how do {K}, satisfied such conditions, come about? It turns out that, toughly speaking, such condition on {K} have equivalent versions when sated in terms of the Fourier transform of {K}, namely the multiper {m}. This is transform the difficulties from physics space to fractional space In the future note, we will find a proof of the following Theorem:

    Theorem 6 For {m=\hat K}.

    If we assume that,

    \displaystyle |(\frac{\partial}{\partial \xi}^{\alpha}m(\xi))|\leq A'_{\alpha}|\xi|^{-n-|\alpha|},\ for\ all \ \ \alpha \ \ \ \ \ (15)

    holds for all {\alpha}, then {K} satisfied 5 for all {\alpha}.

    If we assume that {m} satisfied the above inequality for all {0\leq |\alpha| \leq l}, where {l} is the smallest integer {>\frac{n}{2}}, then {K} satisfied 12

    Remark 2 The multiplier {m} satisfied the second part condition of 6, are called Marcinkiewicz mulltiplier.

     

    4. Maximal function, singular integral, and square functions.

     


    补充说明

    以下是新整理的中文说明;上方旧博客原文保持不变。

    奇异积分的基本理论不只是 Calderon-Zygmund 分解。真正进入一般理论以后,会反复出现三类对象:approximation of the identity、translation invariant singular integrals,也就是 Fourier multipliers,以及 maximal function 和 square function。

    奇异积分一般理论一瞥:近似恒等、Fourier multiplier 与 square function
    奇异积分一般理论中,近似恒等、Fourier multiplier 和 square function 分别控制尺度极限、平移不变结构和多尺度正交性。

    1. 从弱型估计到 $L^p$ 有界性

    如果一个奇异积分算子 $T$ 已知在 $L^2$ 上有界,并且满足弱 $(1,1)$ 估计

    $$|\{x:|Tf(x)|>\lambda\}|\lesssim \frac{\|f\|_1}{\lambda},$$

    那么 Marcinkiewicz interpolation 给出 $1

    2. Approximation of the identity

    取一族核 $\phi_t(x)=t^{-n}\phi(x/t)$,若 $\int\phi=1$,则

    $$\phi_t*f\to f\quad(t\to0).$$

    这类算子看起来温和,但它们和 maximal operator 紧密相连。控制

    $$\sup_{t>0}|\phi_t*f(x)|$$

    本质上就是控制函数在不同尺度上的平均行为。

    3. Translation invariant operators

    若 $T$ 与平移可交换,那么 Fourier transform 会把它对角化:

    $$\widehat{Tf}(\xi)=m(\xi)\widehat f(\xi).$$

    Mikhlin multiplier theorem 给出一套可检验条件:若

    $$|\partial^\alpha m(\xi)|\lesssim |\xi|^{-|\alpha|}$$

    到足够阶数成立,则 $T$ 在 $L^p$ 上有界。

    4. Square functions

    square function 把函数分解到不同频率或尺度:

    $$Sf(x)=\left(\sum_j |P_jf(x)|^2\right)^{1/2}.$$

    它不是只估计每一块,而是用正交性追踪所有尺度的总能量。这是 Littlewood-Paley 理论的核心。

    5. 一条总线

    近似恒等处理尺度极限,multiplier theory 处理平移不变结构,square function 处理多尺度正交性。奇异积分的一般理论,就是在这三种结构之间来回切换。

  • Calderon-Zygmund 分解:奇异积分的实变量入口

    旧博客原文

    原题:Calderon-Zygmund theory of singular integrals.

    1. Calderon-Zygmund decomposition

    The Calderon-Zygmund decomposition is a key step in the real variable analysis of singular integrals. The idea behind this decomposition is that it is often useful to split an arbitrary integrable function into its “small” and “large” parts, and then use different technique to analyze each part.

    The scheme is roughly as follows. Given a unction { f} and an altitude { \alpha}, we write { f=g+b}, where { |g|} is point wise bounded by a constant multiple {\alpha}. While { b} is large, it does enjoy two redeeming features: it is supported in a set of reasonable small measure, and its mean value is zero on each of the ball that constitute its support. To obtain the decomposition { f=g+b}, one might be tempted to “cut” { f} at the height { \alpha}; however, this is not what works. Instead, one bases the composition on the set where the maximal function of { f}has height { \alpha}.

    Theorem 1 (Calderon-Zygmund decomposition)

    Suppose we are given a function { f\in L^1} and a positive number { \alpha}, with {\alpha>\frac{1}{\mu(R^n)}\int_{R^n}|f|d\mu}. Then there exists a decomposition of { f}, {f=g+b}, with { b=\sum_{k}b_k}, and a sequences of balls {\{B_k^*\}}, so that,

    1. { |g(x)|\leq c\alpha}, for a.e. { x}.
    2. Each {latex b_k} is supported in {B_k^*},{ \int|b_k(x)|d\mu(x)\leq c\alpha\mu(B_k^*)}, and { \int b_k(x)d\mu(x)=0}.
    3. { \sum_k\mu(B_k^*)\leq \frac{c}{\alpha}\int|f(x)|d\mu(x)}.

     

    Before proof this theorem, I explain the geometric intuition why this theorem could be true first. Merely speaking, this is just base on cut off the function into two part, the part with high altitude and the part with low altitude and extension the part with high altitude to make the extension one satisfied the condition 2 and 3.

    Proof: In fact this decomposition have a good geometric explain, we just divide the part {\{x: |f(x)|>\alpha\}} and extension it carefully to make they behaviour like several balls, to satisfied the special condition on this part. \Box

    Remark 1 Remark 1: A Calderon-Zygmund decomposition for {L^p} function was done in Charlie Fefferman’s thesis; see Section II of ams.org/mathscinet-getitem?mr=257819  One can also find this in Loukas Grafakos’s Classical Fourier Analysis Classical Fourier Analysis page 303 exercise 4.3.8. The question is broken up into parts that should be easy to handle.

    Several people have considered with this question. An excellent paper that comes to mind is Anthony Carbery’s Variants of the Calderon–Zygmund theory for { L^p}-spaces which appeared in Revista Matematica Iberoamericana, Volume 2, Number 4 in 1986. There are also several useful references that appear in Carbery’s paper.

    Remark 2 We could also consider a variant of Calderon-Zygmund decomposition, such as equipped with a nontrivial weight function { w} or find some different way to decomposition for some special purpose.

    Remark 3 Consider suitable decomposition of the physics space or even both the physics space and fractional space try to gain some reasonable estimate is a fundamental philosophy in harmonic analysis, beside the Calderon-Zegmund decomposition,

    Whitney decomposition. Which is important trick in the proof of fefferman-stein restriction theorem and differential topology.

    Wave packet decomposition. The wave packet decomposition. This decomposition underlies the proof of Carleson’s theorem (this is more explicit in Fefferman’s proof than Carelson’s original proof), Lacey and Thiele’s proof of the boundedness of the bilinear Hilbert transform, as well as a host of follow-up work in multilinear harmonic analysis. The idea of the wave packet decomposition is to decompose a function/operator in terms of an overdetermined basis. This allows one to preserve symmetries (such as modulation symmetries) that aren’t preserved by a classical Calderon-Zygmund decomposition (which endows the frequency with a distinguished role). One might consider using a wave packet decomposition if is working with an operator that has a modulation symmetry. This is discussed in more detailed in Tao’s blog post on the trilinear Hilbert transform.

    Polynomial decomposition. The application of polynomial decomposition to harmonic analysis is more recent, and its full potential still seems unclear. Applications include Dvir’s proof of the finite field Kakeya conjecture, Guth’s proof of the endpoint multilinear Kakeya conjecture (and, indirectly, the Bourgain-Guth restriction theorems), Katz and Guth’s proof of the joints problem and Erdos distance problem, among many other results. Generally, the idea behind the polynomial decomposition is to partition a subset of a vector space over a field into a finite number of cells each of which contains roughly the same fraction of the original set. One further wishes that no low degree algebraic variety can intersect too many of the cells. In Euclidean space, the polynomial ham sandwich decomposition does exactly this. This allows one to, for instance, control linear (or, more generally, `low algebraic degree’) interactions between points in distinct cells. This has so far proven the most useful in incidence-type problems, but many problems in harmonic analysis, thanks to the translation symmetry of the Fourier transform, are inextricably linked with such incidence-type problems. See (again) Tao’s survey of this topic for a more detailed account.

     

    2. Singular integrals

    Have the Calderon-Zegmund decomposition in hand, now we proof a conditional one bounded result for singular integrals.

    The singular integral one is interested in are operator { T}, expressible in the form

    \displaystyle (Tf)(x)=\int_{R^n}K(x,y)f(y)d\mu(y) \ \ \ \ \ (1)

     

    Where the kernel { K} is singular near { x=y}, and so the expression is meaningful only if { K} is treated as a distribution or in some limiting sense. Now the particular regularization of { (Tf)(x)} may be appropriate depends much on the context, and a complete treatment of the issues thereby raised take us quite far afield.

    Let us limit ourselves to two closely related ways of dealing with the questions concerning the definability of the operator. One is to prove estimates for the (dense) subspace where the operator is initially defined. The other is to regularize the given operators by replacing it with a suitable family, and to prove the uniformly estimates for this family. This idea is similar occurring in spectral geometry when we wish to investigate the spectrum of some operator we try to consider some deformation, so deduce to control the spectrum of a seres of paramatrix, for example, consider the wave kernel or heat kernel rather than the passion kernel itself. Common to both methods is a priori approach: We assume some additional properties of the kernel, but then prove estimates that are independent of these “regularity” properties.

    We now carry out the first approach in detail. There will be two kinds of assumptions made about the operator. The first is quantitative: we assume that we are given a bound { A}, so that the operator { T} is defined and bounded on { L^q} with norm { A}; that is,

    \displaystyle \|T(f)\|_q\leq A\|f\|_q, \forall f, f\in L^q \ \ \ \ \ (2)

     

    Moreover, we assume that there is associated to { T} a measurable function { K} (that plays the role of its kernel), so that for the same constant { A} and some constant { c>1},

    \displaystyle \int_{R^n-B(y,c\delta)}|K(x,y)-K(x,\bar y)|d\mu(x)\leq A, \forall \bar y\in B(y,\delta)  \ \ \ \ \ (3)

     

    for all { y\in R^n, \delta>0}.

    The further regularity assumption on the kernel { K} is that for each { f} in {L^q} that has compact surppot, the integral coverages absolutely for almost all { x } in the complement of the support of { f}, and that equality holds for these { x}.

    Theorem 2 (Bounded of singular integral with condition)

    Under the condition 1 and 3 made above on { K}, the operator { T} is bounded in { L^p} norm on { L^p\cap L^q}, when { 1<p<q}. More precisely,

    \displaystyle \|T(f)\|_p\leq A_p\|f\|_p

    For { f\in L^p\cap L^q} with { 1<p<q}, where the bound { A_p} depends only on the constant { A} appearing in 1 and 3 and on { p}, but not on the assumed regularity of { K}, or on { f}.

     

    Proof:

    Now let us begin to prove the conditional theorem. The key point is to use the potential of {T} has been a bounded operator from {L^q\rightarrow L^q}. Said, it already assumed {\exists A>0} such that {\forall f\in L^q} we have {\|T(f)\|_q\leq A\|f\|_q}. Now let us look at the singular integral expression:

    \displaystyle (Tf)(x)=\int_{R^n}K(x,y)f(y)d\mu(y). \ \ \ \ \ (4)

     

     

    The key point is to proof the mapping {f\rightarrow T(f)} is a weak-type {1-1}; that is,

    \displaystyle \mu\{x:|Tf(x)|>\alpha\}\leq \frac{A'}{\alpha}\int |f|d\mu. \ \ \ \ \ (5)

     

    At once we establish 5, then the theorem followed by interpolation. Now we use theorem 1 on {f} get {f=g+b}, thanks to the triangle inequality and something similar we have {g,b \in L^q}, in fact {R^n= A\amalg B, B\cup_{k}B_k}, {g=\chi_A g+\chi_{B}g, b=\chi_A b+\chi_{B} b}, by triangle inequality and {f=g+b}, to proof {g,b \in L^q}, we only need to proof {\chi_A g, \chi_B g, \chi_A b, \chi_B b\in L^q}, but this is easy to proof.

    Now we know the {L^q} bounded of {g,b}, we divide the difficult of establish the weak 1-1 bound of {f} into the difficult of establish the weak 1-1 bound for {g} and {b}. i.e.

    \displaystyle \mu\{x:|Tf(x)|>\alpha\}\leq \mu \{x:|Tg(x)|>\alpha\} +\mu\{x:|Tb(x)|>\alpha\} \ \ \ \ \ (6)

     

    For {g}, if this weak 1-1 bound is not true, we have,

    \displaystyle \mu \{x:|Tg(x)|>\alpha\}\geq \frac{A'}{\alpha}\int |g|d\mu \ \ \ \ \ (7)

     

    thanks to the trivial estimate {\|g\|_q \leq c\alpha^{q-1}\|g\|_1 }. combine this two estimate we have:

    \displaystyle c\alpha^{q-1}\|g\|_1\geq \|g\|^q_q\geq c\|Tg\|^q_q \geq A'\alpha^{q-1} \|g\|_1 \ \ \ \ \ (8)

     

    The first estimate is true on {A} due to {|g|\leq \alpha, a.e. x\in R^n}. But compare the left and the right of 8 lead a contradiction, so 7 follows. For {b}, the thing is more complicated and in fact really involve the structure of the convolution type of the singular integral. The key point is controlling near the diagonal of {K(x,y)}. we warm up with a more refine decomposition {b=\sum b_k}, {\forall k, b_k=b\cdot \chi_{B_k}}. For a large constant {c>>1} choose later define {B^*_k=c B_k}. We know {b\in L^q}, but the really difficult thing occur in the how to combine the following 5 condition to lead a contradiction:

    1. {\|Tb\|_q\leq \|b\|_q}.
    2. property come from the Calderon-Zegmund decomposition, {\int_{B_k}\|b\|\leq c\alpha \mu(B_k),\forall k} and {\int_{B_k}b=0}.
    3. Hormander condition 3 , {\int_{R^n-B(y,c\delta)}|K(x,y)-K(x,\bar y)|d\mu(x)\leq A, \forall \bar y\in B(y,\delta)}
    4. the reverse of weak 1-1 of {b}, {\mu\{x:b(x)>\alpha\}> \frac{A'}{\alpha}\|b\|_1}.
    5. the structure {Tb(x)=\int_{R^n} K(x,y)b(y)dy}

    The first step is to break {b} into {b_k}, and reduce the case of several balls to the case of only one ball, this could be done by triangle inequality or more may be we could do it derectly, but any way it is not difficult.

    Then the thing become intersting, we focus on {b_1}, divide {Tb_1=T\chi_{B_1} b_1+ T\chi_{{\mathbb R}^n-B_1}b}. thanks to the hormander condition 3 we have good control on {T\chi_{{\mathbb R}^n-B_1^*}}, in fact we can proof a weak 1-1 bound on it,

    \displaystyle \mu\{x:|T_{{\mathbb R}^n-B_1^*}b_1|>\alpha\}< \frac{A'}{\alpha}\|b_1\|_1 \ \ \ \ \ (9)

    \displaystyle \begin{array}{rcl} T_{{\mathbb R}^n-B_1^*}b_1(x) & = & \int_{{\mathbb R}^n-B_1^*}K(x,y)b_1(y)dy\\ & = & \int_{{\mathbb R}^n-B_1^*}[K(x,y)-K(x,\bar y)]b_1(y)dy+\int_{{\mathbb R}^n-B_1^*}K(x,\bar y)b_1(y)dy\\ & \leq & \int_{{\mathbb R}^n-B_1^*}Ab_1(y)dy+\int_{{\mathbb R}^n-B_1^*}K(x,\bar y)b_1(y)dy \end{array}

    So we conclude,

    \displaystyle \begin{array}{rcl} \|T_{{\mathbb R}^n-B_1^*}b_1\|_1 & = & \int_{{\mathbb R}^n}|\int_{{\mathbb R}^n-B_1^*}K(x,y)b_1(y)dy|dx\\ & = & \int_{{\mathbb R}^n}\int_{{\mathbb R}^n-B_1^*}|[K(x,y)-K(x,\bar y)]b_1(y)dy|dx+\int_{{\mathbb R}^n}|\int_{{\mathbb R}^n-B_1^*}K(x,\bar y)b_1(y)dy|dx\\ & \leq & A\int_{{\mathbb R}^n}b_1(y)dy+\int_{{\mathbb R}^n-B_1^*}K(x,\bar y)b_1(y)dy=A\int_{{\mathbb R}^n}b_1(y)dy \end{array}

    The last equality used the condition {\int b_1=0}.

    \Box

     


    补充说明

    以下是新整理的中文说明;上方旧博客原文保持不变。

    Calderon-Zygmund 分解是实变量调和分析里最基本的动作之一。它不是简单地把函数按高度截断,而是用 Hardy-Littlewood maximal function 找到真正危险的区域,再把函数分成一个有界的好部分和一族有 cancellation 的坏部分。

    Calderon-Zygmund 分解:奇异积分的实变量入口
    Calderon-Zygmund 分解用 maximal function 找到坏区域,把函数拆成有界部分和带 cancellation 的局部坏块。

    1. 为什么不能直接截断

    给定 $f\in L^1(\mathbb R^n)$ 和高度 $\lambda>0$,我们希望写成

    $$f=g+\sum_j b_j.$$

    其中 $g$ 应该满足 $|g|\lesssim \lambda$,而 $b_j$ 虽然可能很大,但支撑在小集合上,并且有零平均。直接令 $g=f\mathbf 1_{\{|f|\le \lambda\}}$ 并不够,因为奇异积分不是逐点算子,它会感受到坏集合附近的空间结构。

    正确做法是看 maximal function 的超水平集:

    $$\Omega=\{x:Mf(x)>\lambda\}.$$

    然后对 $\Omega$ 做 Whitney decomposition,得到一族彼此控制重叠的 cubes 或 balls。

    2. 分解定理

    Calderon-Zygmund 分解给出

    $$f=g+\sum_j b_j,$$

    并且满足:

    $$|g(x)|\lesssim \lambda \quad\text{a.e.},$$

    $$\operatorname{supp} b_j\subset Q_j,\qquad \int b_j=0,$$

    以及

    $$\sum_j |Q_j|\lesssim \frac{\|f\|_1}{\lambda},\qquad \sum_j\|b_j\|_1\lesssim \|f\|_1.$$

    这里的零平均是关键。奇异积分核在远离 $Q_j$ 的地方可以用光滑性做差分估计,从而把 $b_j$ 的大振幅抵消掉。

    3. 用它证明弱 $(1,1)$

    设 $T$ 是 Calderon-Zygmund singular integral,并已知 $T$ 在 $L^2$ 上有界。为了证明

    $$|\{x:|Tf(x)|>\lambda\}|\lesssim \frac{\|f\|_1}{\lambda},$$

    把 $f=g+b$。好部分用 $L^2$ 有界性处理:

    $$|\{|Tg|>\lambda/2\}|\lesssim \lambda^{-2}\|g\|_2^2\lesssim \frac{\|f\|_1}{\lambda}.$$

    坏部分先丢掉扩大后的坏 cubes,它们总体积可控;在外面利用 $\int b_j=0$ 写

    $$Tb_j(x)=\int_{Q_j}\bigl(K(x,y)-K(x,c_j)\bigr)b_j(y)\,dy,$$

    再用 kernel 的 Holder 光滑性求和。

    4. 更广的分解哲学

    Calderon-Zygmund 分解的精神是:在物理空间中识别坏区域,把坏区域局部化,并为每个坏块制造 cancellation。类似思想在 weighted theory、Hardy space、Whitney decomposition、wave packet decomposition 和 polynomial partitioning 中都会出现。

    不同分解保留不同对称性。Calderon-Zygmund 分解突出空间局部性;wave packet 分解同时追踪空间和频率;polynomial partitioning 则把几何 incidence 信息放进分析估计里。调和分析很多证明,本质上都是在寻找适合当前算子的分解方式。

  • Almost orthogonality:Cotlar-Stein lemma、Schur test 与奇异积分

    旧博客原文

    原题:Almost orthogonality

     

    Motivation and Cotlar’s lemma

    We always need to consider a transform T on Hilbert space l^2(\mathbb Z) (this is a discrete model), or a finite dimensional space V. If under a basis T is given by a diagonal matrix this story is easy,

    \displaystyle A = \begin{pmatrix} \Lambda_1 & 0 & \ldots & 0 \\ 0 & \Lambda_2 & \ldots & 0 \\ \vdots & \vdots & \ddots & \vdots \\ 0 & 0 & \ldots & \Lambda_n \end{pmatrix} \ \ \ \ \ (5)

    Then ||T||=\max_{i}\lambda_i.

    In fact, for T is a transform of a finite dimensional space, T is given by (a_{ij})_{n\times n} by duality we have ||T||=||TT^*||, so we have,

    ||T||=||TT^*||=|(\sum a_{ij}x_j)y_i|\leq |\sum_{i,j}\frac{1}{2}(|a_{ij}(|x_i|^2+|y_j|^2)|\leq M

    If we have given \sum_{i}|a_{ij}|\leq M and \sum_{j}|a_{ij}|\leq M \forall i,j\in \{1,2,...,n\}.

    But in application of this idea, the orthogonal condition always seems to be too restricted and due too this we have the following lemma which is follow the idea but change the orthogonal condition by almost orthogonal.

    Lemma(Catlar-Stein)

    Let \{T_j\}_{j=1}^N be finitely many operators on some Hilbert space H. Such that for some function \gamma : \mathbb Z\to R^+ one has,

    ||T_j^*T_k||\leq \gamma^2(j-k),||T_jT_k^*||\leq \gamma^2(j-k)

    for any 1\leq j,k\leq N. Let \sum_{l=-\infty}^{\infty}\gamma(l)=A<\infty. then ,

    ||\sum_{j=1}^NT_j||\leq A

    Pf:

    tensor power trick + duality ||T||=||TT^*||^{\frac{1}{2}}.

    Singular integrals on L^2

     

    Lemma(Schur)

    Define T is a operator on measure space X\times Y equipped positive product measure \mu\wedge \nu, via,

    (Tf)(x)=\int_YK(x,y)f(y)\nu(dy)

    K is a measurable kernel, then,

    1). ||T||_{1\to 1}\leq \sup_{y\in Y}\int_{X}|K(x,y)|\mu(dx)=:A.

    2). ||T||_{\infty\to \infty}\leq \sup_{x\in X}\int_{Y}|K(x,y)|\nu(dy)=:B.

    3). ||T||_{p\to p}\leq A^{\frac{1}{p}}B^{\frac{1}{p'}},  \forall 1\leq p\leq \infty.

    4). ||T||_{1\to \infty}\leq ||K||_{L^{\infty}(X\times Y)}.

    Pf:

    1),2),4) merely due to Fubini theorem and Bath lemma.

     

    3) proof by the interpolation and combine 1) and 2).

    Theorem

    Let K be a Calderon-Zegmund operator, with the additional assumption

    that |\nabla K(x)|\leq B|x|^{-d-1}. Then

    ||T||_{2\to 2} \leq CB

    with C = C(d).

    Caldero ́n–Vaillancourt theorem

     

    Hardy’s inequality

    Theorem(Hardy inequality)

    For any 0 \leq s < \frac{d}{2} there is a constant C(s, d) with the prop-

    arty that,

    |||x|^{-s} f||_2 \leq C(s,d)||f||_{H^s(R^d)}

    for all f \in H^s(R^d).

     


    补充说明

    以下是新整理的中文说明;上方旧博客原文保持不变。

    正交性是 Hilbert space 中最强的简化机制;almost orthogonality 则是在真实分析问题中更常见的替代品。频率块、空间块和算子族通常不完全正交,但交互足够小。

    Almost orthogonality:Cotlar-Stein lemma、Schur test 与奇异积分
    almost orthogonality 用交互衰减替代严格正交,是奇异积分和伪微分算子估计的基础。

    1. 从正交到几乎正交

    若 $T_j$ 的像彼此正交,则

    $$\left\|\sum_j T_j f\right\|_2^2=\sum_j\|T_jf\|_2^2.$$

    实际中通常只有 $T_i^\ast T_j$ 和 $T_iT_j^\ast$ 随 $|i-j|$ 衰减。

    2. Cotlar-Stein lemma

    $$\|T_i^\ast T_j\|+\|T_iT_j^\ast\|\le a(i-j)$$

    且 $\sum_k a(k)^{1/2}<\infty$,则

    $$\left\|\sum_jT_j\right\|_{2\to2}<\infty.$$

    证明可用 tensor power trick 和对偶性。

    3. Schur test

    对 kernel operator

    $$Tf(x)=\int K(x,y)f(y)\,dy,$$

    $$\sup_x\int |K(x,y)|dy<\infty,\qquad \sup_y\int |K(x,y)|dx<\infty,$$

    则 $T$ 在 $L^2$ 上有界。

    4. 奇异积分中的用途

    Calderon-Zygmund theory 中,常把算子分解成不同尺度的 pieces。尺度相隔很远时,kernel 的光滑性带来交互衰减;相近尺度则只需有限重叠。

    5. Calderon-Vaillancourt 方向

    伪微分算子的 $L^2$ 有界性也可以看成 almost orthogonality 的结果:把相空间切成小块后,不同块之间的交互由 symbol 的导数控制。

  • 一个 Fourier 系数衰减估计:解析延拓、分歧提升与矩阵值函数

    旧博客原文

    原题:An Fourier coefficient decay estimate.

    f is a matrix value analytic function on \mathbb T, we know h(f)>\alpha, this is just mean \forall k\in \mathbb Z , assume | \hat f(k)|\leq e^{-|k|\alpha} , g=log(f) for which we assume g is a lifting of f use the inverse of ramification map ,

    M_{n\times n} \to M_{n\times n} , A \to e^A.

    Then exists \beta=c(\alpha)>0 such that ,

    \forall k \in \mathbb Z, |\hat g(k)| \leq e^{-k\beta}.


    补充说明

    以下是新整理的中文说明;上方旧博客原文保持不变。

    设 $f$ 是圆周或环面上的矩阵值解析函数。解析性最直接的 Fourier 后果是指数衰减:

    $$\|\widehat f(k)\|\le C e^{-\rho |k|}.$$

    这里 $\rho$ 不是一个装饰性常数,而是由 $f$ 能延拓到多宽的复邻域决定。若函数还经过分歧映射的提升,衰减率会被分歧阶数重新缩放。

    一个 Fourier 系数衰减估计:解析延拓、分歧提升与矩阵值函数
    解析函数的 Fourier 系数指数衰减,衰减率由复邻域宽度和可能的分歧提升共同决定。

    1. 一维解析情形

    如果 $f$ 在带状区域 $|\operatorname{Im} z|<\rho$ 内解析且有界,那么

    $$\widehat f(k)=\int_0^1 f(x)e^{-2\pi i kx}\,dx.$$

    当 $k>0$ 时把积分路径向上平移 $i\sigma$,得到因子 $e^{-2\pi k\sigma}$;当 $k<0$ 时向下平移。令 $\sigma<\rho$,就得到指数衰减。

    2. 矩阵值函数没有本质困难

    若 $f(x)$ 取值于矩阵空间,只需要把绝对值换成任意相容矩阵范数。Cauchy 积分和路径平移仍逐项成立,因此

    $$\|\widehat f(k)\|\le \sup_{|\operatorname{Im}z|\le\sigma}\|f(z)\|\,e^{-2\pi\sigma |k|}.$$

    3. 分歧提升的影响

    若 $f$ 通过某个 ramification map 被提升,例如局部写成 $z=w^m$,那么 $w$ 平面中的解析带宽会按比例改变。直观上,绕分歧点一圈的角变量被拉伸,Fourier 频率也随之重标定。

    4. 估计的用途

    这类估计常出现在解析 cocycle、准周期系统和小除数问题中。指数衰减提供了“高频很小”的精确版本,使得截断 Fourier 级数时可以把尾项压到可控范围。

  • 从 linear 到 multilinear:Cauchy-Schwarz、分解与复杂度下降

    旧博客原文

    原题:Linear to Multi-linear

    The technique that transform a problem which is in a linear setting to a multilinear setting is very powerful.

    such like:

    1.The renormalization technique in complex dynamic system, and the generalization

    this is mainly the Ostrowoski representation,and something else.

    2.Fouriour analysis

     

    this can be view when it is difficult to investigate a quality about a function f, it is always easier to take charge with some some part of f, in this case is given by \hat f(\xi),\xi \in R or \hat f(k),k\in Z like cut f into a lot of small parts,deal with every part and use some inequality(always the triangle inequality or similar thing) to glue it into a whole estimate of the quantity of f.

    3.Multi-scales theory

     

    this is used in the improve of Minkowski dimension of 3-dim kakeya set by Katz-Tao.

    4.The proof of Bourgain-Sarnak-Ziegler theorem

     

    Theorem(B-S-Z). Let F : N \to C with |F| \leq 1 and let \nu be a multiplicative function with |\nu| \leq 1. Let \tau > 0 be a small parameter and assume that for all primes p_1, p_2 \leq e^{1/\tau} , p_1 \neq p_2, we have that for M large enough

    |\sum_{m\leq M} F(p_1m)\overline {F(p_2m)}| \leq \tau M.

    Then for N large enough

    |\sum_{m\leq M} \nu(n)F(n) | \leq 2 \sqrt{\tau log(\frac{1}{\tau})}M.

    this theorem is not difficult to prove by bilinear method and Cauchy-Schwarz,you can see the detail in https://arxiv.org/abs/1110.0992v1.
    According to this theorem,to get a good approximation of \sum_{1\leq k\leq x}\mu(k)f(k) we use need a good approximation on \sum_{1\leq k\leq x}f(p_1x)\overline {f(p_2x)}.this will be much easier.but for the RHS f(x) is very complicated so I do not have a non-trivial estimate for \sum_{1\leq k\leq x}f(p_1x)\overline {f(p_2x)} until now.

    5.The multiplier restriction theorem(Tao)

     

    6.Some special construction in additive Combitriocs

     

    Such like when we want to consider some set A\subset Z_1 satisfied \frac{|A-A|}{|A+A|}>>1 it is convenient to consider in a high dimensional linear space Z^N rather than in Z.


    补充说明

    以下是新整理的中文说明;上方旧博客原文保持不变。

    很多问题直接估计一个线性量很难,但把它平方、分解或转成多线性形式后,结构反而清楚。这就是从 linear 到 multilinear 的基本哲学。

    从 linear 到 multilinear:Cauchy-Schwarz、分解与复杂度下降
    从 linear 到 multilinear 的基本动作是用 Cauchy-Schwarz 或分解降低复杂度,同时保留主要信息。

    1. Cauchy-Schwarz 的真正作用

    Cauchy-Schwarz 不只是让估计变粗。好的使用方式是:降低复杂度,同时几乎不损失主要信息。例如

    $$\left|\sum_n a_n b_n\right|^2\le \left(\sum_n|a_n|^2\right)\left(\sum_n|b_n|^2\right).$$

    如果右边的两个平方和更容易理解,这一步就是有效的。

    2. Fourier 分解

    在调和分析中,一个函数被切成频率块:

    $$f=\sum_\theta f_\theta.$$

    线性估计可能无法直接控制 $\sum_\theta f_\theta$,但双线性或多线性估计能利用不同 $\theta$ 之间的 transversal 结构。

    3. Multiscale 理论

    Kakeya 和 restriction 中,多尺度分解把一个集合或函数切成不同尺度上的 pieces。每一层可能只给出局部信息,但层与层之间的递归关系能产生全局维数或范数估计。

    4. BSZ 准则

    Bourgain-Sarnak-Ziegler 准则就是线性到双线性的典型例子。要证明

    $$\sum_{n\le N}\mu(n)a_n=o(N),$$

    可以转而控制

    $$\sum_{n\le N}a_{pn}\overline{a_{qn}}$$

    对不同素数 $p\ne q$ 的相关。这把乘法函数的线性相关转成序列自身的双线性相关。

    5. 方法的边界

    多线性化不是免费午餐。右边的量可能更复杂,分解后还要重新 glue 回整体估计。好的多线性方法,总是在“降低复杂度”和“保持信息”之间找到平衡。

  • Wolff 的 Kakeya maximal estimate:X-ray 估计与三维障碍

    旧博客原文

    原题:Kakeya conjecture (Tomas Wolff 1995)

    There is a main obstacle to improve the kakeya conjecture,remain in dimension 3,and the  result established by Tomas Wolff in 1995 is almost the best result in R^3 even until now.the result establish by Katz and Tao can be view as a corollary of Wolff’s X-ray estimate.

    For f\in L_{loc}^1(R^d),for 0<\delta<1:

    f_{\delta}^*:P^{d-1}\longrightarrow R.f_{\delta}^*(e)=\sup_{T}\frac{1}{|T|}\int_{T}|f|.

    T is varise in all cylinders with length 1.radius \delta.axis in the e direction.

    f_{\delta}^{**}:R^d\longrightarrow R.f^{**}_{\delta}(x)=sup_{T}\frac{1}{|T|}\int_{T}|f|.

    T varise in cylinders contains x,length 1,radius \delta.

    Keeping this two maximal function in mind,we give the statement of the Kakeya maximal function conjecture:

    ||M_{\delta}f||_d\leq C_{\epsilon} \delta^{-\epsilon}||f||_d

    Where M_{\delta}=f_{\delta}^* or M_{\delta}=f_{\delta}^{**}.

    Because we have the obviously 1-\infty estimate:

    ||f^*_{\delta}||_{\infty}\leq\frac{||f||_1}{|T|}=\delta^{1-d}||f||_1.

    ||f^{**}_{\delta}||_{\infty}\leq\frac{||f||_1}{|T|}=\delta^{1-d}||f||_1.

    So by the Riesz-Thorin interpolation we have:

    ||M_{\delta}f||_{q}\leq C_{\epsilon}\delta^{-(\frac{d}{p}-1+\epsilon)}||f||_p.           (*)

    for 1\leq p\leq d,q\leq(d-1)p'.the task is establish (*) for (p,q) as large as posible in the range.

    for the 2 dimension case,the result is well know.the key estimate is:

    \sum_{j}|T_i\cap T_j|\leq log(\frac{1}{\delta})|T_i|

    for d\geq 3 case,the main result of Wolff is:

    ||M_{\delta}f||_q\leq C_{\epsilon}\delta^{-(\frac{d}{p}-1+\epsilon)}||f||_p

    hold for p=\frac{d+2}{2}.q=(d-1)p'. M_{\delta}=f_{\delta}^* or f_{\delta}^{**}.

    Now we sketch the proof.

    prove f_{\delta}^*,f_{\delta}^{**} cases together.

    We can make some reduction:

    the first one is we can assume the sup of f is in a fix compact set.

    the second is instead of consider f_{\delta}^{**},we can consider f_{\delta}^{***}(x)=\sup_{T}\frac{1}{|T|}\int_T|f|.

    where T varies in all cylinder with radius \delta,length 1,axis \frac{\pi}{100} with a fix direction.

    the first reduction is obvious(why?)

    the second reduction rely on a observe:

    ||f_{\delta}^{***}||_q\leq A(\delta)||f||_p          \Longrightarrow    ||f_{\delta}^{**}||_q\leq CA(\delta)||f|_p

    this is just finite cover by rotation of the coordinate and triangle inequality.

    now we begin to establish a frame and put the two situations f_{\delta}^*,f_{\delta}^{***} into it.

    Let M(d,1) be all line in R^d.

    then M(d,1)=R^d\times S^{d-1}/\sim is a 2d-2 dim manifold.

    M(d,1)\longrightarrow P^{d-1}

    l \longrightarrow  e_l

    e_l is the line parallel to l.and the middle point is original.

    dist(l_1,l_2)\sim \theta(l_1,l_2)+d_{mis}(l_1,l_2).

     

    Wolf axiom:

    (A,d) metric space.

    \mu(D(\alpha,\delta)) \sim \delta^m.\alpha\in A.\delta \leq diam(A).

    for certain m\in R^+.

    \forall \alpha\in A.F_{\alpha} \subset M(d,1) is given.and \bar{\cup_{\alpha}F_{\alpha}} is compact.

    d(\alpha,\beta)\lesssim inf_{l\in F_{\alpha};m\in F_{\beta}}dist(l,m) for all \alpha,\beta \in A.

    If f:R^d\longrightarrow R then we define M_{\delta}f:A\longrightarrow R by

    M_{\delta}f(\alpha)=\sup_{l\in F(\alpha)}\frac{1}{|T_{l}^{\delta}|}\int_{|T^{\delta}_l|}|f|.

    Property (**):

    If l_0\in \cup_{\alpha F_{\alpha}}. \Pi is a 2-plane.containing l_0and if \sigma \geq \delta and if \{\alpha_j\}_{j=0}^N is a \delta-seperated subset of A and for each j,there is l_j\in F_{\alpha_j} with dist(l_j,M(\Pi,l))<\delta and dist(l,l_0)<\sigma.then

    N\leq \frac{C\sigma}{\delta}

     


    补充说明

    以下是新整理的中文说明;上方旧博客原文保持不变。

    Wolff 1995 年的 Kakeya maximal estimate 是三维 Kakeya 理论中的经典结果。它把 tubes 的重叠问题转化为 maximal function 和 X-ray transform 的估计。

    Wolff 的 Kakeya maximal estimate:X-ray 估计与三维障碍
    Wolff 的 Kakeya maximal estimate 通过 X-ray 型平均和 incidence 几何控制 tubes 的重叠。

    1. Kakeya maximal function

    给定方向 $\omega$,考虑沿 $\omega$ 方向、半径 $\delta$、长度 $1$ 的 tube。Kakeya maximal function 粗略写为

    $$f_\delta^\ast(\omega)=\sup_{T\parallel\omega}\frac1{|T|}\int_T |f(x)|\,dx.$$

    猜想要求在尽可能大的 $p$ 范围内控制 $\|f_\delta^\ast\|_{L^q(S^{n-1})}$。

    2. 平凡估计与插值

    最直接的 $L^1$ 或 $L^\infty$ 估计通常太弱。通过 Riesz-Thorin interpolation 可以得到一些中间结果,但要接近 Kakeya 猜想,需要利用 tubes 的几何排列。

    3. X-ray transform

    X-ray transform 沿直线积分函数。Kakeya maximal function 可以看作对有限厚度 tubes 的 X-ray 型平均。Wolff 的估计用几何 incidence 控制这些平均的重叠。

    4. 归约步骤

    证明中可以把 tubes 限制在固定紧区域,并把方向分成有限个坐标 patch。这样 maximal function 可由有限方向族上的 tube averages 控制。

    5. 三维障碍

    三维最大的困难是 tubes 可以形成复杂 ruled surface 或 hairbrush 结构,普通 pairwise intersection 估计不够。Wolff 的结果抓住了这类结构的第一层几何约束,为 Katz-Tao 的后续改进提供了基础。

  • 三维 Kakeya 猜想:hairbrush、密度递降与 Katz-Tao 思路

    旧博客原文

    原题:Kakeya conjecture in R^3

    Kakeya conjecture in R^3 is very subtle.in fact wolff stay the best(but not very difficult to get,just use the structure so-called hairbrush)result \frac{5}{2} until the result of Katz and Tao \frac{5}{2}+\epsilon.Where \epsilon is a constant independent with kakeya set.and in the article of Tao,they proved \epsilon>\frac{1}{10^{10}}.

    Two-dimensional case

    first we overview the case of dimension 2,these is the only case that is proved.and the key point is the estimate:

    \mu(T_{i}\cap (\cup_{j\in I,j\neq i}T_j))<log(\frac{1}{\delta})\mu(T_i).

    where T_i=T_i(x_i,\theta_i) satisfied \cap_{i\in I}T_i is a \delta-neibeihood of kakeya set X.to remember one thing:this is equivalent to the maximal function version of kakeya conjecture,but for the minkoski version,there is a extra structure for the group I_{\delta} in different scales(this can be view as a multi-scale apporoach).

    this inequality is easy to proof.just observed that \mu(T_i\cap T_j)\sim \frac{1}{\theta_i-\theta_j}\delta^2.and to remember one thing:the inequality can be view as a uniformly estimate of overlap of the kakeya set,that is just mean the overlap would not concentrate to much at a lonely stick.this is enough to get a proof of the 2 dimension case just by a density decrement trick:we just not consider about the whole set I,but a low density subset \hat I\subset I,where \frac{|\hat I|}{|I|}\sim \delta^{\lambda},and make \lambda\to 0^+.

    Kakeya estimates

    Let \sigma\leq \delta\leq \theta<<1,and let T_{\delta} be a collection of \delta-tubes.whose set of directions all lie in a cap of radius \theta. Let 2<d<3 be fixed.
    • If we have a Kakeya estimate at some dimension d, and if the collection     T_{\delta} is direction-separated, then

     

    ||\sum_{i\in I}\chi_{T_i}||_{d'}\lesssim \delta^{\frac{d-3}{d}}\theta^{\frac{d+1}{d}}(1)
    • If we have an X-ray estimate at some dimension d, and if T_{\delta} consists ofessentially distinct tubes, then

     

     ||\sum_{i\in I}\chi_{T_i}||_{d'}\lesssim \delta^{\frac{d-3}{d}}\theta^{\frac{d+1}{d}}m^{1-\beta}
    for some β > 0, where m is the directional multiplicity of T_{\delta}.(2)

    So obviously the X-ray estimate is stronger than the kakeya estimate.it is just give the information of the overlap of the sticks with the same direction.

    in fact wolff have establish the X-ray estimate at dimension \frac{5}{2},so (2) just come from a rescaling argument.

    The sticky reduction

    renormalization process,just consider the process to make the thin sticks to be fat.and to proof this structure nearly has Markov property.but with a very small error term when change the scale.this is proved by the X-ray estimate.

    Triple intersection estimate

    Use Hardy-Litterwood-Soblev inequality,we can get a so called triple intersection estimate in general,said the triple intersection is smaller than the situation the 3 lines move together.and we just accosiate this to the cap-cup principle to get some information of the volume of X_{\delta}=\mu(\cup_{i\in I}T_i).img_0012

    Reduce to additive combination problem

    The right problem is just you have a n\times n cubes,and there is some sticks according them,if the distance of sticks is

    img_0011

     


    补充说明

    以下是新整理的中文说明;上方旧博客原文保持不变。

    Kakeya 猜想问:包含每个方向单位线段的集合是否必须有 full Hausdorff 或 Minkowski dimension?二维情形已经解决,三维情形则出现了 hairbrush、multiscale 和 sum-product 等深层结构。

    三维 Kakeya 猜想:hairbrush、密度递降与 Katz-Tao 思路
    三维 Kakeya 问题的困难来自不同方向细 tubes 的高重叠结构。

    1. Tube formulation

    令 $E_\delta$ 是 Kakeya set 的 $\delta$-neighborhood。Minkowski 维数问题等价于估计

    $$|E_\delta|\gtrsim_\varepsilon \delta^{\varepsilon}$$

    在合适维数归一化下的下界。更常用的语言是研究方向分离的 $\delta$-tubes 的重叠。

    2. 二维情形

    二维中,任意两个不同方向 tubes 的交叠容易控制。核心估计可以理解为:重叠不能集中在一根孤立 tube 附近。由此配合 density decrement,可推出二维 Kakeya 集合有 full dimension。

    3. Hairbrush 结构

    三维中,许多 tubes 可以围绕一根 tube 形成 hairbrush。Wolff 的观察是:如果很多 tubes 都与一根 tube 相交,那么它们的方向和位置仍然受到几何限制。这给出非平凡维数下界。

    4. Multiscale 难点

    Minkowski 版本比单尺度 maximal function 更微妙,因为不同尺度之间的结构会相互传递。一个尺度上的高重叠可能在下一尺度分裂,反过来又影响全局维数。

    5. Katz-Tao 方向

    Katz-Tao 的改进把 Kakeya 问题与 sum-product 现象联系起来。大致图像是:若 tubes 太集中,会诱导出同时具有加法和乘法结构的集合;sum-product 阻止这种集合太小。这是三维 Kakeya 后续发展的关键思想。