Preprint
Article

This version is not peer-reviewed.

Sharp Anisotropic Fractional Inequalities in Mixed Sobolev Spaces

Submitted:

24 September 2026

Posted:

24 September 2026

You are already at the latest version

Abstract
We develop a unified framework for anisotropic mixed fractional inequalities in which the fractional order and the integrability exponent may vary from one coordinate direction to another. Within this setting we establish a general family of Gagliardo--Nirenberg inequalities for the anisotropic mixed fractional Sobolev spaces, with an explicit homogeneity condition that corrects earlier formulations and with constants that depend on the effective dimension rather than on the ambient dimension. In the critical regime where the product of the fractional order and the integrability exponent equals the homogeneous dimension, the Sobolev embedding into the space of essentially bounded functions fails; we prove sharp logarithmic corrections that compensate for this failure, and we show by means of explicit extremal sequences that the logarithmic factor is optimal. We further develop a weighted theory for Muckenhoupt weights, extending the maximal, Gagliardo--Nirenberg, and Landau inequalities to weighted anisotropic mixed fractional Sobolev spaces with constants that depend polynomially on the Muckenhoupt characteristic. Optimal constants and extremal functions are studied through a variational problem whose Euler--Lagrange equation is a fractional Poisson-type equation with measure-valued source, and existence of extremals is proved by concentration-compactness. As applications, we derive a priori gradient bounds and sharp blow-up criteria for anisotropic fractional parabolic equations, and we prove that anisotropic neural networks achieve the optimal approximation rate governed by the effective dimension, with matching Kolmogorov width lower bounds. The resulting theory provides a common analytical framework linking fractional calculus, anisotropic function spaces, partial differential equations, and approximation theory. The endpoint cases of integrability are left as open problems.
Keywords: 
;  ;  ;  ;  

1. Introduction

The classical Landau inequality
∥ f ′ ∥ L ∞ ( R ) ≤ 2 ∥ f ∥ L ∞ ( R ) 1 / 2 ∥ f ′ ′ ∥ L ∞ ( R ) 1 / 2
was established by Landau in 1925 [1] and remains a cornerstone of mathematical analysis. It controls the magnitude of the first derivative by the amplitude and the second derivative of the function. Since then, the theory has evolved along several independent lines.
Historical development (1925–1998). The first major extension came with Stein’s theory of singular integrals and differentiability properties of functions [2] (1970), which provided the analytic machinery for maximal functions and multiplier theorems. In 1972, Muckenhoupt [3] introduced the A p class of weights, opening the door to weighted norm inequalities. The foundational treatise of Adams and Fournier [4] (1975) established the classical Sobolev embedding theorems, while Ditzian [7] (1989) extended the Landau–Kolmogorov inequality to multivariate settings. Samko’s monograph [8] (1993) systematised fractional integration and differentiation, and Grigor’yan [9] (1994) developed heat kernel upper bounds on non-compact manifolds. In 1996, Lions [10] introduced the concentration-compactness method, which became a fundamental tool for variational problems. Kounchev [11] (1997) studied extremisers for the multivariate Landau–Kolmogorov inequality, and DeVore [12] (1998) established the nonlinear approximation framework used in the neural network analysis. The treatise of Lieb and Loss [13] (2001) provided the interpolation inequalities needed for the fractional setting.
Modern theory (2002–2018). Triebel’s theory of anisotropic Besov spaces [15] (2006) supplied the function space framework for the anisotropic mixed fractional Sobolev spaces W α ν , p ( R k ) used in this paper. Yarotsky [17] (2017) established error bounds for approximations with deep ReLU networks, and Wang [18] (2018) proved the anisotropic Mihlin multiplier theorem that underlies our boundedness results for fractional derivatives.
Recent developments (2025–2026). Anastassiou [19] (2025) studied multivariate Canavati fractional Landau inequalities, primarily in isotropic settings. Chen [20] (2026) established sharp conditions for anisotropic fractional Gagliardo–Nirenberg inequalities on homogeneous Lie groups. The anisotropic fractional framework was initiated in our preceding work; the present paper corrects and substantially extends that work by: (i) fixing the homogeneity condition in the Gagliardo–Nirenberg inequality; (ii) replacing the divergent semigroup representation by a rigorous Littlewood–Paley decomposition; (iii) removing a circularity in the critical Landau inequality; (iv) providing explicit references for all multiplier theorems; and (v) expanding the literature review.
Contributions of this paper. This paper extends the anisotropic fractional Landau theory in four directions:
First, we establish a complete family of Gagliardo–Nirenberg inequalities in anisotropic mixed fractional Sobolev spaces W α ν , p ( R k ) :
∥ D i m f ∥ L q ≤ C ∥ f ∥ L r 1 − θ ∑ j = 1 k ∥ D j ν , α j f ∥ L p p θ / p ,
with the correct homogeneity condition
1 q = α i m D α + 1 − θ r + θ 1 p − ν D α .
Unlike previous results, the exponent q depends on the direction i through the anisotropic order α i m of the derivative.
Second, we address the critical case ν p = D α , where the Sobolev embedding W α ν , p ↪ L ∞ fails. We prove sharp logarithmic corrections using a correct Littlewood–Paley decomposition, avoiding the divergent semigroup representation used in earlier works.
Third, we develop a weighted theory for Muckenhoupt weights w ∈ A p , using the anisotropic Mihlin multiplier theorem of Wang [18].
Fourth, we study optimal constants and extremal functions via concentration-compactness [10].
Organisation. The paper is organised as follows. Section 2 sets up the anisotropic scaling structure, mixed fractional Sobolev spaces, fractional derivatives, the anisotropic Mihlin multiplier theorem, the anisotropic Hardy–Littlewood maximal operator, and the Littlewood–Paley decomposition. Section 3 establishes the general Gagliardo–Nirenberg hierarchy, with the corrected homogeneity condition, the first-order L ∞ and mixed-derivative corollaries, and sharpness of the exponents. Section 4 treats the critical case with sharp logarithmic corrections, including the critical embedding and the critical Landau inequality. Section 5 develops the weighted theory for Muckenhoupt weights. Section 6 is devoted to optimal constants and extremal functions, with the Euler–Lagrange equation and existence via concentration-compactness. Section 7 applies the theory to fractional parabolic PDEs, and Section 8 addresses neural network approximation. Section 9 summarises the main results, and Section 10 presents concluding remarks and open problems.

2. Preliminaries

2.1. Anisotropic Scaling Structure

Let α = ( α 1 , … , α k ) ∈ ( 0 , ∞ ) k . The anisotropic dilation group is
T λ α x = ( λ α 1 x 1 , … , λ α k x k ) , λ > 0 .
The anisotropic homogeneous dimension is D α = ∑ i = 1 k α i , and the effective dimension is d α = ∑ i = 1 k α i − 1 . The anisotropic distance is
ρ α ( x ) = ∑ i = 1 k | x i | 2 / α i 1 / 2 .
We use the notation 〈 β 〉 = ∑ i = 1 k α i β i for the anisotropic order of a multi-index β ∈ N 0 k .

2.2. Mixed Fractional Sobolev Spaces

Definition 1. 
For ν > 0 , 1 ≤ p < ∞ , and α ∈ ( 0 , ∞ ) k , the mixed fractional Sobolev space W α ν , p ( R k ) is the completion of C c ∞ ( R k ) under the norm
∥ f ∥ W α ν , p = ∥ f ∥ L p + | f | W ˙ α ν , p ,
where the homogeneous seminorm is
| f | W ˙ α ν , p = ∑ j = 1 k ∥ D j ν , α j f ∥ L p p 1 / p ,
and the directional fractional derivative D j ν , α j is defined in Definition 2.
Remark 1. 
The homogeneous seminorm (7) scales uniformly under T λ α :
| f λ | W ˙ α ν , p = λ ν − D α / p | f | W ˙ α ν , p , f λ ( x ) = f ( T λ α x ) .
This uniformity is essential for the homogeneity condition (3) and fails for the weighted sum ∑ j ∥ D j ν , α j f ∥ L p α j / ν when the α j are not all equal.
Lemma 1. 
W α ν , p ( R k ) is a Banach space, and C c ∞ ( R k ) is dense in it.
Proof. 
Completeness follows from the standard argument: a Cauchy sequence in W α ν , p is Cauchy in L p , so it has an L p limit f. The difference quotients converge in L p ( R k × R ) , establishing that f has finite seminorm. Density of C c ∞ follows from mollification; see [15] for the anisotropic case. □

2.3. Anisotropic Fractional Derivatives

Definition 2. 
The Riemann–Liouville fractional derivative in direction e i is
D i ν , α i f ( x ) = 1 Γ ( n − ν / α i ) ∂ n ∂ x i n ∫ − ∞ x i f ( x 1 , … , t , … , x k ) ( x i − t ) ν / α i − n + 1 d t ,
where n = ⌊ ν / α i ⌋ + 1 . The Fourier transform characterisation is
F [ D i ν , α i f ] ( ξ ) = ( i ξ i ) ν / α i f ^ ( ξ ) .
Remark 2. 
The fractional derivatives D i ν , α i are bounded operators from W α ν , p ( R k ) to L p ( R k ) . This follows from the Fourier multiplier characterisation (10) and the anisotropic Mihlin multiplier theorem (Theorem 1).
Theorem 1 
(Anisotropic Mihlin multiplier theorem). Let m ∈ C ∞ ( R k ∖ { 0 } ) satisfy the anisotropic Mihlin condition: there exists C N > 0 such that for all multi-indices γ with | γ | ≤ ⌊ d α / 2 ⌋ + 1 ,
| ∂ γ m ( ξ ) | ≤ C N | ξ | − 〈 γ 〉 , 〈 γ 〉 = ∑ j = 1 k α j γ j .
Then the multiplier operator T m f = F − 1 ( m f ^ ) extends to a bounded operator on L p ( R k ) for 1 < p < ∞ , with norm depending on C N , p, and α.
Proof. 
See [18] for the anisotropic Hardy space version, which implies the L p boundedness by interpolation and duality. The condition (11) is the anisotropic analogue of the classical Mihlin condition; the number of derivatives required, ⌊ d α / 2 ⌋ + 1 , reflects the homogeneous dimension d α of the anisotropic dilation group. □

2.4. The Anisotropic Hardy–Littlewood Maximal Operator

For f ∈ L loc 1 ( R k ) , define
M α f ( x ) = sup r > 0 1 | B α ( x , r ) | ∫ B α ( x , r ) | f ( y ) | d y ,
where the anisotropic ball is B α ( x , r ) = { y ∈ R k : | y i − x i | < r 1 / α i } .
Theorem 2 
(Anisotropic Hardy–Littlewood maximal inequality). For 1 < p ≤ ∞ , there exists C ( p , α ) > 0 such that
∥ M α f ∥ L p ≤ C ( p , α ) ∥ f ∥ L p .
Proof. 
The proof follows the standard Vitali covering argument. The anisotropic balls satisfy the doubling property | B α ( x , 3 r ) | = 3 d α | B α ( x , r ) | , which follows from the change of variables y = T 3 α x . The weak ( 1 , 1 ) estimate is
| { M α f > λ } | ≤ 3 d α λ ∥ f ∥ L 1 .
Marcinkiewicz interpolation then yields the strong ( p , p ) bound. A detailed proof is in [2]. □

2.5. Littlewood–Paley Decomposition

The following decomposition replaces the divergent semigroup representation f = ∫ 0 ∞ ( I − e t Δ α ) f d t t , whose Fourier multiplier ∫ 0 ∞ ( 1 − e − t Φ ( ξ ) ) d t t diverges logarithmically at t → ∞ for every ξ ≠ 0 .
Lemma 2 
(Littlewood–Paley representation). Let ϕ ∈ C c ∞ ( R ) satisfy ϕ ( t ) = 0 for t ≤ 1 / 2 , ϕ ( t ) = 1 for t ≥ 1 , and let ψ ( t ) = ϕ ( t ) − ϕ ( 2 t ) . Define ψ j ( ξ ) = ψ ( 2 − j Φ ( ξ ) ) , where Φ ( ξ ) = ∑ i = 1 k | ξ i | 2 / α i . Then for f ∈ W α ν , p ( R k ) ,
f = ∑ j ∈ Z ψ j ( D α ) f
in the sense of distributions, and the series converges in W α ν , p . Moreover, the following norm equivalence holds:
∥ f ∥ W α ν , p ≃ ∑ j ∈ Z 2 j ν p | ψ j ( D α ) f | p 1 / p L p .
Proof. 
The functions ψ j form a partition of unity: ∑ j ∈ Z ψ ( 2 − j λ ) = 1 for λ > 0 . The convergence and the norm equivalence (16) follow from the standard Littlewood–Paley theory in anisotropic Besov spaces; see [15]. □
Lemma 3 
(Anisotropic Bernstein inequality). For every j ∈ Z , every multi-index β ∈ N 0 k , and every 1 ≤ p ≤ q ≤ ∞ ,
∥ ψ j ( D α ) D β f ∥ L q ≤ C 2 j ( 〈 β 〉 + D α ( 1 / p − 1 / q ) ) ∥ ψ j ( D α ) f ∥ L p ,
where 〈 β 〉 = ∑ i α i β i .
Proof. 
The operator ψ j ( D α ) has symbol ψ ( 2 − j Φ ( ξ ) ) supported in the anisotropic annulus 2 j ≤ Φ ( ξ ) ≤ 2 j + 1 . On this annulus, | ξ | ≃ 2 j / 2 in the anisotropic sense. The derivative D β contributes a factor of order 2 j 〈 β 〉 / 2 in the anisotropic symbol. The L p → L q bound follows from Young’s inequality and the fact that the kernel of ψ j ( D α ) has L 1 norm bounded uniformly in j, see [15]. □

3. Generalised Gagliardo–Nirenberg Inequalities

3.1. The Main Interpolation Inequality

Theorem 3 
(Anisotropic Gagliardo–Nirenberg inequality). Let m ∈ N , ν > 0 , 1 < p , r < ∞ , and 1 ≤ q ≤ ∞ . Let α ∈ ( 0 , ∞ ) k and suppose
θ ∈ α i m ν , 1 , 1 q = α i m D α + 1 − θ r + θ 1 p − ν D α ,
with the convention that 1 / ∞ = 0 when q = ∞ . Then for every f ∈ W α ν , p ( R k ) ∩ L r ( R k ) ,
∥ D i m f ∥ L q ≤ C ( ν , p , r , α ) ∥ f ∥ L r 1 − θ | f | W ˙ α ν , p θ .
The constant is explicit:
C = C HL ν ν − α i m / θ 1 − θ ∑ j = 1 k α j − 1 θ C Bern ( ν , p , α ) θ ,
where C HL is the constant from Theorem 2 and C Bern is the Bernstein constant from Lemma 3.
Proof. 
We prove the result via the Littlewood–Paley decomposition (Lemma 2) and the anisotropic Bernstein inequality (Lemma 3). Write
f = ∑ j ∈ Z Δ j α f , Δ j α : = ψ j ( D α ) .
Then
∥ D i m f ∥ L q ≤ ∑ j ∈ Z ∥ D i m Δ j α f ∥ L q .
By the Bernstein inequality (17) with β = m e i (so that 〈 β 〉 = α i m ),
∥ D i m Δ j α f ∥ L q ≤ C 2 j ( α i m + D α ( 1 / p − 1 / q ) ) ∥ Δ j α f ∥ L p .
We also need the interpolation inequality for the LP pieces:
∥ Δ j α f ∥ L p ≤ ∥ Δ j α f ∥ L r 1 − η ∥ Δ j α f ∥ L ∞ η
for some η ∈ [ 0 , 1 ] . Alternatively, we use the equivalent form
∥ Δ j α f ∥ L q ≤ C 2 j D α [ ( 1 − θ ) ( 1 / r − 1 / q ) + θ ( 1 / p − 1 / q ) ] ∥ Δ j α f ∥ L r 1 − θ ∥ Δ j α f ∥ L p θ ,
which follows from interpolating between L r and L p at frequency 2 j .
Combining (23) and (25), we obtain
∥ D i m Δ j α f ∥ L q ≤ C 2 j θ ν ∥ Δ j α f ∥ L r 1 − θ ∥ Δ j α f ∥ L p θ ,
where we used the identity
α i m + D α [ ( 1 − θ ) ( 1 / r − 1 / q ) + θ ( 1 / p − 1 / q ) ] = θ ν ,
which follows from the homogeneity condition (3).
Now we sum over j. By Hölder’s inequality in the sequence space,
∑ j ∈ Z 2 j θ ν ∥ Δ j α f ∥ L r 1 − θ ∥ Δ j α f ∥ L p θ ≤ ∑ j ∈ Z ∥ Δ j α f ∥ L r 1 − θ ∑ j ∈ Z 2 j ν p ∥ Δ j α f ∥ L p p θ / p ≤ C ∥ f ∥ L r 1 − θ | f | W ˙ α ν , p θ ,
where we used the Littlewood–Paley characterisation of L r (for 1 < r < ∞ ) and of the homogeneous Sobolev seminorm (7). This completes the proof.
The condition θ ≥ α i m / ν ensures that the exponent θ ν − α i m ≥ 0 in the Bernstein step, so that the sum over j converges. The condition θ ≤ 1 is natural for interpolation. □
Remark 3 
(On the direction dependence of q). The target exponent q in (19) depends on the direction i through the term α i m / D α in (3). This is a genuine feature of the anisotropic setting: differentiating in a direction with larger weight α i requires a larger exponent q to balance the scaling. When α 1 = … = α k = 1 (isotropic case), we recover the classical condition
1 q = m k + 1 − θ r + θ 1 p − ν k .

3.2. Consequences and Special Cases

Corollary 1 
(First-order L ∞ inequality). Let ν ∈ ( 1 , 2 ) , 1 < p < ∞ , and α ∈ ( 0 , ∞ ) k such that ν p > D α . Set
θ = α i ν − D α / p .
Then for every f ∈ W α ν , p ( R k ) and every direction i,
∥ D i 1 f ∥ L ∞ ≤ C ( ν , p , α ) ∥ f ∥ L ∞ 1 − θ | f | W ˙ α ν , p θ .
Proof. 
Take m = 1 , q = ∞ , r = ∞ in Theorem 3. The homogeneity condition becomes
0 = α i D α + θ 1 p − ν D α ,
which gives (30). The condition θ ≤ 1 is equivalent to ν − D α / p ≥ α i , which holds when ν p > D α and α i ≤ ν − D α / p . The subcritical embedding W α ν , p ↪ L ∞ holds when ν p > D α (see Theorem 4). □
Remark 4. 
In the isotropic case α i = 1 , k = 1 , p = ∞ , Corollary 1 gives θ = 1 / ν , recovering the classical fractional Landau exponent. For finite p, the exponent is larger: θ = 1 / ( ν − 1 / p ) > 1 / ν , reflecting the stronger norm on the right-hand side.
Corollary 2 
(Mixed derivatives). Let β = ( β 1 , … , β k ) with β i ∈ N 0 and 〈 β 〉 = ∑ i α i β i < ν . Define θ : = 〈 β 〉 / ν . Assume the exponents satisfy
1 q = 〈 β 〉 D α + 1 − θ r + θ 1 p − ν D α .
Then for every f ∈ W α ν , p ( R k ) ∩ L r ( R k ) ,
∥ D β f ∥ L q ≤ C ∥ f ∥ L r 1 − θ | f | W ˙ α ν , p θ .
Proof. 
The proof proceeds by induction on | β | . For a single directional derivative, this is Theorem 3. For the induction step, apply Theorem 3 to g = D β ′ f with β ′ = β − e i , and use the fact that D i ν , α i commutes with D β ′ . □

3.3. Sharpness of the Exponents

The exponents in Theorem 3 are sharp in the sense of scaling invariance. Under the anisotropic scaling f λ ( x ) = f ( T λ α x ) , we have
∥ D i m f λ ∥ L q = λ α i m − D α / q ∥ D i m f ∥ L q ,
∥ f λ ∥ L r = λ − D α / r ∥ f ∥ L r ,
and
| f λ | W ˙ α ν , p = λ ν − D α / p | f | W ˙ α ν , p .
The homogeneity of both sides of (19) requires
α i m − D α q = − D α r ( 1 − θ ) + θ ν − D α p ,
which is precisely the condition (3). This confirms that the exponent relation is not only sufficient but also necessary for the inequality to be scale-invariant.

4. The Critical Case: Logarithmic Corrections

4.1. The Critical Embedding

Theorem 4 
(Critical embedding). Let ν p = D α . Then
W α ν , p ( R k ) ↪ L q ( R k ) for all p ≤ q < ∞ ,
but W α ν , p ( R k ) ¬ ↪ L ∞ ( R k ) . Moreover, for every f ∈ W α ν , p ( R k ) ,
∥ f ∥ L ∞ ≤ C ∥ f ∥ W α ν , p 1 + log 1 / p 1 + ∥ f ∥ W α ν , p ∥ f ∥ L p 1 / p .
Proof. 
We use the Littlewood–Paley decomposition from Lemma 2 and avoid the divergent semigroup representation. Write
f = ∑ j ≤ J Δ j α f + ∑ j > J Δ j α f = : f low + f high .
Low-frequency estimate. By the Bernstein inequality (17) with q = ∞ and p as given,
∥ Δ j α f ∥ L ∞ ≤ C 2 j D α / p ∥ Δ j α f ∥ L p = C 2 j ν ∥ Δ j α f ∥ L p ,
where we used D α / p = ν in the critical case. Summing over j ≤ J ,
∥ f low ∥ L ∞ ≤ C ∑ j ≤ J 2 j ν ∥ Δ j α f ∥ L p ≤ C 2 J ν ∥ f ∥ L p ,
where the last inequality uses the fact that ∥ Δ j α f ∥ L p ≤ C ∥ f ∥ L p for all j (by the L p boundedness of Δ j α ) and the geometric sum ∑ j ≤ J 2 j ν ≤ C 2 J ν .
High-frequency estimate. For j > J , we use the decay of the LP pieces:
∥ Δ j α f ∥ L p ≤ C 2 − j ν ∥ f ∥ W ˙ α ν , p ,
which follows from the norm equivalence (16). Combining with the Bernstein inequality,
∥ Δ j α f ∥ L ∞ ≤ C 2 j ν ∥ Δ j α f ∥ L p ≤ C ∥ f ∥ W ˙ α ν , p .
This bound is uniform in j, so the sum over j > J diverges linearly. To obtain a convergent estimate, we use the refined decay: for j > J with 2 J ν ≥ ∥ f ∥ W ˙ α ν , p / ∥ f ∥ L p ,
∥ Δ j α f ∥ L p ≤ C 2 − j ν ∥ f ∥ W ˙ α ν , p .
Summing over j ∈ ( J , J + N ] for N ≃ log 2 ( ∥ f ∥ W ˙ α ν , p / ( 2 J ν ∥ f ∥ L p ) ) ,
∑ j = J + 1 J + N Δ j α f L ∞ ≤ C ∑ j = J + 1 J + N ∥ Δ j α f ∥ L ∞ ≤ C ∑ j = J + 1 J + N 2 j ν · 2 − j ν ∥ f ∥ W ˙ α ν , p = C N ∥ f ∥ W ˙ α ν , p .
For j > J + N , the LP pieces decay geometrically, and the sum is bounded by C 2 − ( J + N ) ν ∥ f ∥ W ˙ α ν , p .
Optimisation. We now choose the truncation parameter J to balance the low- and high-frequency contributions. By the Littlewood–Paley norm equivalence (16), the sequence
a j : = 2 j ν ∥ Δ j α f ∥ L p
belongs to ℓ p ( Z ) with
∥ ( a j ) ∥ ℓ p ≤ C ∥ f ∥ W ˙ α ν , p .
For the high-frequency part, we use the elementary estimate for sequences in ℓ p . To derive it, fix a block length N ≥ 1 and split the tail according to
{ j > J } = { J < j ≤ J + N } ∪ { j > J + N } .
On the finite block, the discrete Hölder inequality with conjugate exponents p and p ′ gives
∑ J < j ≤ J + N a j ≤ N 1 / p ′ ∑ J < j ≤ J + N a j p 1 / p ≤ N 1 / p ′ ∥ ( a j ) ∥ ℓ p .
The remaining tail is controlled by the ℓ p summability of ( a j ) , which ensures that its contribution is bounded by a constant multiple of ∥ ( a j ) ∥ ℓ p provided the block length is chosen large enough to capture the geometric decay of the LP pieces. Balancing the two contributions by taking N ≃ 1 + log 2 ( 2 J ) yields
∑ j > J a j ≤ C ∥ ( a j ) ∥ ℓ p 1 + log 2 ( 2 J ) 1 / p ′ .
Consequently,
∥ f high ∥ L ∞ ≤ C ∑ j > J 2 j ν ∥ Δ j α f ∥ L p ≤ C ∥ f ∥ W ˙ α ν , p 1 + log 2 ( 2 J ) 1 / p ′ .
Combining this with the low-frequency estimate (43), we obtain
∥ f ∥ L ∞ ≤ C 2 J ν ∥ f ∥ L p + ∥ f ∥ W ˙ α ν , p ( 1 + log 2 ( 2 J ) ) 1 / p ′ .
We now choose J such that
2 J ν ∥ f ∥ L p ≃ ∥ f ∥ W ˙ α ν , p , i . e . J = 1 ν log 2 ∥ f ∥ W ˙ α ν , p ∥ f ∥ L p .
With this choice, the two terms are comparable, and since 1 / p ′ = 1 − 1 / p ≤ 1 , we deduce
∥ f ∥ L ∞ ≤ C ∥ f ∥ W α ν , p 1 + log 1 / p 1 + ∥ f ∥ W α ν , p ∥ f ∥ L p 1 / p ,
which is (40).
Failure of the L ∞ embedding. The sequence f j ( x ) = ϕ ( 2 j T 2 − j α x ) satisfies ∥ f j ∥ L ∞ = 1 but ∥ f j ∥ W α ν , p → ∞ ; see [15]. This shows that the embedding into L ∞ fails in the critical case. □
Remark 5. 
The logarithmic correction in (40) is sharp: it cannot be replaced by a power of the logarithm with exponent smaller than 1 / p . This follows from the standard example of a function whose LP pieces are all of comparable size; see [5].

4.2. Critical Landau Inequality

Theorem 5 
(Critical mixed fractional Landau inequality). Let ν ∈ ( 1 , 2 ) , 1 < p < ∞ , and α ∈ ( 0 , ∞ ) k such that ν p = D α . Then for every f ∈ W α ν , p ( R k ) and every direction i,
∥ D i 1 f ∥ L ∞ ≤ C ( ν , p , α ) ∥ f ∥ L ∞ 1 − θ | f | W ˙ α ν , p θ 1 + log 1 / p 1 + | f | W ˙ α ν , p ∥ f ∥ L ∞ ,
where θ = α i / ν in the critical case (obtained formally from (30) by setting D α / p = ν and taking the limit).
Proof. 
We follow the fractional Taylor expansion argument, avoiding the circularity of earlier versions. Fix x ∈ R k and h > 0 . The fractional Taylor formula
f ( x + h e i ) = f ( x ) + h D i 1 f ( x ) + 1 Γ ( ν ) ∫ 0 h ( h − t ) ν − 1 D i ν , α i f ( x + t e i ) d t
is valid for f ∈ C 1 ( R k ) ∩ W α ν , p ( R k ) ; for general f ∈ W α ν , p , it follows by density and the absolute continuity of fractional integrals, see [8].
Rearranging (57) and taking absolute values,
h | D i 1 f ( x ) | ≤ 2 ∥ f ∥ L ∞ + 1 Γ ( ν ) ∫ 0 h ( h − t ) ν − 1 | D i ν , α i f ( x + t e i ) | d t .
Estimating the integral using Hölder’s inequality with exponents p and p ′ ,
∫ 0 h ( h − t ) ν − 1 | D i ν , α i f ( x + t e i ) | d t ≤ ∫ 0 h ( h − t ) p ′ ( ν − 1 ) d t 1 / p ′ ∫ 0 h | D i ν , α i f ( x + t e i ) | p d t 1 / p .
The first factor is C h ν − 1 + 1 / p ′ , and the second is h 1 / p M i ( p ) ( D i ν , α i f ) ( x ) , where M i ( p ) is the one-dimensional p-maximal function in direction i. Thus
| D i 1 f ( x ) | ≤ 2 ∥ f ∥ L ∞ h + C h ν − 1 Γ ( ν ) ( ν − 1 / p ) 1 / p ′ M i ( p ) ( D i ν , α i f ) ( x ) .
Optimising over h > 0 gives
| D i 1 f ( x ) | ≤ C ∥ f ∥ L ∞ 1 − 1 / ν ( M i ( p ) ( D i ν , α i f ) ( x ) ) 1 / ν .
Relating the one-dimensional maximal function to the anisotropic maximal operator (Theorem 2), we obtain
∥ D i 1 f ∥ L ∞ ≤ C ∥ f ∥ L ∞ 1 − 1 / ν ∥ D i ν , α i f ∥ L ∞ 1 / ν .
Key step (avoiding circularity). Instead of invoking the critical embedding for D i ν , α i f (which would be circular, since the critical embedding fails for W α ν , p ), we apply the Littlewood–Paley estimate directly to g = D i ν , α i f . Note that g ∈ L p but not necessarily in W α ν , p ; however, the LP decomposition of g is still valid, and we have the norm equivalence
∥ g ∥ L ∞ ≤ C ∥ g ∥ L p 1 − η ∥ g ∥ B ˙ p , ∞ ν η
for some η ∈ ( 0 , 1 ) , where B ˙ p , ∞ ν is the anisotropic Besov space. In the critical case, B ˙ p , ∞ ν ↪ L ∞ with a logarithmic correction:
∥ g ∥ L ∞ ≤ C ∥ g ∥ B ˙ p , ∞ ν 1 + log 1 / p 1 + ∥ g ∥ B ˙ p , ∞ ν ∥ g ∥ L p 1 / p .
Since ∥ g ∥ B ˙ p , ∞ ν ≤ C ∥ g ∥ W α ν , p ≤ C | f | W ˙ α ν , p and ∥ g ∥ L p = ∥ D i ν , α i f ∥ L p ≤ C | f | W ˙ α ν , p , we obtain
∥ D i ν , α i f ∥ L ∞ ≤ C | f | W ˙ α ν , p 1 + log 1 / p 1 + | f | W ˙ α ν , p ∥ f ∥ L ∞ 1 / p .
Substituting into (62) and using θ = α i / ν (which follows from the homogeneity condition in the critical case), we obtain (56). □
Remark 6. 
The key difference from the earlier version is that we donotapply the critical embedding to D i ν , α i f (which would require D i ν , α i f ∈ W α ν , p , a non-trivial assumption). Instead, we use the Besov-space logarithmic embedding (64), which is valid for all g ∈ B ˙ p , ∞ ν and does not require g ∈ W α ν , p .

4.3. Sharpness of the Logarithmic Correction

Theorem 6 
(Optimality of the logarithmic correction). Let ν ∈ ( 1 , 2 ) , 1 < p < ∞ , and α ∈ ( 0 , ∞ ) k such that ν p = D α . Then there exists a sequence { f n } n ∈ N ⊂ W α ν , p ( R k ) such that
∥ D i 1 f n ∥ L ∞ ∥ f n ∥ L ∞ 1 − θ | f n | W ˙ α ν , p θ ≳ log 1 / p n , θ = α i ν .
Proof. 
Let ϕ ∈ C c ∞ ( R ) be supported in [ 0 , 1 ] with ϕ ( t ) = 1 on [ 1 / 4 , 3 / 4 ] . For each n ∈ N , define
f n ( x ) = ϕ ( ρ α ( x ) ) ρ α ( x ) ψ n ( ρ α ( x ) ) ,
where ψ n ∈ C ∞ ( [ 0 , ∞ ) ) satisfies
ψ n ( t ) = n , 0 ≤ t ≤ e − 2 n , log ( 1 / t ) , e − n ≤ t ≤ 1 ,
with a smooth monotone interpolation on ( e − 2 n , e − n ) .
We establish three estimates. First, ∥ f n ∥ L ∞ ≃ 1 because ρ α ψ n ( ρ α ) ≲ 1 and for ρ α ∈ [ 1 / 2 , 1 ] , ρ α ψ n ( ρ α ) ≳ 1 .
Second, for the derivative, in the region ρ α ∈ [ e − n , 1 / 2 ] ,
D i 1 f n ( x ) = ( log ( 1 / ρ α ) − 1 ) ∂ i ρ α .
Since | ∂ i ρ α | ≳ 1 for ρ α bounded away from 0, we have ∥ D i 1 f n ∥ L ∞ ≳ n .
Third, for the fractional norm, the function g ( x ) = ϕ ( ρ α ( x ) ) ρ α ( x ) log ( 1 / ρ α ( x ) ) belongs to the critical Besov space B p , p D α / p ≃ W α ν , p , see [15]. The truncation ψ n differs from log ( 1 / ρ ) only on ρ ≤ e − n , where the contribution to the Gagliardo seminorm is
∫ 0 e − n ρ D α − 1 ( ρ n ) p d ρ = n p p + D α e − n ( p + D α ) → 0 .
Thus | f n | W ˙ α ν , p ≃ 1 .
Combining these estimates yields the ratio ≳ n ≳ log 1 / p n . □

5. Weighted Inequalities

5.1. Muckenhoupt Weights and Weighted Spaces

Definition 3 
(Muckenhoupt A p weights). Let 1 < p < ∞ . A non-negative locally integrable function w : R k → [ 0 , ∞ ) belongs to A p (relative to anisotropic balls) if
[ w ] A p : = sup Q 1 | Q | ∫ Q w ( x ) d x 1 | Q | ∫ Q w ( x ) − 1 / ( p − 1 ) d x p − 1 < ∞ ,
where Q ranges over anisotropic balls B α ( x , r ) .
Proposition 1 
(Doubling property). Let w ∈ A p . Then there exists C d > 0 such that for every anisotropic ball B α ( x , r ) ,
w ( B α ( x , 2 r ) ) ≤ C d w ( B α ( x , r ) ) .
The constant depends on [ w ] A p and α.
Proof. 
This follows from the definition of the A p characteristic; see [6]. The doubling constant satisfies C d ≤ C k , p , α [ w ] A p . □
Definition 4 
(Weighted spaces). For w ∈ A p , define L p ( w ) by ∥ f ∥ L p ( w ) = ( ∫ | f | p w ) 1 / p , and W α ν , p ( w ) as the completion of C c ∞ under
∥ f ∥ W α ν , p ( w ) = ∥ f ∥ L p ( w ) + | f | W ˙ α ν , p ( w ) ,
where
| f | W ˙ α ν , p ( w ) = ∑ j = 1 k ∥ D j ν , α j f ∥ L p ( w ) p 1 / p .

5.2. Weighted Maximal Function Estimates

Theorem 7 
(Weighted anisotropic maximal inequality). Let w ∈ A p with 1 < p < ∞ . Then
∥ M α f ∥ L p ( w ) ≤ C ( p , α , [ w ] A p ) ∥ f ∥ L p ( w ) .
Proof. 
The proof follows the standard structure adapted to the measure w ( x ) d x . Since w ∈ A p , it is doubling (Proposition 1), so the Vitali covering lemma holds with respect to w. The weak ( 1 , 1 ) estimate becomes
w ( { M α f > λ } ) ≤ C ( [ w ] A p , α ) λ ∥ f ∥ L 1 ( w ) .
The strong ( p , p ) bound follows from Marcinkiewicz interpolation. See [6] for details. □

5.3. Weighted Gagliardo–Nirenberg Inequalities

Theorem 8 
(Weighted Gagliardo–Nirenberg inequality). Let w ∈ A p , 1 < p < ∞ , ν > 0 , m ∈ N . Suppose
θ ∈ α i m ν , 1 , 1 q = α i m D α + 1 − θ r + θ 1 p − ν D α .
Then for every f ∈ W α ν , p ( w ) ∩ L r ( w ) ,
∥ D i m f ∥ L q ( w ) ≤ C ∥ f ∥ L r ( w ) 1 − θ | f | W ˙ α ν , p ( w ) θ .
The constant depends on p , r , ν , α , [ w ] A p .
Proof. 
The proof follows the same strategy as Theorem 3, with all L p norms replaced by weighted L p ( w ) norms. The key additional ingredient is the weighted Littlewood–Paley estimate:
∥ f ∥ W α ν , p ( w ) ≃ ∑ j ∈ Z 2 j ν p | ψ j ( D α ) f | p 1 / p L p ( w ) ,
which follows from the weighted Mihlin multiplier theorem [18] and the fact that w ∈ A p implies the weighted Hardy–Littlewood maximal inequality. The remainder of the proof is identical to the unweighted case, with constants depending on [ w ] A p . □
Corollary 3 
(Weighted Landau inequality). Let w ∈ A p , ν ∈ ( 1 , 2 ) , 1 < p < ∞ , and ν p > D α . Set θ = α i / ( ν − D α / p ) . Then
∥ D i 1 f ∥ L ∞ ≤ C ∥ f ∥ L ∞ 1 − θ | f | W ˙ α ν , p ( w ) θ .
Proof. 
Since w ∈ A p is locally integrable and strictly positive a.e., the weighted and unweighted L ∞ -norms coincide: ∥ f ∥ L ∞ ( w ) = ∥ f ∥ L ∞ . The result follows from Theorem 8 with m = 1 , q = ∞ , r = ∞ . □

6. Optimal Constants and Extremal Functions

6.1. Variational Formulation

Consider the single-term Landau inequality
∥ D i 1 f ∥ L ∞ ≤ C opt ∥ f ∥ L ∞ 1 − θ ∥ D i ν , α i f ∥ L p θ ,
with θ = α i / ( ν − D α / p ) . The optimal constant is
C opt : = inf f ∈ W α ν , p f ¬ ≡ 0 , D i 1 f ¬ ≡ 0 ∥ D i 1 f ∥ L ∞ ∥ f ∥ L ∞ 1 − θ ∥ D i ν , α i f ∥ L p θ .
By homogeneity, impose ∥ f ∥ L ∞ = ∥ D i 1 f ∥ L ∞ = 1 . The variational problem becomes
Λ opt : = inf ∥ D i ν , α i f ∥ L p : f ∈ W α ν , p , ∥ f ∥ L ∞ = 1 , ∥ D i 1 f ∥ L ∞ = 1 ,
with C opt = Λ opt − 1 / θ .

6.2. The Euler–Lagrange Equation

For a smooth extremal f, the Fréchet derivative of the energy functional J ( f ) = ∥ D i ν , α i f ∥ L p p is
J ′ ( f ) = − p ( − Δ α , i ν ) | D i ν , α i f | p − 2 D i ν , α i f .
The subdifferential of the L ∞ -norm at g is
∂ ∥ g ∥ L ∞ = μ ∈ M ( R k ) : supp μ ⊂ { | g | = ∥ g ∥ L ∞ } , ∫ sgn ( g ) d μ = ∥ μ ∥ M .
The Lagrange multiplier rule gives
− ( − Δ α , i ν ) | D i ν , α i f | p − 2 D i ν , α i f = ∑ m = 1 M λ m δ x m ,
where x m ∈ S f ∪ S D i 1 f and M ≥ 2 . For p = 2 ,
− ( − Δ α , i ν ) f = ∑ m = 1 M λ m δ x m .

6.3. Existence of Extremal Functions

Theorem 9 
(Existence of extremals). Let ν ∈ ( 1 , 2 ) , 1 < p < ∞ , and ν p > D α . Then the variational problem admits a solution f ∈ W α ν , p ( R k ) satisfying ∥ f ∥ L ∞ = 1 , ∥ D i 1 f ∥ L ∞ = 1 , and achieving Λ opt . Moreover, f is smooth away from a finite set S and satisfies the Euler–Lagrange equation.
Proof. 
The proof uses the concentration-compactness principle [10]. Let { f n } be a minimising sequence. The anisotropic Gagliardo–Nirenberg inequality (Theorem 3) gives a uniform bound ∥ f n ∥ W α ν , p ≤ C . After passing to a subsequence, f n ⇀ f weakly in W α ν , p .
Define concentration functions
Q n ( t ) = sup x ∈ R k ∫ B α ( x , t ) | f n ( y ) | p d y .
By the concentration-compactness principle, one of three alternatives holds: vanishing, dichotomy, or tightness. Vanishing would imply ∥ f n ∥ L ∞ → 0 , contradicting ∥ f n ∥ L ∞ = 1 . Dichotomy would yield a strict subadditivity contradiction to minimality. Thus tightness holds.
By the anisotropic Rellich–Kondrachov theorem, W α ν , p ( B α ( 0 , R ) ) ↪ L p ( B α ( 0 , R ) ) compactly. Combining tightness with compactness gives strong convergence f n → f in L p ( R k ) . Since ∥ f n ∥ L ∞ = 1 uniformly, ∥ f ∥ L ∞ = 1 .
Weak lower semicontinuity gives ∥ D i ν , α i f ∥ L p ≤ lim inf ∥ D i ν , α i f n ∥ L p = Λ opt , so f is a minimiser. The Euler–Lagrange equation follows from the Lagrange multiplier rule. Regularity away from the contact set follows from the Schauder estimates for the fractional Laplacian. □

6.4. Explicit Sharp Constants in Special Cases

Proposition 2 
(One-dimensional classical case). For k = 1 , α 1 = 1 , ν = 2 , p = ∞ , the optimal constant is C opt = 2 , and the extremals are f ( x ) = A sin ( ω x + φ ) .
Proposition 3 
(Fractional one-dimensional case). For k = 1 , α 1 = 1 , ν ∈ ( 1 , 2 ) , p = ∞ ,
C ( ν ) = 2 1 − 1 / ν Γ ( ν ) ν ν − 1 1 − 1 / ν 1 ( ν − 1 ) 1 / ν .

7. Applications to Fractional Parabolic PDEs

7.1. The Model Equation

Consider
∂ t u + ( − Δ α ) ν / 2 u = F ( u , ∇ u ) , u ( 0 , x ) = u 0 ( x ) .
The operator ( − Δ α ) ν / 2 has Fourier symbol ( ∑ i | ξ i | 2 / α i ) ν / 2 .

7.2. Local Existence and Uniqueness

Theorem 10 
(Local existence and uniqueness). Assume F is locally Lipschitz and satisfies
| F ( s , z ) | ≤ C ( 1 + | s | q + | z | r )
with q , r ≥ 1 , q , r < p * : = D α p / ( D α − ν p ) . Then for every u 0 ∈ W α ν , p , there exists T > 0 and a unique solution
u ∈ C ( [ 0 , T ] ; W α ν , p ) ∩ L p ( 0 , T ; W α ν , p ) .
Proof. 
Define Φ ( u ) ( t ) = e − t ( − Δ α ) ν / 2 u 0 + ∫ 0 t e − ( t − s ) ( − Δ α ) ν / 2 F ( u ( s ) , ∇ u ( s ) ) d s . The semigroup estimates
∥ e − t ( − Δ α ) ν / 2 g ∥ W α ν , p ≤ C ∥ g ∥ W α ν , p , ∥ e − t ( − Δ α ) ν / 2 g ∥ W α ν , p ≤ C t − ν / 2 ∥ g ∥ L p
follow from the heat kernel bounds; see [9]. Using the growth condition on F and the anisotropic Sobolev embedding, Φ maps X T = C ( [ 0 , T ] ; W α ν , p ) ∩ L p ( 0 , T ; W α ν , p ) into itself and is a contraction on a sufficiently small time interval. The contraction mapping principle gives existence and uniqueness. □

7.3. A Priori Estimates

Theorem 11 
(Gradient bound for parabolic solutions). Let u be a smooth solution of (90) on [ 0 , T ] . Assume F satisfies the growth condition and ∥ u ( t ) ∥ L ∞ is bounded. Then
∥ ∇ u ( t , · ) ∥ L ∞ ≤ C ( t ) 1 + ∥ u ( t , · ) ∥ L ∞ + | u 0 | W ˙ α ν , p ,
where C ( t ) < ∞ for t > 0 .
Proof. 
Apply the Landau inequality (Corollary 1) to f = u ( t , · ) :
∥ ∇ u ( t ) ∥ L ∞ ≤ C ∥ u ( t ) ∥ L ∞ 1 − θ | u ( t ) | W ˙ α ν , p θ , θ = α i ν − D α / p .
It remains to control | u ( t ) | W ˙ α ν , p . Multiply the equation by | u | p − 2 u and integrate:
1 p d d t ∥ u ∥ L p p + ∑ j ∥ D j ν , α j u ∥ L p p = ∫ F ( u , ∇ u ) | u | p − 2 u d x .
Using the growth condition and the anisotropic Sobolev embedding,
∫ F ( u , ∇ u ) | u | p − 2 u d x ≤ C ( 1 + ∥ u ∥ W α ν , p p − 1 + q + r ) .
This gives a differential inequality for ∥ u ∥ L p p , which implies boundedness. The semigroup representation then yields ∥ D j ν , α j u ( t ) ∥ L p ≤ C ( t ) . □

7.4. Blow-up Criteria

Theorem 12 
(Blow-up criterion). Let u be a solution of
∂ t u + ( − Δ α ) ν / 2 u = | u | γ u , γ > 0 .
If the maximal existence time T * is finite, then
∫ 0 T * ∥ u ( t ) ∥ L ∞ γ d t = ∞ .
Proof. 
Suppose ∫ 0 T * ∥ u ( t ) ∥ L ∞ γ d t < ∞ . The energy estimate gives
1 p d d t ∥ u ∥ L p p + ∑ j ∥ D j ν , α j u ∥ L p p ≤ ∥ u ∥ L ∞ γ ∥ u ∥ L p p .
By Gronwall, ∥ u ( t ) ∥ L p ≤ C on [ 0 , T * ) . The semigroup representation then yields ∥ D j ν , α j u ( t ) ∥ L p ≤ C , so ∥ u ( t ) ∥ W α ν , p remains bounded, allowing extension beyond T * . Contradiction. □

8. Neural Network Approximation

8.1. Approximation Rates

Theorem 13 
(Neural network approximation). Let f ∈ W α ν , p ( R k ) with ν p > D α . Then there exists a sequence of anisotropic neural networks f N with at most N parameters such that
∥ f − f N ∥ L ∞ ≤ C ( ν , p , α ) N − ν / d α ∥ f ∥ W α ν , p .
Proof. 
The proof uses the anisotropic wavelet basis from [15] and the universal approximation theorem for ReLU networks [17]. The wavelets ψ j , m are tensor products:
ψ j , m ( x ) = 2 j d α / 2 ∏ i = 1 k ψ ( 2 j α i x i − m i ) ,
with support in ∏ i [ 2 − j α i m i , 2 − j α i ( m i + C ψ ) ] . The wavelet coefficients satisfy
∥ f ∥ W α ν , p ≃ ∑ j 2 j ν p ∑ m | c j , m | p 1 / p .
For J ∈ N , the wavelet projection f J = ∑ j ≤ J ∑ m c j , m ψ j , m satisfies
∥ f − f J ∥ L ∞ ≤ C 2 − J ν ∥ f ∥ W α ν , p .
Each wavelet is approximated by a neural network with error C 2 − j ν and width independent of j , m . The number of parameters is N J ≤ C ∑ j ≤ J 2 j d α ≤ C 2 J d α . Eliminating J gives the rate N − ν / d α . □

8.2. Lower Bounds

Theorem 14 
(Kolmogorov N-width lower bound). Let F = { f ∈ W α ν , p : ∥ f ∥ W α ν , p ≤ 1 } . Then
d N ( F , L ∞ ) ≳ N − ν / d α .
Proof. 
Let ϕ ∈ C c ∞ ( R ) be supported in [ 0 , 1 ] with ∥ ϕ ∥ L ∞ = 1 . For scale j ∈ N , define
ϕ j , m ( x ) = 2 − j d α ∏ i = 1 k ϕ ( 2 j α i x i − m i ) ,
for m i = 1 , … , 2 j α i . These functions are supported on disjoint anisotropic boxes. Normalising ψ j , m = ϕ j , m / ∥ ϕ j , m ∥ W α ν , p ∈ F , the functions are disjoint in support, so the L ∞ distance between any two is 2 − j d α . The number of such functions is 2 j d α . The standard volume argument [12] gives d N ( F , L ∞ ) ≳ 2 − j ν with N ≃ 2 j d α , hence d N ( F , L ∞ ) ≳ N − ν / d α . □

9. Results

We now summarise the main contributions of the paper in continuous form. The results are stated without proofs, which are given in the preceding sections. The central object of study is the anisotropic mixed fractional Sobolev space W α ν , p ( R k ) , endowed with the homogeneous seminorm
| f | W ˙ α ν , p = ∑ j = 1 k ∥ D j ν , α j f ∥ L p p 1 / p ,
whose scaling under the anisotropic dilation T λ α is uniform and equal to λ ν − D α / p . This uniform scaling is the key structural property that allows a consistent formulation of the Gagliardo–Nirenberg inequalities in the anisotropic setting.
The first main result is the anisotropic Gagliardo–Nirenberg inequality, Theorem 3. For m ∈ N , 1 < p , r < ∞ , 1 ≤ q ≤ ∞ , and α ∈ ( 0 , ∞ ) k , if
θ ∈ α i m ν , 1 , 1 q = α i m D α + 1 − θ r + θ 1 p − ν D α ,
then every f ∈ W α ν , p ( R k ) ∩ L r ( R k ) satisfies
∥ D i m f ∥ L q ≤ C ( ν , p , r , α ) ∥ f ∥ L r 1 − θ | f | W ˙ α ν , p θ .
The constant is explicit and depends on the anisotropic Hardy–Littlewood constant, the Bernstein constant, and the effective dimension d α = ∑ i α i − 1 . The proof is based on the Littlewood–Paley decomposition and the anisotropic Bernstein inequality, and it replaces the divergent semigroup representation used in earlier formulations. The exponent relation is not only sufficient but also necessary for scaling invariance, as shown by the homogeneity analysis in Section 3.3. A distinctive feature of the anisotropic result is that the target exponent q depends on the direction i through the anisotropic order α i m , and the weighted sum in the seminorm is the natural quantity that preserves homogeneity.
The second main result concerns the critical regime ν p = D α . Theorem 4 establishes that
W α ν , p ( R k ) ↪ L q ( R k ) ( p ≤ q < ∞ ) ,
but
W α ν , p ( R k ) ↪ L ∞ ( R k ) .
The failure of the L ∞ embedding is compensated by a sharp logarithmic estimate:
∥ f ∥ L ∞ ≤ C ∥ f ∥ W α ν , p 1 + log 1 / p 1 + ∥ f ∥ W α ν , p ∥ f ∥ L p 1 / p .
This estimate is obtained from a rigorous Littlewood–Paley decomposition, avoiding the divergent integral ∫ 0 ∞ ( 1 − e − t Φ ( ξ ) ) d t / t . In the same critical regime, Theorem 5 gives the corresponding Landau inequality
∥ D i 1 f ∥ L ∞ ≤ C ( ν , p , α ) ∥ f ∥ L ∞ 1 − θ | f | W ˙ α ν , p θ 1 + log 1 / p 1 + | f | W ˙ α ν , p ∥ f ∥ L ∞ , θ = α i ν ,
where the logarithmic factor is unavoidable. The proof avoids the circularity of earlier versions by applying the logarithmic Besov embedding to g = D i ν , α i f rather than to f itself. Theorem 6 proves that the logarithmic correction is sharp: there exists a sequence { f n } ⊂ W α ν , p ( R k ) such that the quotient between the left-hand side and the subcritical part grows like log 1 / p n . This establishes the optimality of the correction and closes a gap in the previous literature.
The third main result is the weighted theory for Muckenhoupt weights w ∈ A p , developed in Section 5. The weighted anisotropic maximal inequality, Theorem 7, states that
∥ M α f ∥ L p ( w ) ≤ C ( p , α , [ w ] A p ) ∥ f ∥ L p ( w ) , 1 < p < ∞ .
The weighted Gagliardo–Nirenberg inequality, Theorem 8, extends the unweighted hierarchy to W α ν , p ( w ) :
∥ D i m f ∥ L q ( w ) ≤ C ∥ f ∥ L r ( w ) 1 − θ | f | W ˙ α ν , p ( w ) θ ,
with the same exponent relation as in the unweighted case. The constant depends polynomially on the Muckenhoupt characteristic [ w ] A p . The weighted Landau inequality follows as Corollary 3. These results provide a foundation for studying anisotropic fractional PDEs with singular or degenerate coefficients.
The fourth main result is the study of optimal constants and extremal functions in Section 6. The variational problem for the sharp constant in the single-term Landau inequality is formulated as
Λ opt = inf ∥ D i ν , α i f ∥ L p : f ∈ W α ν , p , ∥ f ∥ L ∞ = 1 , ∥ D i 1 f ∥ L ∞ = 1 ,
with C opt = Λ opt − 1 / θ . For smooth extremals, the Euler–Lagrange equation is
− ( − Δ α , i ν ) | D i ν , α i f | p − 2 D i ν , α i f = ∑ m = 1 M λ m δ x m ,
where the points x m lie in the contact sets and M ≥ 2 . For p = 2 , this reduces to a fractional Poisson equation with measure-valued source. Theorem 9 proves existence of extremals via concentration-compactness. Explicit constants are recovered in special cases: the classical Landau constant C = 2 , the fractional one-dimensional constant
C ( ν ) = 2 1 − 1 / ν Γ ( ν ) ν ν − 1 1 − 1 / ν 1 ( ν − 1 ) 1 / ν ,
and the limiting behaviours lim ν → 2 − C ( ν ) = 2 and C ( ν ) ≍ ( ν − 1 ) − 1 as ν → 1 + .
The fifth main result is the application to anisotropic fractional parabolic equations, developed in Section 7. For the model equation
∂ t u + ( − Δ α ) ν / 2 u = F ( u , ∇ u ) , u ( 0 , x ) = u 0 ( x ) ,
Theorem 10 establishes local existence and uniqueness under natural Lipschitz and growth conditions on F. Theorem 11 provides a priori gradient bounds by combining the Landau inequality with energy estimates. Theorem 12 gives a sharp blow-up criterion for power-type nonlinearities: if the maximal existence time T * is finite, then
∫ 0 T * ∥ u ( t ) ∥ L ∞ γ d t = ∞ .
Conversely, global existence follows if this integral is finite.
The sixth main result is the neural network approximation theory of Section 8. Theorem 13 proves that for f ∈ W α ν , p ( R k ) with ν p > D α , there exist anisotropic neural networks f N with at most N parameters such that
∥ f − f N ∥ L ∞ ≤ C ( ν , p , α ) N − ν / d α ∥ f ∥ W α ν , p .
Theorem 14 provides the matching Kolmogorov N-width lower bound
d N ( F , L ∞ ) ≳ N − ν / d α , F = { f ∈ W α ν , p : ∥ f ∥ W α ν , p ≤ 1 } ,
confirming that the rate is optimal. The effective dimension d α replaces the ambient dimension k, quantifying the dimensional reduction induced by anisotropy.
Taken together, these results provide a unified analytical framework linking fractional calculus, anisotropic function spaces, partial differential equations, and approximation theory. The effective dimension d α = ∑ i α i − 1 emerges as the key parameter governing the rates, constants, and logarithmic corrections throughout the theory. The endpoint cases p = 1 and p = ∞ , the explicit computation of C opt for general parameters, the extension to bounded domains and manifolds, and the optimal depth–width tradeoffs for neural networks remain open and are discussed in the final section.

10. Conclusions and Open Problems

We have developed a comprehensive theory of anisotropic mixed fractional Landau and Gagliardo–Nirenberg inequalities. The main contributions are as follows.
1.
A general Gagliardo–Nirenberg hierarchy for W α ν , p (Theorem 3), covering arbitrary integer orders and non-endpoint Lebesgue exponents, with the correct homogeneity condition (3) and explicit constants depending on d α .
2.
Sharp logarithmic corrections in the critical case ν p = D α (Theorems 4 and 5), with optimality proved via explicit extremal sequences (Theorem 6).
3.
A weighted theory for A p weights (Theorems 7 and 8), extending all estimates to W α ν , p ( w ) .
4.
Existence of extremals for the variational problem characterizing the optimal constant (Theorem 9), together with the derivation of the Euler–Lagrange equation (86).
5.
Applications to fractional parabolic PDEs and neural network approximation, including optimal rates N − ν / d α with matching lower bounds.
Several questions remain open and warrant further investigation.
1.
The endpoint cases p = 1 and p = ∞ , where the present theory does not apply and new techniques are required.
2.
The explicit computation of the optimal constant C opt for general parameters ( ν , p , α ) , which is equivalent to solving the nonlinear spectral problem associated with the Euler–Lagrange equation.
3.
The extension of the theory to bounded domains with Dirichlet or Neumann boundary conditions, where the interaction between fractional derivatives and the boundary introduces additional difficulties.
4.
The generalization to Riemannian and sub-Riemannian manifolds with anisotropic structures, which would require developing a suitable notion of anisotropic fractional Laplacian on manifolds.
5.
The development of numerical algorithms for solving the variational problem and computing optimal constants, based on the Euler–Lagrange equation and concentration-compactness methods.
6.
Optimal depth–width tradeoffs for neural networks, determining the precise dependence of the approximation error on both depth and width, beyond the parameter count N.
Addressing these open problems would further strengthen the theory and broaden its applicability across analysis, PDEs, and machine learning.

Acknowledgments

The authors would like to thank IPEN of the National Nuclear Energy Commission – São Paulo (CNEN/IPEN–SP) for the institutional and financial support during the postgraduate (postdoctoral) period.

Conflicts of Interest

The authors declare that they have no competing financial, professional, or personal interests that could have influenced the research presented in this paper.

References

  1. Landau, E. (1925). Die Ungleichungen für zweimal differentiierbare Funktionen (Vol. 6). AF Høst &amp; søn.
  2. Stein, E. M. (1970). Singular integrals and differentiability properties of functions (No. 30). Princeton university press.
  3. Muckenhoupt, B. Weighted norm inequalities for the Hardy maximal function. Trans. Am. Math. Soc. 1972, 165, 207–226. [Google Scholar] [CrossRef]
  4. Adams, R.; Fournier, J. F. (1975). Sobolev spaces, acad. Press, New York, 19(5). 5.
  5. Brezis, H.; Gallouet, T. Nonlinear Schrödinger evolution equations. Nonlinear Anal. Theory Methods Appl. 1980, 4(4), 677–681. [Google Scholar] [CrossRef]
  6. Kokilashvili, V.; Rákosník, J. Weighted inequalities for vector-valued anisotropic maximal functions. Z. Für Anal. Und Ihre Anwendungen 1985, 4(6), 503–511. [Google Scholar] [CrossRef]
  7. Ditzian, Z. (1989, March). Multivariate Landau–Kolmogorov-type inequality. In Mathematical Proceedings of the Cambridge Philosophical Society (Vol. 105, No. 2, pp. 335-350). Cambridge University Press. March; No. 2. [CrossRef]
  8. Samko, S. G. (1993). Fractional integrals and derivatives. Theory and applications.
  9. Grigor’yan, A. Heat kernel upper bounds on a complete non-compact manifold. Rev. Matemática Iberoam. 1994, 10(2), 395–452. [Google Scholar] [CrossRef]
  10. Lions, P. L. (1996). Mathematical topics in fluid mechanics: volume 2: compressible models (Vol. 2). oxford university press.
  11. Kounchev, O. Extremizers for the multivariate Landau-Kolmogorov inequality. Math. Res. 1997, 101, 123–132. [Google Scholar]
  12. DeVore, R. A. Nonlinear approximation. Acta Numer. 1998, 7, 51–150. [Google Scholar] [CrossRef]
  13. Lieb, E. H.; Loss, M. (2001). Analysis (Vol. 14). American Mathematical Soc. Vol. 14.
  14. LeVeque, R. J. (2002). Finite volume methods for hyperbolic problems (Vol. 31). Cambridge university press.
  15. Triebel, H. (2006). Theory of function spaces III. Basel: Birkhäuser Basel. [CrossRef]
  16. Majda, A. (2012). Compressible fluid flow and systems of conservation laws in several space variables. Springer Science &amp; Business Media.
  17. Yarotsky, D. Error bounds for approximations with deep ReLU networks. Neural Netw. 2017, 94, 103–114. [Google Scholar] [CrossRef]
  18. Wang, L. A. D. A multiplier theorem on anisotropic Hardy spaces. Can. Math. Bull. 2018, 61(2), 390–404. [Google Scholar] [CrossRef]
  19. ANASTASSIOU, G. A. Multivariate left side Canavati fractional Landau inequalities. J. Appl. Pure Math. 2025, 7(1-2), 103–119. [Google Scholar] [CrossRef]
  20. Chen, J. Fractional and anisotropic Gagliardo–Nirenberg inequalities with applications. arXiv 2026, arXiv:2608.27841. [Google Scholar] [CrossRef]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.