§ Publications

Research Archive

Reading notes, technical commentaries, and methodological annotations on selected peer-reviewed and preprint literature in quantitative finance and market microstructure. Each note reviews a single published work and records its bearing on desk methodology. Materials are interpretive and informational; they are not investment advice.

2026-06-16
ARR-RN-2026-052
Reading Note

Path-Dependence of the At-the-Money-Forward Implied Volatility Term Structure

Reviews: H. Andrès, A. Boumezoued, B. JourdainThe implied volatility surface (also) is path-dependent · arXiv:2312.15950 (2023)

Andrès, Boumezoued, and Jourdain extend the path-dependent volatility (PDV) programme of Guyon and Lekeufack from the level of instantaneous or realized variance to the entire at-the-money-forward (ATMF) implied volatility term structure. Fitting the SSVI (Surface SVI) parameterization to daily SPX option data and extracting the maturity-dependent ATMF total variance parameter θ(T) for maturities T ranging from one month to two years, the authors regress each θ(T) onto the same two path features that drive the PDV model — the signed trend kernel R_{1,t} = Σ_{s ≤ t} K_1(t−s) r_s and the unsigned activity kernel R_{2,t} = Σ_{s ≤ t} K_2(t−s) r_s² — and find that the maturity-specific path regression explains a substantial fraction of the variance of θ(T) at each maturity, with the explanatory power ranging from approximately 80% at short maturities to 60–70% at two years. The regression therefore extends the two-feature sufficiency claim to the full ATMF term structure: the term structure of implied volatility is, to a first approximation, a deterministic function of the past return path via two scalar features, rather than an independent state variable characterizing the market's forward variance expectations. The Bergomi forward variance curve ξ(T) = E_t[σ²_T] provides the theoretical framing: the PDV features R_1 and R_2 correspond to projections of the Bergomi curve onto the two dominant modes of its variation, so that the path regression is implicitly extracting the two leading principal components of the curve's dynamics directly from the observable return path. The maturity-dependent coefficient functions α_j(T) are estimated independently at each T, revealing that the loading on R_2 (activity) dominates at short maturities — consistent with the near-term variance predictability of realized variance — while the loading on R_1 (trend) grows in relative importance at long maturities — consistent with the contribution of directional drift to longer-horizon expected variance. The residual term-structure covariance, estimated from the regression residuals across maturities, reveals a structured two-factor pattern suggesting that a third path feature — at an intermediate timescale between those of K_1 and K_2 — would substantially close the remaining gap. For the desk, the conclusion is that the ATMF term structure is not an autonomous information source but is largely redundant with the observable return path: a surface read conducted independently of the path-conditioning step is reading largely the same information twice, and any apparent disagreement between the surface-implied term structure and the path-predicted term structure is the anomaly to investigate rather than the surface being a genuine additional signal. The surface's independent information is concentrated in the residual ε(T) — the component of each θ(T) unexplained by the path regression — and in the shape parameters of the SSVI fit (skew, wing asymmetry) that are not constrained by the ATMF regression.

· path-dependent volatility· implied volatility surface· SSVI· term structureRead Note →
2026-06-11
ARR-TC-2026-031
Technical Commentary

Signature-Linear Volatility and the Survival of the Riccati Structure

Reviews: E. Abi Jaber, L.-A. GérardSignature volatility models: pricing and hedging with Fourier · arXiv:2402.01820 · SIAM J. Financial Math. (2025)

Abi Jaber and Gérard build a class of volatility models in which the instantaneous variance σ_t² is expressed as a linear functional of the time-extended signature of a driving Brownian motion — σ_t² = ⟨ℓ, 𝕊(B̃)_{0,t}⟩ for B̃_t = (t, B_t) and a coefficient ℓ in the dual of the tensor algebra T((R²)) — and prove that the characteristic functional of the log-price and integrated variance retains an exponential-affine form analogous to the Heston characteristic function, with the scalar Riccati ODE of classical affine models replaced by a tensor-algebra-valued Riccati equation. The universality theorem underlying this class, which asserts that every continuous linear functional on the space of continuous paths of bounded p-variation can be expressed as an inner product with the signature, guarantees that the signature-linear family contains all path-dependent volatility specifications of practical interest as special cases — including Stein–Stein, Bergomi, and Heston in the Markovian limit, and rough and path-dependent variants in the infinite-dimensional regime — while maintaining a single unified Fourier pricing framework across all these models. The key technical contribution is the identification of the tensor-algebra Riccati equation as the correct generalization of the classical Heston Riccati ODE to path-dependent volatility. In the Heston model, the log of the characteristic function satisfies dH/dt = F(H, u) with F quadratic in H and the solution given by an analytic formula; in the signature-linear model, H(t) is valued in the tensor algebra T((R²)) and satisfies the same quadratic-in-H structure, now interpreted in the tensor product sense, with the solution no longer analytic but computable by numerical integration of the tensor-valued ODE. The connection to polynomial diffusions — where the expected signature E[𝕊(B̃)_{0,t}] satisfies a linear ODE on the tensor algebra, expressible as matrix exponentiation for any finite truncation level N — provides an additional computational pathway that is complementary to the Fourier method and more efficient for computing quadratic hedging ratios. The practical consequence for a derivatives desk is that path-dependence and Fourier tractability are not in opposition: the same FFT infrastructure used for Heston pricing can be adapted to signature-linear models by replacing the scalar Riccati solve with a tensor-valued solve, maintaining compatibility with the full infrastructure of fast calibration, smile interpolation, and greeks computation that the classical affine framework provides.

· rough paths· signature· characteristic functional· Fourier pricingRead Note →
2026-06-04
ARR-TC-2026-029
Technical Commentary

Joint SPX/VIX Calibration as Linear Optimization over Signature Coefficients

Reviews: C. Cuchiero, G. Gazzani, J. Möller, S. Svaluto-FerroJoint calibration to SPX and VIX options with signature-based models · arXiv:2301.13235 · Mathematical Finance (2025)

Cuchiero, Gazzani, Möller, and Svaluto-Ferro construct signature market models in which the SPX log-price is expressed as a polynomial functional of the signature of a primary diffusion process X, and demonstrate that within this class the joint calibration to SPX vanilla options and VIX options and futures — historically the hardest simultaneous fitting problem in derivatives modeling — can be cast as optimization over the signature coefficient vector ℓ, with the VIX² available in closed form as a signature functional of the conditional expected signature. The approach transforms the joint calibration problem from a nonlinear search over the parameters of a bespoke SDE — which requires a separate parameterization for the VIX dynamics beyond what the SPX dynamics specify — into a linear-algebraic problem in the signature coefficient space, where both SPX option prices and VIX option prices are computed as linear functions of ℓ (after appropriate linearization) and the optimization is therefore convex or nearly convex. The central computational device is the polynomial diffusion embedding: the expected signature E[𝕊(X)_{s,t} | F_s] satisfies the linear ODE d/dT E[𝕊(X)_{s,s+T}] = A · E[𝕊(X)_{s,s+T}] where A is the generator matrix of the associated polynomial diffusion on the truncated tensor algebra, enabling efficient computation of the VIX² = (1/Δ) E_t[∫_T^{T+Δ} σ_u² du] = ⟨ℓ^VIX, E[𝕊(X)_{T,T+Δ} | F_T]⟩ as a matrix-exponential times the current signature state. The VIX is therefore an affine-polynomial function of the state 𝕊(X)_{0,T} at the VIX fixing date T, guaranteeing that VIX options can be priced by simulation of the polynomial diffusion on the truncated tensor algebra without auxiliary SDE parameters. The universality of the signature basis ensures that the model class is rich enough to approximate any consistent joint SPX/VIX dynamics to arbitrary accuracy: the universal approximation theorem for signatures asserts that every continuous path functional can be approximated by a linear functional on the signature, and the polynomial diffusion structure on the truncated tensor algebra provides the computational machinery to evaluate these functionals and their conditional expectations efficiently. The result is a model class that achieves the long-sought joint calibration goal through linear algebra rather than through ad hoc parameter engineering.

· signature· joint calibration· VIX· convex optimizationRead Note →
2026-05-28
ARR-RN-2026-041
Reading Note

Roughness of the Log-Volatility Process and the Failure of Markovian Calibration

Reviews: J. Gatheral, T. Jaisson, M. RosenbaumVolatility is rough · arXiv:1410.3394 · Quantitative Finance 18(6), 2018

Gatheral, Jaisson, and Rosenbaum address a foundational question in stochastic volatility theory: what is the regularity of the log-volatility process as measured directly from high-frequency equity and index data? The central object of study is the scaling behavior of the q-th absolute moment of log-volatility increments, defined as m(q, Δ) = E[|log σ_{t+Δ} − log σ_t|^q], evaluated across a range of moment orders q and time lags Δ spanning several decades of scale. If this quantity scales as a power law Δ^{ζ(q)} with ζ(q) ≈ qH for a single constant H, the process belongs to the class of monofractal processes with Hurst index H, and the empirical content of the paper is precisely that H ≈ 0.1 across a wide collection of equity indices, with this estimate proving remarkably stable across instruments, sample periods, and estimation methodologies. The significance of H < 1/2 is that the log-volatility increments exhibit mean-reversion at all observable scales — the process is rougher than Brownian motion, and a model consistent with this finding cannot be Markovian in volatility. The theoretical vehicle proposed to capture this roughness is the Rough Fractional Stochastic Volatility (RFSV) model, in which log σ_t is driven by a fractional Brownian motion W^H with H ≈ 0.1 rather than a standard Brownian motion. Fractional Brownian motion, constructed via the Mandelbrot–van Ness representation as a moving-average integral of Brownian increments with a power-law kernel of exponent H − 1/2, is the unique Gaussian process with stationary increments and self-similar paths, and for H ∈ (0, 1/2) its paths are almost surely Hölder continuous of any order strictly less than H, which is significantly less regular than the Brownian case H = 1/2. The RFSV model sets log σ_t = ν W^H_t up to a mean-reverting correction, and the key structural consequence is that the covariance function of the log-volatility satisfies E[log σ_t log σ_s] ∼ |t − s|^{2H} for small |t − s|, a form that encodes extremely long memory of volatility at short scales despite eventual decorrelation at long scales. The multiscaling question — whether ζ(q) is truly linear in q or exhibits convexity consistent with multifractal models — is addressed empirically, and the data support the monofractal description within estimation uncertainty, though the authors are careful to note that the sample sizes available do not definitively exclude weak multiscaling. For desk purposes this paper functions as the empirical anchor for an entire family of non-Markovian models. It establishes that the standard toolkit of Markovian stochastic volatility — including Heston, SABR, and their calibrations — is structurally incapable of reproducing the empirically observed scaling, because any Markovian diffusion driven by a standard Brownian motion necessarily has H = 1/2 locally. The practical implication is not merely academic: any conditioning scheme that treats the current volatility level as a sufficient statistic for the future evolution of the smile is discarding the majority of the information available in the realized volatility path. This reading therefore motivates the use of path functionals of realized log-volatility as primary conditioning inputs, and it places the desk's entire rough-vol adjacent methodology on a quantitative empirical footing.

· rough volatility· fractional Brownian motion· Hurst exponent· realized varianceRead Note →
2026-05-21
ARR-RN-2026-047
Reading Note

A Four-Factor Markovian Path-Dependent Volatility Model as a Pricing Engine

Reviews: G. Gazzani, J. GuyonPricing and calibration in the 4-factor path-dependent volatility model · arXiv:2406.02319 (2024)

Gazzani and Guyon develop the operational pricing and calibration infrastructure for the four-factor Markovian path-dependent volatility (PDV) model — the low-dimensional approximation to Guyon–Lekeufack's non-Markovian PDV specification in which each power-law kernel is replaced by a sum of two exponentials, reducing the infinite-dimensional path history to a four-dimensional exponentially-weighted state vector — and demonstrate that the resulting model achieves competitive joint SPX/VIX smile calibration using neural-network-based approximation of the pricing functionals. The core construction is the approximation of each power-law kernel K_j(u) = c_j(1 + u/θ_j)^{−α_j} by a two-exponential sum Σ_{k=1}^2 w^j_k exp(−β^j_k u) with weights and decay rates calibrated to minimize the L² approximation error over the relevant timescale range, yielding four state variables Y^k_t = ∫_0^t e^{−β^j_k(t−s)} r_s^{p_j} ds (p_1 = 1, p_2 = 2) that together encode the trend and activity path features in a four-dimensional Markov process driven by the return increments. The Markovian state Y_t = (Y^1_t, Y^2_t, Y^3_t, Y^4_t) ∈ R^4 satisfies a linear SDE of the form dY^k_t = −β^k Y^k_t dt + c^k dX^k_t where X^k_t depends on r_t (for the trend factors, k=1,2) or r_t² (for the activity factors, k=3,4), making the full state process (S_t, Y_t) a five-dimensional Markov system amenable to PDE methods. The option price V(y, t; K, T) at current state y and calendar time t is approximated by a neural network V_θ(y, t; K, T) trained to satisfy the pricing PDE L V_θ = 0 via a physics-informed neural network (PINN) penalty, and the VIX is computed as the expectation of the variance swap rate over the 30-day forward window, itself approximated by a second neural network VIX_θ(y, T). The paper establishes that the four-factor Markovian approximation achieves implied volatility accuracy within 0.1–0.3 vol points relative to the true non-Markovian PDV model (as estimated by Monte Carlo of the full path-dependent model), and that the neural network pricing functions, once trained offline, provide inference at millisecond timescales per option price, enabling real-time joint SPX/VIX recalibration that would be infeasible with Monte Carlo simulation of the non-Markovian process.

· path-dependent volatility· Markovian approximation· neural pricing· SPX/VIXRead Note →
2026-05-12
ARR-RN-2026-038
Reading Note

Endogenous Path-Dependence and the Sufficiency of Past Returns for Implied Volatility

Reviews: J. Guyon, J. LekeufackVolatility is (mostly) path-dependent · SSRN 4174589 · Quantitative Finance 23(9), 2023

Guyon and Lekeufack construct a Path-Dependent Volatility (PDV) regression model in which implied volatility at any given strike and maturity is expressed as a deterministic function of two real-valued path functionals computed from the recent return history: a signed trend kernel R_{1,t} capturing the directional drift of log-prices over a power-law-weighted window, and an unsigned activity kernel R_{2,t} capturing the weighted sum of squared returns over the same window. The central empirical result is that a parsimonious nonparametric regression of SPX ATM implied volatility onto (R_{1,t}, R_{2,t}) achieves an in-sample R² exceeding 90%, a figure that remains robust across different regimes and is only marginally improved by including additional features. The paper therefore makes the two-feature sufficiency claim: the implied volatility surface, to a first approximation, is a measurable function of the past return path encoded through these two scalar summaries, and adding model-specific latent state variables such as a stochastic volatility factor delivers negligible marginal explanatory power once R_1 and R_2 are included. The kernel structure is central to the result. Both R_{1,t} and R_{2,t} use power-law decay kernels K_i(τ) = (1 + τ/θ_i)^{−α_i}, which are capable of simultaneously weighting recent returns heavily while retaining long-horizon information at a rate that decays slowly enough to capture the well-documented long memory of realized variance. The authors show that these power-law kernels can be approximated to arbitrary accuracy by a sum of four exponential functions — a 4-factor Markovian decomposition — enabling the path-dependent state (R_{1,t}, R_{2,t}) to be embedded in a four-dimensional Markovian system without loss of explanatory power at the level of the regression. This Markovian approximation is the bridge between the non-Markovian PDV description and practical implementation in a simulation or hedging framework. From the perspective of this desk's reading programme, the paper occupies a complementary position to Gatheral–Jaisson–Rosenbaum on roughness: both identify non-Markovian path structure as the dominant driver of volatility dynamics, but where roughness is a statement about the local Hölder regularity of the latent vol path, PDV is a statement about the global explanatory power of path integrals against the observable return process. The two are different coordinate representations of the same underlying non-Markovian data-generating process, and reconciling them quantitatively — asking how much of the RFSV path information is captured by (R_1, R_2) — is an open question this desk treats as material.

· path-dependent volatility· implied volatility· SPX/VIX· long memoryRead Note →
2026-05-07
ARR-RN-2026-044
Reading Note

Mean-Variance Allocation as a Linear Functional on the Signature of the Augmented Path

Reviews: O. Futter, B. Horvath, M. WieseSignature Trading: a path-dependent extension of the mean-variance framework with exogenous signals · arXiv:2308.15135 (2023)

Futter, Horvath, and Wiese generalize the classical Markowitz mean-variance framework to path-dependent strategies by representing the trading position at time t as a linear functional of the signature of a lead-lag-embedded augmented path — the joint path of asset prices and exogenous signals, augmented by a lead-lag construction that explicitly encodes the causal temporal ordering of signal values preceding subsequent price moves. The position process π_t = ⟨ℓ, 𝕊(Ẑ^{ll})_{0,t}⟩ for ℓ an element of the dual tensor algebra transforms the mean-variance objective into a quadratic programme in ℓ whose necessary and sufficient optimality conditions yield a closed-form solution ℓ* = A^{-1} b / λ, where A is the covariance matrix of the signature and b is its cross-covariance with the cumulative PnL — an exact analogue of the Markowitz weight vector but in the infinite-dimensional signature coefficient space. Drawdown control, momentum tilts, and mean-reversion strategies emerge as particular signature coefficients selected by the optimization, rather than requiring separate parametric models for each effect. The lead-lag embedding is the construction that enables the signature to capture causal predictability: without it, the signature of the joint path (S_t, Z_t) contains only contemporaneous cross-moments between price and signal, while the lead-lag version includes cross-iterated integrals of the form ∫_0^t Z_{s-Δ} dS_s that encode the lagged cross-correlation between signal and subsequent price move — the fundamental quantity that any momentum or mean-reversion strategy exploits. At signature degree 2 the resulting strategy includes terms proportional to ∫_0^t Z_{s-Δ} dS_s (signal-return cross-correlation), (S_t − S_0)² (variance of path), and ∫_0^t S_s ds (time-average of price, encoding mean-reversion), exactly recovering classical trend-following, variance-timing, and Bollinger-band strategies as special cases of the degree-2 signature strategy. At higher degrees the strategy captures non-linear interactions between signal and price history that cannot be expressed as any simple parametric strategy. The paper's technical elegance is that by reducing all of these effects to coefficients in a single linear functional on the signature, it provides a systematic and unified way to discover, represent, and estimate an optimal trading strategy from historical data without pre-specifying which combination of momentum, mean-reversion, and signal effects to include; the optimization jointly determines which effects are present and their optimal magnitudes, constrained only by the choice of signature truncation level N and the risk aversion parameter λ.

· signature· portfolio optimization· mean-variance· exogenous signalsRead Note →
2026-04-24
ARR-TC-2026-019
Technical Commentary

The Fractional Riccati Equation and Tractable Pricing under Rough Heston

Reviews: O. El Euch, M. RosenbaumThe characteristic function of rough Heston models · arXiv:1609.02108 · Mathematical Finance 29(1), 2019

El Euch and Rosenbaum establish the theoretical bridge between two independently motivated strands of the market microstructure and derivatives literature: the observation that near-critical Hawkes processes with power-law kernels generate rough variance paths in the macroscopic limit, and the need for a tractable characteristic-function formula for option pricing under non-Markovian rough volatility. The paper derives the rough Heston model as the limit in distribution of a rescaled sequence of nearly-critical Hawkes processes — where criticality means that the branching ratio approaches one and the kernel is power-law with exponent α ∈ (1/2, 1) — and shows that the macroscopic variance process V_t satisfies a fractional stochastic Volterra integral equation driven by a correlated Brownian motion. This micro-to-macro derivation is significant because it provides a mechanistic justification for rough volatility: the empirically observed roughness of index variance is a consequence of the near-critical, self-exciting nature of high-frequency order flow, and H = α − 1/2 ∈ (0, 1/2) is determined by the tail exponent of the microscopic excitation kernel. The characteristic function E_Q[e^{iuX_T}], where X_T = log(S_T/S_0) is the log-price at maturity T, is shown to admit a semi-closed-form expression analogous to the classical Heston formula. In the classical Heston model, the characteristic function is exp(g(T)u + ∫_0^T h(T−s)V_s ds) where h solves a linear Riccati ODE; in the rough Heston extension, the same exponential-affine structure is preserved but h now solves a fractional Riccati equation D^α h = F(h) where D^α is the Caputo fractional derivative of order α = H + 1/2. This is the central analytical result: the Volterra-affine structure of the variance process is rigid enough to generate a characteristic function that can be evaluated by numerical integration of the fractional ODE, enabling fast Fourier inversion for European option prices. The Volterra structure of the variance process deserves emphasis: V_t is not a semimartingale and does not have a standard Itô decomposition, yet the conditional characteristic function of X_T given the path of V retains the exponential-affine form in the running variance integral. This preservation of quasi-affine structure under fractional extension is the technical achievement that makes rough Heston analytically tractable at all, and it is what distinguishes it from more general rough volatility models such as the Bergomi model, where no such closed form exists.

· rough Heston· fractional Riccati· Hawkes limit· affine structureRead Note →
2026-04-16
ARR-TC-2026-024
Technical Commentary

Path-Dependent PDEs for Volterra Models and the Weak Error of Discretization

Reviews: O. Bonesini, A. Jacquier, A. PannierRough volatility, path-dependent PDEs and weak rates of convergence · arXiv:2304.03042 (2023)

Bonesini, Jacquier, and Pannier establish two connected results for stochastic Volterra models: first, that conditional expectations of smooth functionals of the solution path — the option pricing functions in a rough volatility model — are the unique classical solutions of associated path-dependent partial differential equations (PPDEs) in the sense of Dupire's functional Itô calculus; and second, that the weak error of Euler-type discretization schemes for smooth functionals of the Riemann–Liouville fractional Brownian motion W^H with H ∈ (0, 1/2) converges to zero at a rate of order Δt^H, where Δt is the time step, providing a rigorous quantitative bound on the discretization error introduced by any numerical scheme for rough volatility option pricing. The path-dependent PDE is the central theoretical object: in a classical Markovian setting, the conditional expectation V(t, x) = E_t[Φ(S_T) | S_t = x] satisfies the standard backward Kolmogorov PDE ∂_t V + (1/2)σ²(x)∂²_x V = 0, an equation in the finite-dimensional state (t, x). In a non-Markovian rough volatility setting, the state at time t is the entire path (X_s : s ≤ t) of the underlying process, and the conditional expectation V(t, X_{·∧t}) is a function of this infinite-dimensional path state, satisfying the PPDE ∂_t V + ∂^h_X V · σ(X_t) + (1/2)∂²_{XX} V · σ²(X_t) = 0, where ∂^h_X V is the horizontal derivative (rate of change under horizontal path extension) and ∂_{XX} V is the vertical derivative (rate of change under path perturbation at the endpoint) in the sense of Dupire's functional Itô calculus. The existence and uniqueness of classical solutions to this PPDE under the rough Volterra dynamics is the first main result, providing a rigorous PDE description of rough volatility option pricing. The weak error bound is the paper's second main contribution and its most practically significant result for numerical implementation: the H-dependent rate O(Δt^H) characterizes precisely how much of the option pricing error in any Euler scheme for rough vol is attributable to the coarse time-stepping, as opposed to the variance-reduction methods or payoff regularity, and it shows that the discretization error worsens as H decreases (rougher paths), with the extreme case H → 0 producing an error rate approaching O(1) — no convergence at all. For H = 0.1, the practical implication is that the weak convergence rate is Δt^{0.1}, meaning that reducing the time step by a factor of 2 reduces the weak error by only 2^{0.1} ≈ 1.07 — a dramatic slowdown relative to the Brownian case (H = 1/2) where the same step reduction cuts error by 2^{0.5} ≈ 1.41.

· rough volatility· path-dependent PDE· weak convergence· fractional Brownian motionRead Note →
2026-04-08
ARR-RN-2026-033
Reading Note

Endogeneity Near Criticality in Self-Exciting Mid-Price Dynamics

Reviews: S. J. Hardiman, N. Bercot, J.-P. BouchaudCritical reflexivity in financial markets: a Hawkes process analysis · arXiv:1302.1405 · Eur. Phys. J. B 86, 2013

Hardiman, Bercot, and Bouchaud model the sequence of mid-price changes in E-mini S&P 500 futures as a Hawkes point process, parameterizing the self-excitation kernel as a power law φ(τ) ∝ τ^{−(1+β)} with exponent β ∈ (0, 1), and estimating the branching ratio n = ∫_0^∞ φ(τ) dτ from data spanning 1998 through 2011. The central finding is that n remains persistently close to unity — in the range 0.9 to 1.0 — across the entire sample period, independent of market regimes, volatility levels, or significant structural changes in market microstructure over that decade. A branching ratio near one means that each exogenous event — a genuine news arrival or an externally motivated order — generates on average nearly one additional endogenously triggered event via the self-excitation mechanism, and the cascade of endogenous events then generates further cascades, producing a near-critical amplification of the original shock by a factor of approximately 1/(1−n). Near criticality, this amplification diverges, and the system's response becomes scale-free: the distribution of endogenous cascade sizes follows a power law with exponent related to β, there is no characteristic time scale for the decay of the self-excitation, and the market operates in a state that maximizes information transmission but also maximizes susceptibility to large endogenous fluctuations. The scale-free nature of the endogenous cascade under power-law kernels is the key structural feature that connects this paper to the rough volatility literature: the same power-law tail that generates near-criticality in the Hawkes branching structure is the kernel that, in the El Euch–Rosenbaum limit, produces the fractional Brownian motion driving rough volatility. The temporal structure of the cascades — with cross-event correlations decaying as τ^{−β} — maps directly onto the Hurst index H = 1 − β/2 of the macroscopic variance process, providing a microstructure interpretation of the roughness parameter. For regime interpretation at the desk level, the persistent proximity of n to unity suggests that the market operates in a state of permanent near-criticality rather than transitioning between clearly subcritical and clearly supercritical regimes, and this has implications for how the desk should interpret apparent momentum signals: much of what appears as directional momentum at short horizons is in fact the tail of an endogenous Hawkes cascade triggered by a small exogenous shock, and conditioning on the cascade's estimated duration is more informative than conditioning on the direction of the triggering move alone.

· Hawkes processes· reflexivity· branching ratio· endogeneityRead Note →
2026-04-01
ARR-TC-2026-021
Technical Commentary

Mesh-Free Solution of Path-Dependent PDEs in a Signature-Kernel RKHS

Reviews: A. Pannier, C. SalviA path-dependent PDE solver based on signature kernels · arXiv:2403.11738 (2024)

Pannier and Salvi develop a mesh-free numerical solver for path-dependent PDEs by representing the solution in the reproducing kernel Hilbert space (RKHS) generated by the signature kernel — the positive definite kernel on path space defined by k_sig(x, y) = ⟨𝕊(x), 𝕊(y)⟩_{T((R^d))} — and collocating the PPDE residual at a finite set of sampled paths, transforming the infinite-dimensional functional PDE into a finite-dimensional linear system in the kernel coefficients. The signature kernel, introduced by Salvi and collaborators in the context of statistical learning on sequential data, is the inner product between the signatures of two paths in the full (untrunacted) tensor algebra, and can be computed efficiently via a PDE on the two-parameter domain [0,T]² without explicit computation of the signature components; this makes it possible to work in an effectively infinite-dimensional signature space without incurring the curse of dimensionality that would arise from explicit signature truncation. The reproducing kernel Hilbert space framework for PPDE solution generalizes classical kernel collocation methods for PDEs (the Kansa method, radial basis function collocation) from finite-dimensional state spaces to the infinite-dimensional path space that is the natural domain of PPDEs. In the finite-dimensional case, the solution V(t, x) is approximated in the RKHS of a kernel k(x, x') on R^d, and the PDE is collocated at a finite set of state-space points; in the path-dependent case, the solution V(t, x_{·∧t}) is approximated in the RKHS of the signature kernel k_sig on the space of continuous paths, and the PPDE is collocated at a finite set of sampled paths. The convergence theory for kernel collocation in finite dimensions — which guarantees that the RKHS approximant converges to the true PDE solution as the number of collocation points grows and their fill distance shrinks — extends to the path-space setting under appropriate completeness and density conditions on the signature kernel RKHS. The computational workflow consists of three phases: (1) sampling a set of n collocation paths from a reference measure on path space (typically Monte Carlo paths from the underlying dynamics); (2) computing the kernel matrix K ∈ R^{n×n} with K_{ij} = k_sig(x^{(i)}, x^{(j)}) via the signature kernel PDE; and (3) solving the linear system (K + λ_reg I) α = F for the kernel coefficients α ∈ R^n, where F encodes the PPDE residual at each collocation path and λ_reg is a Tikhonov regularization parameter. The solution at a new path x is then V(t, x_{·∧t}) = Σ_{i=1}^n α_i k_sig(x^{(i)}, x_{·∧t}), computable in O(n) inner product evaluations.

· signature kernel· path-dependent PDE· RKHS· collocationRead Note →
2026-03-24
ARR-TC-2026-011
Technical Commentary

Cross-Impact Propagators, Matrix Volterra Equations, and Optimal Multi-Asset Execution

Reviews: E. Abi Jaber, E. Neuman, E. TuschmannOptimal portfolio choice with cross-impact propagators · arXiv:2403.10273 (2024)

Abi Jaber, Neuman, and Tuschmann develop a rigorous framework for optimal multi-asset portfolio execution in the presence of cross-impact — the phenomenon that trading in one asset shifts the price of other correlated assets — modeled through a matrix-valued Volterra propagator. The single-asset price impact literature (Almgren–Chriss, Gatheral's propagator model) has established that a trader who executes a large order shifts the market price of the traded asset by an amount proportional to the order's signed volume, with the impact decaying over time according to a mean-reverting kernel (the propagator). In the multi-asset case, the impact of trading asset j on the price of asset i is encoded by the off-diagonal entry G_{ij}(t) of a matrix-valued propagator G(t) ∈ R^{d×d}, so that the mid-price vector S(t) ∈ R^d evolves as S_t = S_0 + ∫_0^t G(t-s) dQ_s + M_t where Q_s is the cumulative trading vector at time s (so dQ_s is the vector of instantaneous trading rates), G(t-s) is the propagator matrix evaluated at lag t-s, and M_t is a martingale representing the exogenous price innovations unrelated to the trader's activity. The cross-impact propagator G_{ij}(t) quantifies both the own-impact of trading i on i (the diagonal G_{ii}) and the cross-impact of trading j on i (the off-diagonal G_{ij}), and the key empirical observation motivating the paper is that the off-diagonal elements G_{ij} for highly correlated assets — such as equities in the same sector, or assets linked by index arbitrage — are a significant fraction of the diagonal elements, meaning that ignoring cross-impact leads to substantially suboptimal execution strategies that trade too aggressively in correlated assets and incur more total impact cost than necessary. The optimal portfolio execution problem is to find the trading rate process θ_t (the vector of instantaneous trading rates dQ_t/dt) that minimizes the expected total execution cost C(θ) = E[∫_0^T (θ_t · S_t + (1/2) θ_t · Λ θ_t) dt] — the sum of the price impact cost (the price S_t is pushed away from its initial value by past trades) and the instantaneous execution cost (the bid-ask spread and market impact of the current trade θ_t, modeled by the positive definite matrix Λ) — subject to the inventory constraint ∫_0^T θ_t dt = q_0 (the total quantity q_0 must be executed by time T). The paper establishes that the optimal trading rate θ* satisfies a matrix Fredholm integral equation of the second kind: θ*_t + ∫_t^T K(s-t) θ*_s ds = λ(t) where K is the matrix resolvent kernel of the propagator G and λ(t) is a Lagrange multiplier vector determined by the inventory constraint. This is the direct multi-asset generalization of the scalar Fredholm equation arising in single-asset propagator models (Gatheral's formula for the optimal liquidation rate in the Almgren–Chriss model with power-law impact), and its solution provides the globally optimal trading schedule across all d assets simultaneously. The technical novelty lies in the operator-theoretic analysis of the matrix Fredholm equation: the resolvent kernel K must be constructed from the propagator G via the matrix Volterra equation G + K * G = K (where * denotes matrix convolution), a vector-space generalization of the scalar Volterra resolvent that requires G to satisfy a positive semi-definiteness condition (the no-statistical-arbitrage condition) for the resolvent K to exist and for the trading strategy θ* to be admissible (finite cost). The paper proves that the matrix propagator G satisfies this condition if and only if the spectral measure of the Fourier transform ĝ(ω) = ∫_0^∞ e^{-iωt} G(t) dt satisfies Re[ĝ(ω)] ≥ 0 for all ω — a generalization of the Bochner condition for positive definiteness — and that under this condition the Fredholm equation for θ* has a unique solution in L²([0,T], R^d).

· price impact· cross-impact· Volterra propagator· optimal executionRead Note →
2026-03-19
ARR-MA-2026-012
Methodological Annotation

Self- and Mutually-Exciting Processes across the Microstructure Stack

Reviews: E. Bacry, I. Mastromatteo, J.-F. MuzyHawkes processes in finance · arXiv:1502.04592 · Market Microstructure and Liquidity 1(1), 2015

Bacry, Mastromatteo, and Muzy provide a comprehensive survey of Hawkes processes applied across the full microstructure stack, unifying under a single mathematical framework several applications that are often treated as methodologically disjoint. The survey covers: tick-level volatility estimation via the identification of quadratic variation with the integral of the Hawkes intensity; measurement of market endogeneity via the scalar branching ratio; cross-asset and cross-venue systemic contagion analysis via the spectral radius of the kernel matrix; optimal execution under Hawkes-driven market impact in the Almgren–Chriss framework; and full order-book modeling via bid- and ask-side excitation intensities. The central thesis is that a single mathematical object — the matrix kernel Φ(τ) of a multivariate Hawkes process — encodes the cross-event dynamics at all of these levels, and that calibrating Φ from high-frequency data provides a unified basis for volatility modeling, contagion measurement, and momentum signal interpretation. The survey's taxonomic value for a quantitative desk is substantial. By demonstrating that the branching ratio (Hardiman–Bouchaud), the rough-vol kernel (El Euch–Rosenbaum), and the order-book excitation (various authors) are all special cases of the multivariate Hawkes framework at different levels of aggregation, the paper creates a common vocabulary and a common set of estimation tools that can be applied across the desk's different analytical functions without requiring the construction of separate theoretical foundations for each application. The spectral radius ρ(∫Φ) — the largest eigenvalue of the matrix of integrated kernels — emerges as a single aggregate endogeneity indicator that collapses the multivariate excitation structure to a scalar, providing a natural generalization of the scalar branching ratio to the cross-asset setting. For this desk's methodology, the Bacry–Mastromatteo–Muzy survey occupies the role of an architectural reference: it is the document that establishes the common process-level language shared across microstructure, volatility, and cross-asset analysis, and it is the background against which all specific Hawkes-based methodological choices should be understood and justified.

· Hawkes processes· market microstructure· contagion· order flowRead Note →
2026-03-17
ARR-TC-2026-010
Technical Commentary

Non-Adversarial Neural SDE Calibration via Strictly Proper Signature Kernel Scoring Rules

Reviews: Z. Issa, B. Horvath, M. Lemercier, C. SalviNon-adversarial training of Neural SDEs with signature kernel scores · arXiv:2305.16274 (2023)

Issa, Horvath, Lemercier, and Salvi address a fundamental methodological problem in the training of neural stochastic differential equations for financial market simulation: the instability and mode collapse endemic to generative adversarial network (GAN) approaches to path-space distribution matching. Neural SDEs — SDEs of the form dX_t = b_θ(t, X_t) dt + σ_θ(t, X_t) dW_t where the drift b_θ and diffusion σ_θ are neural networks parameterized by θ — are natural candidates for flexible market simulators because they combine the inductive bias of the SDE framework (continuous paths, no-arbitrage structure, interpretable dynamics) with the universal approximation capacity of neural networks. The standard training approach matches the distribution of simulated paths X under the neural SDE to the distribution of historical market paths via a GAN objective — a discriminator neural network D_ψ is trained to distinguish simulated from real paths while the generator (the neural SDE) is trained to fool D_ψ — but this min-max formulation is notoriously unstable in practice, suffering from discriminator collapse, gradient vanishing, and failure to converge to a unique equilibrium. The paper replaces the GAN objective with a strictly proper scoring rule on path space: a functional S(P, ω) such that E_{ω' ∼ P}[S(P, ω')] ≥ E_{ω ∼ Q}[S(P, ω)] for all P ≠ Q, with equality if and only if P = Q, where ω denotes a realized path. A strictly proper scoring rule defines a proper divergence D_S(P, Q) = E_Q[S(P, ω)] - E_P[S(P, ω)] ≥ 0 that is zero only when P = Q, meaning minimizing D_S(P_θ, P_data) over the neural SDE parameters θ achieves training by minimizing a well-defined loss function — not a min-max game — that has a unique global minimum when the neural SDE correctly captures the data distribution. The key innovation is the construction of a strictly proper scoring rule on the infinite-dimensional space of continuous paths using the signature kernel: the scoring rule S_{sig}(P, ω) = E_{X ∼ P}[k_{sig}(X, ω)] - (1/2) E_{X,X' ∼ P}[k_{sig}(X, X')] + const, where k_{sig}(X, Y) is the signature kernel (the inner product of the path signatures in the tensor algebra), is strictly proper because the signature kernel is a characteristic kernel on path space — it distinguishes all distinct path distributions — and the scoring rule S_{sig} is the kernel score associated to k_{sig}, which is strictly proper by the Gretton–Sriperumbudur–Fukumizu theorem for characteristic kernels. The computational advantage over GAN training is substantial: the training loss D_S(P_θ, P_data) = E_{X ∼ P_θ, ω ∼ P_data}[k_{sig}(X, ω)] - (1/2) E_{X,X' ∼ P_θ}[k_{sig}(X, X')] is a simple expected value that can be estimated by Monte Carlo without any discriminator network, the gradient ∂D_S/∂θ can be computed by differentiating through the SDE simulation (using adjoint sensitivity methods for the SDE) and through the signature kernel computation, and the optimization problem is a standard minimization rather than a min-max game, enabling stable gradient descent with well-understood convergence guarantees. The signature kernel k_{sig}(X, Y) = Σ_{n=0}^{∞} 〈X^{⊗n}, Y^{⊗n}〉 / n! can be computed exactly for polynomial kernels or via the kernel PDE (∂_{s,t} k_{sig}(X_{[0,s]}, Y_{[0,t]}) = k_{sig}(X_{[0,s]}, Y_{[0,t]}) · dX_s · dY_t), and the Gram matrix (k_{sig}(X_i, X_j))_{i,j} for a batch of simulated paths can be assembled on a GPU with matrix-parallel computation whose cost is comparable to the SDE simulation itself.

· neural SDE· signature kernel· scoring rules· market simulationRead Note →
2026-03-12
ARR-TC-2026-016
Technical Commentary

A Rough-Path PDE Representation for Local Stochastic Volatility

Reviews: P. Bank, C. Bayer, P. K. Friz, L. PelizzariRough PDEs for local stochastic volatility models · arXiv:2307.09216 (2023)

Bank, Bayer, Friz, and Pelizzari develop a rough-PDE representation of option pricing functions in the class of local stochastic volatility (LSV) models, proving that the pricing functional V(t, s; 𝐖_{[0,t]}) — where 𝐖 is the geometric rough-path lift of the driving noise process — satisfies a linear PDE driven by the rough path 𝐖 in the sense of Lyons' rough-path integration theory, with the PDE coefficients being measurable functions of the current state (t, s) and the rough path 𝐖 up to time t. The Feynman–Kac correspondence is extended from the classical Itô setting — where the pricing PDE is a deterministic PDE driven by the current state — to the rough-path setting, where the pricing PDE is a random PDE driven by 𝐖 as a rough-path signal, with the solution given by the conditional expectation of the payoff under the rough-path-driven stochastic flow. The local stochastic volatility model combines a local volatility function σ_loc(t, S_t) with a stochastic volatility factor ξ_t driven by rough or diffusive noise, so that the instantaneous variance is σ_t² = σ_loc(t, S_t)² · ξ_t. The local volatility function σ_loc is calibrated to vanilla option prices via the Dupire formula, ensuring exact smile calibration at each maturity, while ξ_t introduces the stochastic dynamics required to explain the dynamics of the smile and the correlation between spot and volatility that a pure local volatility model misses. The rough-PDE framework handles the case where ξ_t is driven by rough noise (ξ_t = exp(ν W^H_t − ν²t^{2H}/2) in the rough Bergomi convention) by treating the rough path 𝐖 = (W^H, (W^H)^{⊗2}) — the rough path lift including the iterated integral — as a fixed input signal and writing the pricing equation as a PDE conditional on 𝐖. The extension of Feynman–Kac to the rough-path setting requires careful treatment of the composition of rough-path integration with the conditional expectation: the classical proof of Feynman–Kac uses Itô's formula to show that the conditional expectation solves the backward Kolmogorov PDE, and in the rough-path setting this step requires the rough-path chain rule (the Lyons–Gubinelli change of variables formula) applied to the composition of the pricing functional with the rough-path flow, producing the rough-PDE with drivers given by the controlled rough-path structure of the solution.

· rough paths· local stochastic volatility· Feynman–Kac· rough PDERead Note →
2026-03-10
ARR-TC-2026-009
Technical Commentary

Dispersion-Constrained Martingale Schrödinger Bridges and the Joint SPX/VIX Calibration Problem

Reviews: J. GuyonDispersion-constrained martingale Schrödinger problems and the exact joint S&P 500/VIX smile calibration puzzle · Finance and Stochastics 28(1), 2024

Guyon's paper addresses the longstanding calibration puzzle of jointly fitting the implied volatility surface of the S&P 500 (SPX) and the VIX options market simultaneously under a single coherent model. The puzzle arises because the VIX — defined as VIX²_T = (1/Δ) E_Q[∫_{T}^{T+Δ} σ²_t dt | F_T] under the risk-neutral measure Q — constrains the instantaneous variance forward curve in a way that is difficult to reconcile with the smile shapes observed in SPX vanilla options, particularly for short maturities where the SPX smile is steep and the VIX smile is convex. Standard stochastic volatility models calibrated to one surface typically fail to match the other — rough Heston and SABR models reproduce SPX term structure and skew but systematically misprice VIX options, while mean-reverting variance models that fit VIX smiles produce SPX smiles with incorrect slope and term structure. The core contribution is to embed the joint calibration problem within a variational framework in which the risk-neutral measure Q* is identified as the solution to a dispersion-constrained martingale Schrödinger problem: minimize the relative entropy H(Q|Q_ref) subject to three classes of constraints — the static SPX marginal constraint (the risk-neutral distribution of S_T under Q equals the market-implied marginal μ_T for each maturity T), the VIX constraint (the conditional variance E_Q[VIX²_T | F_T] matches the market price of VIX futures and the distributional constraint from VIX option prices defines the marginal of VIX_T under Q), and the martingale constraint (the undiscounted price process (S_t)_{t≥0} is a local martingale under Q). Together these three constraints form the dispersion constraint because the VIX definition ties the conditional expected quadratic variation to the observable VIX futures level, restricting the set of admissible measures Q in a way that is more stringent than either the marginal or martingale constraint alone. The Schrödinger bridge formulation — minimize H(Q|Q_ref) over Q in M(μ, ν) subject to the dispersion constraint — has the critical property that the optimizer Q* is unique whenever it exists (by strict convexity of relative entropy) and is characterized by an exponential tilting of Q_ref: Q*(dω) = exp(φ(ω_{T_1}) + ψ(ω_{T_2}) + λ · g(ω)) Q_ref(dω) where φ, ψ are the Lagrange multipliers for the two marginal constraints, λ is the Lagrange multiplier for the dispersion constraint, and g(ω) is a functional encoding the VIX definition in terms of the path ω. The dual problem is: maximize inf_{Q} [H(Q|Q_ref) - E_Q[φ(X_{T_1}) + ψ(X_{T_2}) + λ · g]] over (φ, ψ, λ), which has a smooth concave dual function in the Lagrange multipliers. The exponential tilting structure means that Q* takes the form of a reference process Q_ref — typically a continuous-time diffusion or a local volatility model — tilted by an exponential density that incorporates the market constraints; this is the martingale analogue of the classical Sinkhorn–Knopp theorem for coupling probability measures, and it establishes that the joint SPX/VIX calibration problem admits a unique solution in the minimum-entropy class. The algorithmic implementation proceeds via a Sinkhorn-type fixed-point iteration on the Lagrange multipliers: initialize (φ_0, ψ_0, λ_0), then alternate between updating φ to match the SPX marginal constraint (a single-period Schrödinger equation solved by a path-integral over Q_ref), updating ψ to match the VIX marginal constraint, and updating λ to match the VIX futures level and dispersion target. Each iteration step requires a Monte Carlo or PDE solve over the reference process Q_ref, and convergence is established by the strict convexity of the dual function and the invertibility of the constraint Jacobian at the optimum. The rate of convergence is geometric in the number of Sinkhorn sweeps with contraction rate determined by the relative entropy distance between Q_ref and Q*, meaning that a well-chosen reference process Q_ref — one that is close in relative entropy to the market measure — accelerates convergence significantly.

· martingale Schrödinger· relative entropy· joint calibration· VIXRead Note →
2026-03-02
ARR-RN-2026-027
Reading Note

Semi-Static Hedging and the Duality of Model-Free Option Bounds

Reviews: M. Beiglböck, P. Henry-Labordère, F. PenknerModel-independent bounds for option prices: a mass transport approach · arXiv:1106.5929 · Finance and Stochastics 17(3), 2013

Beiglböck, Henry-Labordère, and Penkner recast the problem of model-independent option bounds as a martingale optimal transport problem: given the terminal marginal distributions μ and ν of a price process at two dates T_1 < T_2, implied by the vanilla option prices at those maturities, what is the supremum (respectively infimum) of the price of an exotic payoff Φ(S_{T_1}, S_{T_2}) over all martingale measures consistent with those marginals? The primal problem is a linear programme over the set M(μ, ν) of martingale couplings — joint distributions of (S_{T_1}, S_{T_2}) that have μ and ν as marginals and satisfy the martingale constraint E_Q[S_{T_2} | S_{T_1}] = S_{T_1} — and the dual problem is a static superreplication strategy: a combination of vanilla options at T_1 and T_2 and a dynamic hedge in the underlying between T_1 and T_2. The principal result of the paper is the no-duality-gap theorem: under mild regularity conditions on Φ, the primal supremum equals the dual infimum, and the dual minimizer is an explicit semi-static hedge that achieves the bound. The Kantorovich duality structure of the problem mirrors classical optimal transport, with the key modification that the coupling must be a martingale. The role of the convex order condition — μ ≤_c ν, meaning ∫φ dμ ≤ ∫φ dν for all convex functions φ — as the necessary and sufficient condition for the existence of any martingale coupling is established via Strassen's theorem, which states that the convex order is equivalent to the existence of a measure-preserving martingale from μ to ν. The convex order condition is therefore the no-arbitrage condition on the vanilla marginals: if it is violated, the market prices of vanillas at T_1 and T_2 are mutually inconsistent. When the condition holds, the set M(μ, ν) is non-empty and the transport problem is well-posed, and the bound gives the sharpest possible constraint on the exotic price consistent with the observed vanilla data. The desk reading of this paper is motivated by the need for model-free guardrails on exotic exposures during events, where the vanilla surface provides reliable information about marginal distributions but the model-specific coupling assumption is the source of greatest uncertainty. The martingale transport bound replaces the model-specific price with the worst-case price over all consistent couplings, and the gap between this bound and the model price quantifies the model-assumption dependence of the exotic valuation.

· martingale optimal transport· robust pricing· convex duality· semi-static hedgingRead Note →
2026-02-24
ARR-MA-2026-008
Methodological Annotation

Geometry of Martingale Optimal Transport and the Cost of Robust Bounds

Reviews: J. Z.-G. Hiew, T. Lim, B. Pass, M. Cruz de SouzaDimension reduction in martingale optimal transport: geometry and robust option pricing · arXiv:2309.04947 (2023)

Hiew, Lim, Pass, and Cruz de Souza investigate the geometry of optimal solutions to the martingale optimal transport problem, proving that for a broad class of payoff functions Φ(x, y) — including all convex-concave payoffs and many exotic option payoffs — the optimizer Q* ∈ M(μ, ν) is supported on a low-dimensional set in the product space R × R, specifically on the graph of a measurable function or on a set of dimension at most one (a curve) in each fiber {x} × R. This concentration of the optimizer's support — called the martingale transport map or the irreducible martingale transport — provides two simultaneous benefits: it gives structural insight into what the model-free worst-case dynamics look like (they are supported on a degenerate coupling rather than a diffuse distribution over all of R²), and it provides a route to tractable numerical computation of the bounds since the optimization can be restricted to the low-dimensional support class rather than over all of M(μ, ν). The dimension-reduction result exploits the geometry of the convex order condition and the martingale constraint to show that the optimizer must satisfy a necessary condition — the co-monotonicity or left-curtain property — that forces the support of Q* to be contained in a specific one-dimensional set determined by the payoff Φ and the marginals (μ, ν). For the class of payoffs that are convex in y — including calls, puts, and variance swap payoffs — the optimizer of the upper bound UB = sup_{Q ∈ M(μ,ν)} E_Q[Φ(X,Y)] is the Kellerer left-curtain coupling, which transports each atom of μ to a two-point distribution supported on adjacent quantile levels of ν; for payoffs convex in x and concave in y (including some corridor options and spreads), a different degenerate coupling structure arises. The key insight is that the optimizer is always extremal in the convex set M(μ,ν) — it lies on the boundary rather than the interior — and extreme points of M(μ,ν) have the low-dimensional support property by the Choquet representation theorem applied to the compact convex set M(μ,ν). The computational consequence is that the numerical MOT problem can be solved on the low-dimensional support rather than by discretizing the full product measure on a grid, reducing the problem from O(n²) variables (a coupling over n × n grid points) to O(n) variables (a one-dimensional transport on the support manifold), enabling exact computation of model-free bounds for realistic option-market grids without the scaling limitations that make the full MOT problem computationally expensive.

· martingale optimal transport· robust pricing· convex geometry· model uncertaintyRead Note →
2026-02-16
ARR-TC-2026-011
Technical Commentary

Static-Arbitrage Constraints on the SVI Parameterization

Reviews: J. Gatheral, A. JacquierArbitrage-free SVI volatility surfaces · arXiv:1204.0646 · Quantitative Finance 14(1), 2014

Gatheral and Jacquier characterize precisely the conditions under which the Stochastic Volatility Inspired (SVI) parameterization of the implied volatility smile, and its surface extension across maturities, is free of static arbitrage. Three types of static arbitrage are relevant: butterfly arbitrage, which requires the total implied variance as a function of log-moneyness k to be convex (∂²w/∂k² ≥ 0 everywhere) so that the risk-neutral density implied by the smile is non-negative; calendar-spread arbitrage, which requires that the total implied variance w(k, T) be non-decreasing in T at every fixed k; and the large-moneyness (Lee) moment formula constraints, which require the right and left wings of the smile to grow at most linearly in |k| as |k| → ∞. The raw SVI parameterization, with five parameters (a, b, ρ, m, σ), does not automatically satisfy any of these constraints, and the paper derives the explicit parameter-space conditions under which each constraint holds, providing a practical algorithm for butterfly-free and calendar-spread-free calibration. The SSVI (Surface SVI) sub-parameterization introduced in the paper imposes a specific functional relationship between the ATM total variance θ = w(0, T) and the remaining shape parameters, yielding a two-parameter family φ(θ) (plus the correlation ρ) such that the resulting surface is arbitrage-free by construction for any choice of φ satisfying three explicit inequalities. The power-law and Heston-compatible choices of φ(θ) are analyzed in detail, providing both a flexible fitting parameterization and one that is anchored to the Heston model's structural predictions for the term structure of the ATM variance. The Lee moment formula — which constrains the maximum slope of the smile wings in terms of the moments of the risk-neutral distribution — is recovered as a special case of the SSVI constraints. For the desk, the significance of this paper is primarily methodological: any quantitative analysis of implied volatility surface features — skew level, term structure slope, wing behavior — is meaningful only if conducted on a surface that has first been verified or fitted to be free of static arbitrage. An arbitrage-laden surface is not the output of any consistent pricing model, and conditioning signals derived from its features are therefore conditioning on numerical artifacts rather than on genuine market information. The SSVI parameterization is accordingly adopted as the desk's canonical surface representation for any analysis that conditions on surface shape.

· SVI· implied volatility surface· static arbitrage· calibrationRead Note →
2026-01-29
ARR-RN-2026-022
Reading Note

The Order Book as a Markov Queuing System under a Fixed Reference Price

Reviews: W. Huang, C.-A. Lehalle, M. RosenbaumSimulating and analyzing order book data: the queue-reactive model · arXiv:1312.0563 · J. Am. Stat. Assoc. 110(509), 2015

Huang, Lehalle, and Rosenbaum propose a model in which the limit order book, conditional on a fixed reference price, is represented as a continuous-time Markov chain on the state space of non-negative integer queue vectors, with order-flow intensities — submission, cancellation, and execution — allowed to depend on the full current configuration of standing liquidity across price levels. The key conceptual separation is between within-regime dynamics, in which the reference price (the mid or best bid-ask midpoint) remains fixed and the queue configurations evolve according to the Markov chain, and regime transitions, in which an event depletes a best-level queue to zero and causes a reference-price jump. Within each regime the system reaches approximate stationarity before the next price jump, which justifies fitting the state-dependent intensities from empirical data under a piecewise-stationarity hypothesis and then assembling the lower-frequency price dynamics from the renewal process of regime transitions. The state-dependent intensity structure is the empirical core of the paper: at each price level i relative to the reference price, the arrival rate λ^+_i and cancellation rate λ^-_i are estimated as functions of the current queue depths (Q_1, ..., Q_K), revealing that the system is genuinely non-homogeneous — deep queues suppress further arrivals at that level and attract arrivals at adjacent levels in ways that a constant-intensity Poisson model cannot capture. The mean-field approximation, which factors the joint queue distribution into a product of marginals evaluated at conditional mean depths, provides a computationally tractable approximation whose predictions for the stationary book shape and the first-passage distribution to boundary (price-jump time) match simulated and empirical data closely. Critically, the model reproduces not just the time-averaged book shape but also the autocorrelation structure of queue imbalance, inter-trade durations, and the conditional distribution of price moves given book configuration. For the desk, the queue-reactive model establishes a precise definition of what constitutes genuine price-formation pressure at the microstructure level: only reference-price transitions carry directional information about the underlying price process, while queue-level dynamics within a regime are essentially mechanical, mean-reverting fluctuations around the stationary book shape. This decomposition has direct consequences for how short-horizon order-book signals should be filtered: apparent imbalance at the best bid and ask is predominantly within-regime queue noise and not informative about the direction of the next reference-price move, while persistent depletion of a best-level queue approaching boundary conditions is the genuine price-formation signal.

· limit order book· queue-reactive model· Markov queues· microstructureRead Note →
2026-01-12
ARR-TC-2026-006
Technical Commentary

Hedging under Frictions as Convex Risk Minimization over Neural Policies

Reviews: H. Buehler, L. Gonon, J. Teichmann, B. WoodDeep Hedging · arXiv:1802.03042 · Quantitative Finance 19(8), 2019

Buehler, Gonon, Teichmann, and Wood recast the problem of derivative replication under market frictions as the minimization of a convex, law-invariant risk measure ρ over policies parameterized by feedforward neural networks, providing a computationally tractable framework that accommodates proportional and fixed transaction costs, liquidity constraints, position limits, and market impact simultaneously, without requiring a closed-form pricing formula or an analytically tractable Markov structure for the underlying dynamics. The core theoretical result is a universal approximation theorem for the policy class: for any ε > 0 and any admissible Markovian hedging policy in the classical sense, there exists a neural network of finite width and depth such that its output, evaluated on the observable information set at each rebalancing date, approximates the optimal action to within ε in Lp norm, guaranteeing that the deep hedging approximation class is dense in the set of Markovian strategies as network size grows. The Fenchel–Legendre duality of convex risk measures provides the theoretical scaffolding: for any convex law-invariant ρ satisfying the Fatou property, the dual representation ρ(X) = sup_{Q ∈ M_ρ} {E^Q[−X] − ρ*(Q)} holds, where M_ρ is the risk envelope of ρ and ρ* its conjugate functional. This dual form identifies the optimal hedge δ* as the strategy that simultaneously minimizes the primal objective and achieves the dual supremum — a saddle-point condition that connects deep hedging to the classical theory of equivalent martingale measures and provides an interpretation of the trained policy in terms of the least-favorable measure Q* ∈ M_ρ induced by the risk measure. The indifference price of a contingent claim Φ, defined as the shift p* for which the risk of the hedged position plus the claim equals the risk of the unhedged zero-PnL position, is a further byproduct of the duality structure and can be read off from the optimal policy's terminal distribution. The paper's governance prescription — that the objective function, the friction model, and the training distribution must each be explicitly chosen and documented before the policy is trained — is as important as its technical content. A neural policy is only as disciplined as its training specifications, and in the absence of closed-form guidance, the choices of ρ, C_k, and the simulation measure collectively determine whether the learned strategy is robust or fragile, generalizing or overfitting, auditable or opaque.

· deep hedging· transaction costs· convex risk measures· reinforcement learningRead Note →
2025-12-18
ARR-MA-2025-048
Methodological Annotation

Random-Matrix Limits on the Information Content of Empirical Correlation Matrices

Reviews: L. Laloux, P. Cizeau, J.-P. Bouchaud, M. PottersNoise dressing of financial correlation matrices · arXiv:cond-mat/9810255 · Phys. Rev. Lett. 83, 1999

Laloux, Cizeau, Bouchaud, and Potters compare the empirical eigenvalue spectrum of sample correlation matrices constructed from S&P 500 returns against the prediction of the Marchenko–Pastur law — the limiting spectral distribution of a Wishart matrix (1/T) X X^T for X a matrix of i.i.d. entries — and find that the bulk of eigenvalues is statistically indistinguishable from pure noise, with only a small number of outlier eigenvalues above the upper spectral edge λ_+ = σ²(1 + √q)² carrying genuine factor structure. The ratio q = N/T, where N is the number of assets and T the number of observations, determines the noise band: for typical estimation problems with N = 400–500 stocks and T = 250–1000 daily returns, q ranges from 0.4 to 2, placing between 40% and the majority of eigenvalues inside the noise band. The eigenvalue cleaning procedure replaces in-band eigenvalues with their Marchenko–Pastur bulk mean while retaining the outlier eigenvalues and their eigenvectors, producing a cleaned covariance matrix that is a materially better estimator of the population covariance matrix in every matrix norm. The Marchenko–Pastur theorem, established for large Wishart matrices under i.i.d. sub-Gaussian entries, provides the limiting spectral density ρ_MP(λ) = (1/(2πσ²qλ))√{(λ_+ − λ)(λ − λ_-)} on the support [λ_-, λ_+] with λ_± = σ²(1 ± √q)², and the result extends to more general population covariance structures via the companion equation of Marchenko–Pastur as sharpened by Silverstein and Choi. The identification of outlier eigenvalues with genuine financial factors relies on the Baik–Ben Arous–Péché phase transition: population eigenvalues exceeding the critical threshold σ²(1 + √q)² produce spikes in the empirical spectrum that are detectable above the noise band in the large-(N,T) limit, while weaker factors below this threshold are absorbed into the bulk and are statistically indistinguishable from noise regardless of their economic significance. The desk reads this paper as establishing a hard statistical ceiling on the information extractable from any finite-sample cross-sectional analysis: the eigenvalue cleaning result quantifies exactly how much of the sample correlation matrix is noise, and any portfolio optimization or risk attribution that uses the uncleaned matrix is inflating noise by a factor of order T/N in the precision matrix entries, producing minimum-variance portfolios and risk decompositions that are dominated by estimated correlations that are statistically meaningless.

· random matrix theory· correlation matrix· Marchenko–Pastur· factor structureRead Note →