ARR-RN-2026-038·Reading Note·2026-05-12

Endogenous Path-Dependence and the Sufficiency of Past Returns for Implied Volatility

· path-dependent volatility· implied volatility· SPX/VIX· long memory
§ Reviewed Work
Volatility is (mostly) path-dependent
J. Guyon, J. Lekeufack
SSRN 4174589 · Quantitative Finance 23(9), 2023
View source ↗
§01

Abstract

Guyon and Lekeufack construct a Path-Dependent Volatility (PDV) regression model in which implied volatility at any given strike and maturity is expressed as a deterministic function of two real-valued path functionals computed from the recent return history: a signed trend kernel R_{1,t} capturing the directional drift of log-prices over a power-law-weighted window, and an unsigned activity kernel R_{2,t} capturing the weighted sum of squared returns over the same window. The central empirical result is that a parsimonious nonparametric regression of SPX ATM implied volatility onto (R_{1,t}, R_{2,t}) achieves an in-sample R² exceeding 90%, a figure that remains robust across different regimes and is only marginally improved by including additional features. The paper therefore makes the two-feature sufficiency claim: the implied volatility surface, to a first approximation, is a measurable function of the past return path encoded through these two scalar summaries, and adding model-specific latent state variables such as a stochastic volatility factor delivers negligible marginal explanatory power once R_1 and R_2 are included. The kernel structure is central to the result. Both R_{1,t} and R_{2,t} use power-law decay kernels K_i(τ) = (1 + τ/θ_i)^{−α_i}, which are capable of simultaneously weighting recent returns heavily while retaining long-horizon information at a rate that decays slowly enough to capture the well-documented long memory of realized variance. The authors show that these power-law kernels can be approximated to arbitrary accuracy by a sum of four exponential functions — a 4-factor Markovian decomposition — enabling the path-dependent state (R_{1,t}, R_{2,t}) to be embedded in a four-dimensional Markovian system without loss of explanatory power at the level of the regression. This Markovian approximation is the bridge between the non-Markovian PDV description and practical implementation in a simulation or hedging framework. From the perspective of this desk's reading programme, the paper occupies a complementary position to Gatheral–Jaisson–Rosenbaum on roughness: both identify non-Markovian path structure as the dominant driver of volatility dynamics, but where roughness is a statement about the local Hölder regularity of the latent vol path, PDV is a statement about the global explanatory power of path integrals against the observable return process. The two are different coordinate representations of the same underlying non-Markovian data-generating process, and reconciling them quantitatively — asking how much of the RFSV path information is captured by (R_1, R_2) — is an open question this desk treats as material.

§02

Notation / Conceptual Frame

The model sets σ_t = f(R_{1,t}, R_{2,t}) where f is estimated nonparametrically and the two regressors are defined as convolutions of the return path against power-law kernels: R_{1,t} = Σ_{s < t} K_1(t − s) r_s and R_{2,t} = Σ_{s < t} K_2(t − s) r_s², where r_s = log(S_s / S_{s−1}) is the log-return at time s. The kernel forms are K_i(τ) = c_i (1 + τ/θ_i)^{−α_i} with normalizing constants c_i chosen so that Σ_τ K_i(τ) = 1 in discrete time; the four scalar parameters (θ_i, α_i) for each kernel are calibrated by maximizing the regression R². The power-law structure of K_i implies that R_{1,t} is a signed, long-memory weighted moving average of past returns, capturing the persistent trend effect, while R_{2,t} is an unsigned, long-memory weighted moving average of past squared returns, capturing persistent variance clustering. The 4-factor Markovian approximation expresses each power-law kernel as Σ_{k=1}^4 w_k exp(−β_k τ), yielding an expanded state vector Y_t = (Y_{1,t}^{(1)}, ..., Y_{4,t}^{(1)}, Y_{1,t}^{(2)}, ..., Y_{4,t}^{(2)}) ∈ R^8 that satisfies a linear SDE driven by r_t and r_t², enabling exact Markovian simulation and allowing the PDV framework to be embedded in classical derivative-pricing architectures.

§03

Commentary

The signed versus unsigned distinction between R_1 and R_2 is the structural insight at the core of the PDV framework. R_1 encodes the directional character of recent price action — whether the market has been trending up or down — and its signed contribution to σ_t captures the well-known leverage effect: negative trends elevate implied volatility, positive trends suppress it. R_2 by contrast captures the aggregate level of return variability regardless of direction, encoding the volatility clustering effect that HAR-type models have long identified as the dominant driver of realized variance at intermediate horizons. The two-feature regression dominates alternatives with more features not because the feature set is optimally chosen in any information-theoretic sense, but because these two functionals span the most informative directions in the space of return path statistics for the specific task of predicting ATM implied volatility at short-to-medium maturities. The relationship to HAR (Heterogeneous AutoRegressive) models is instructive: HAR models regress realized variance on daily, weekly, and monthly realized variance averages, which is a particular discrete approximation to the kind of multi-scale path integration that R_2 performs. The PDV framework generalizes this by using a continuously parametrized power-law kernel and by incorporating the signed trend channel R_1, which HAR models omit. The PDV regression therefore sits at the intersection of the HAR econometric tradition and the rough-vol theoretical tradition: it is empirical enough to require no latent state beyond the observable return path, and structural enough to connect to the Volterra integral representation underlying RFSV. The endogenous versus exogenous variance decomposition — what fraction of variance is explained by past returns versus by genuinely external shocks — is quantified by the R² itself and by the residual of the regression, and both rough-vol and PDV can be understood as projections of one non-Markovian data-generating process onto different observable features: rough-vol projects onto the latent vol path, PDV projects onto the return path, and the connection between the two is mediated by the leverage correlation between returns and volatility.

§04

Implications for Research Methodology

The two-feature sufficiency result has an immediate implementation consequence: it provides a concise, real-time-computable low-dimensional feature set — (R_{1,t}, R_{2,t}) — that serves as a near-sufficient statistic for the implied volatility surface. Because both features are defined as causal convolutions of the past return path, they can be updated recursively with each new return observation at negligible computational cost, and their Markovian 4-factor approximations can be pre-calibrated to the historical power-law parameters and then run forward in real time. For desk conditioning, this means that surface reads should be conducted not against a background of spot level and current realized variance alone, but against the full pair (R_1, R_2), which encodes both trend direction and activity level in a way that maximally summarizes the surface information available from the path. The practical implementation of R_1 and R_2 as state variables in the desk's signal-generation framework requires a calibrated choice of kernel parameters (θ_i, α_i) that may themselves be regime-dependent. The paper's cross-sample stability analysis suggests these parameters are relatively stable for SPX over multi-year windows, but the desk should treat single-name applications with higher parameter uncertainty and correspondingly wider conditioning bands. The insight that path features dominate level features for surface conditioning at rough-path horizons is consistent with the Gatheral–Jaisson–Rosenbaum roughness finding, and the two papers together motivate a unified path-feature conditioning architecture.

§05

Limitations

The 90% R² figure is an in-sample, index-level result for SPX, and its applicability to single-name equities is substantially more limited. Single-name implied volatility is driven by a combination of index-level path features and idiosyncratic factors — earnings, credit events, flow positioning — that are not captured by return path functionals, and out-of-sample R² figures for single names are substantially lower. The desk should treat the PDV feature set as a necessary but not sufficient conditioning input for single-name surface reads, supplementing (R_1, R_2) with name-specific flow and event signals. The distinction between explanatory R² and forecasting R² is material and the paper does not fully address it: a high in-sample R² with well-chosen kernel parameters does not guarantee that R_1 and R_2 constructed from past returns will forecast the next-period implied volatility with comparable accuracy, because the kernel parameters themselves were chosen to maximize the in-sample fit. The possibility of kernel parameter look-ahead bias — where the optimal (θ_i, α_i) estimated over the full sample period have implicitly absorbed future information — is a concern for any backtest of PDV-based conditioning signals, and the desk applies a strict expanding-window calibration discipline to guard against this source of overfitting.

§ Related Notes
This note is informational and interpretive. It does not constitute personalized investment advice. Market activity involves risk. Historical analysis and model outputs do not guarantee future results.