Abstract
Futter, Horvath, and Wiese generalize the classical Markowitz mean-variance framework to path-dependent strategies by representing the trading position at time t as a linear functional of the signature of a lead-lag-embedded augmented path — the joint path of asset prices and exogenous signals, augmented by a lead-lag construction that explicitly encodes the causal temporal ordering of signal values preceding subsequent price moves. The position process π_t = ⟨ℓ, 𝕊(Ẑ^{ll})_{0,t}⟩ for ℓ an element of the dual tensor algebra transforms the mean-variance objective into a quadratic programme in ℓ whose necessary and sufficient optimality conditions yield a closed-form solution ℓ* = A^{-1} b / λ, where A is the covariance matrix of the signature and b is its cross-covariance with the cumulative PnL — an exact analogue of the Markowitz weight vector but in the infinite-dimensional signature coefficient space. Drawdown control, momentum tilts, and mean-reversion strategies emerge as particular signature coefficients selected by the optimization, rather than requiring separate parametric models for each effect. The lead-lag embedding is the construction that enables the signature to capture causal predictability: without it, the signature of the joint path (S_t, Z_t) contains only contemporaneous cross-moments between price and signal, while the lead-lag version includes cross-iterated integrals of the form ∫_0^t Z_{s-Δ} dS_s that encode the lagged cross-correlation between signal and subsequent price move — the fundamental quantity that any momentum or mean-reversion strategy exploits. At signature degree 2 the resulting strategy includes terms proportional to ∫_0^t Z_{s-Δ} dS_s (signal-return cross-correlation), (S_t − S_0)² (variance of path), and ∫_0^t S_s ds (time-average of price, encoding mean-reversion), exactly recovering classical trend-following, variance-timing, and Bollinger-band strategies as special cases of the degree-2 signature strategy. At higher degrees the strategy captures non-linear interactions between signal and price history that cannot be expressed as any simple parametric strategy. The paper's technical elegance is that by reducing all of these effects to coefficients in a single linear functional on the signature, it provides a systematic and unified way to discover, represent, and estimate an optimal trading strategy from historical data without pre-specifying which combination of momentum, mean-reversion, and signal effects to include; the optimization jointly determines which effects are present and their optimal magnitudes, constrained only by the choice of signature truncation level N and the risk aversion parameter λ.
Notation / Conceptual Frame
Let S_t ∈ R^{d_S} be asset prices and Z_t ∈ R^{d_Z} be exogenous signals; the augmented path is Ẑ_t = (S_t, Z_t) ∈ R^{d_S + d_Z}. The lead-lag embedding Ẑ^{ll} : [0, 2n] → R^{2(d_S+d_Z)} on a discrete time grid {t_0, ..., t_n} is defined by linearly interpolating between the points (Ẑ_{t_k}, Ẑ_{t_{k-1}}) and (Ẑ_{t_k}, Ẑ_{t_k}), creating a piecewise-linear path whose signature contains the cross-iterated integrals ∫ ΔẐ^(1)_{t_k} ⊗ ΔẐ^(2)_{t_k} encoding the lagged covariance between the first and second copies of Ẑ. The position process is π_t = ⟨ℓ, 𝕊(Ẑ^{ll})_{0,t}⟩ ∈ R^{d_S} and the cumulative PnL is G_T = ∫_0^T π_t · dS_t. The mean-variance objective is J(ℓ) = E[G_T] − (λ/2)Var[G_T] = ⟨b, ℓ⟩ − (λ/2)⟨Aℓ, ℓ⟩ where b = E[𝕊(Ẑ^{ll})_{0,T} · G_T] ∈ T((R^D))^* and A = E[𝕊(Ẑ^{ll})_{0,T} ⊗ 𝕊(Ẑ^{ll})_{0,T}] ∈ (T((R^D))^*)^{⊗2} with D = 2(d_S + d_Z). The optimal coefficient is ℓ* = A^{-1} b / λ, computed by solving the linear system Aℓ* = b/λ in the truncated tensor algebra of dimension N_sig = Σ_{k=0}^N D^k.
Commentary
The Riesz representation theorem for Hilbert spaces is the foundational principle underlying the signature trading result: the map ℓ → G_T(ℓ) is a continuous linear functional on the Hilbert space completion of the tensor algebra under the A-inner product, and the Riesz theorem guarantees the existence of a unique element ℓ* in this Hilbert space representing the functional b, which is precisely the optimal coefficient vector. The quadratic programme in ℓ is therefore not merely a computational convenience but reflects the geometric structure of the mean-variance problem in the function space of path-dependent strategies. The condition number of A — the ratio of its largest to smallest eigenvalue — determines the stability of the optimal solution ℓ* = A^{-1} b/λ: a large condition number indicates near-collinearity of some signature features, which would produce large and unstable coefficient estimates that overfit the historical data. The Marchenko–Pastur analysis of A, treating the N_sig × n matrix of signature feature vectors as a large random matrix, predicts that the sample A has a substantial fraction of near-zero eigenvalues when N_sig / n is not small, exactly the situation arising in practice with high truncation levels N and limited historical data n; eigenvalue cleaning of A provides the same regularization as Tikhonov penalization of ||ℓ||² and yields a closed-form Ledoit–Wolf type shrinkage formula for ℓ* that does not require cross-validation. The interpretation of individual signature coefficients as recognizable trading strategies holds only at low degrees N = 1, 2: at degree 1 the strategy is a constant position (static holding), at degree 2 it includes all classical one-signal strategies, and at degree 3 it includes all pairwise signal interactions. Beyond degree 3 the individual coefficients lose interpretability and the strategy should be understood as a non-parametric combination of all possible lead-lag and higher-order interactions, whose behavior can only be characterized by simulating the strategy on test data rather than by reading individual coefficient values.
Implications for Research Methodology
For the desk's systematic quantitative strategies, the signature mean-variance framework provides a principled way to incorporate path-dependent signals — both price-based (momentum, mean-reversion, realized variance trajectory) and exogenous (news sentiment, volume anomalies, options market metrics) — into position sizing without the need to separately engineer feature interactions and tune position-sizing formulas for each combination. The offline computation of (A, b) from historical data requires O(n · N_sig) operations where n is the number of historical path observations and N_sig the number of signature coefficients, and the online position computation at time t requires a single inner product ⟨ℓ*, 𝕊(Ẑ^{ll})_{0,t}⟩ computable in O(N_sig) time, making the live strategy inference fast enough for intraday rebalancing at any desired frequency. The unified handling of the risk aversion parameter λ — which appears as a single scalar in the closed-form solution ℓ* = A^{-1} b/λ — enables clean position scaling: the expected PnL and variance of the strategy are E[G_T] = ⟨b, ℓ*⟩ = ||b||²_{A^{-1}} / λ and Var[G_T] = ⟨Aℓ*, ℓ*⟩ = ||b||²_{A^{-1}} / λ², so the Sharpe ratio is ||b||_{A^{-1}} / (λ^{1/2} · Var^{1/2}) = ||b||_{A^{-1}} / λ^{1/2}, which is monotone in the signal strength ||b||_{A^{-1}} and decreasing in λ; targeting a specific Sharpe or position volatility determines λ analytically as λ = ||b||²_{A^{-1}} / (target Sharpe)², providing a transparent risk-management dial.
Limitations
The estimation of the second-moment tensor A from historical data is the primary statistical bottleneck: for D = 2(d_S + d_Z) = 8 (two assets and two signals, each duplicated in the lead-lag embedding) and N = 4, N_sig = Σ_{k=0}^4 8^k = 4681, yielding a 4681 × 4681 matrix requiring approximately n >> 4681 observations for reliable inversion, which at daily frequency corresponds to more than 18 years of data. The ridge regression solution ℓ* = (A + μI)^{-1} b/λ reduces the effective dimensionality via regularization but introduces a bias whose magnitude grows with μ and whose direction depends on the unknown true ℓ* in a circular way; optimal μ from cross-validation depends on the test-set distribution, and non-stationarity means that the test-set performance of ℓ*_μ on the most recent data may be poor regardless of cross-validation on the historical window. The stationarity assumption — that the joint distribution of (Ẑ_t)_{t∈[0,T]} is time-stationary, enabling consistent estimation of (A, b) from historical data — is violated by most financial paths at multiple timescales: structural breaks in market microstructure, volatility regime changes, and slow drift in signal-return predictability all induce non-stationarity that causes the historically estimated ℓ* to be suboptimal for the current period. The rolling re-estimation of (A, b) over a sliding window of length L attempts to track the non-stationarity but introduces a lookback horizon parameter L whose optimal value trades off against variance (short L → high variance) and bias (long L → includes stale data), a trade-off with no theoretically optimal solution absent a parametric model for how (A, b) evolve in calendar time.
- Mesh-Free Solution of Path-Dependent PDEs in a Signature-Kernel RKHS· Technical Commentary
- Random-Matrix Limits on the Information Content of Empirical Correlation Matrices· Methodological Annotation