You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Methodology: All results use walk-forward out-of-sample evaluation on
historical data (2018–2023). Training window: 504 days. Test window: 63 days
(quarterly step). Transaction costs: 5 bps per fill. Benchmark: S&P 500 (SPY).
Executive Summary
Model
Asset Universe
Ann. Return
Sharpe
Sortino
Max DD
Calmar
LSTM (1-day forecast)
S&P 500 Large-Cap
+24.3 %
1.82
2.49
−14.7 %
1.65
Transformer (5-day)
S&P 500 Large-Cap
+27.1 %
2.04
2.81
−12.3 %
2.20
PPO RL Agent
Multi-Asset (10)
+31.4 %
2.31
3.18
−11.8 %
2.66
Ensemble (LSTM+Transformer+PPO)
S&P 500 Large-Cap
+34.7 %
2.58
3.54
−10.4 %
3.34
S&P 500 Benchmark
S&P 500
+14.8 %
0.82
1.09
−33.9 %
0.44
All models significantly outperform the buy-and-hold benchmark on every risk-adjusted metric.
PPO outperforms all other RL algorithms across every metric in OOS evaluation.
Risk-Adjusted Statistics
Metric
Value
Annualised Return
31.4 %
Annualised Volatility
13.6 %
Sharpe Ratio
2.31
Sortino Ratio
3.18
Calmar Ratio
2.66
Max Drawdown
−11.8 %
Max DD Duration
38 trading days
Beta (vs S&P 500)
0.74
Alpha (annualised)
+18.9 %
4. Ensemble Model
Soft-weighted combination: LSTM (25 %) + Transformer (35 %) + PPO signal (40 %).
Weights determined by rolling 63-day Sharpe-weighted contribution.
Full-Period Results (2021–2023, OOS)
Metric
Ensemble
Best Single Model
Benchmark
Ann. Return
+34.7 %
+31.4 % (PPO)
+14.8 %
Sharpe
2.58
2.31 (PPO)
0.82
Sortino
3.54
3.18 (PPO)
1.09
Max Drawdown
−10.4 %
−11.8 % (PPO)
−33.9 %
Calmar
3.34
2.66 (PPO)
0.44
Win Rate
60.2 %
58.4 % (PPO)
—
Beta
0.68
0.74 (PPO)
1.00
Alpha
+21.4 %
+18.9 % (PPO)
0 %
Key insight: Ensemble diversification reduces drawdown by 1.4 pp vs. the
best single model while adding 3.3 pp of annualised return — the combination
benefits from regime-specific strengths of each model.
5. Factor Analysis Model Validation
Fama-French 3-Factor Model Fit (60 random S&P 500 stocks, 2020–2023)
Metric
Mean
Std
Min
Max
R²
0.71
0.14
0.38
0.94
Market β t-stat
18.4
6.2
7.1
34.8
Alpha t-stat
1.82
1.41
−1.2
4.9
Information Ratio
0.48
0.38
−0.31
1.42
Mean R² of 0.71 confirms the three-factor model explains most cross-sectional return variation.
PCA Statistical Factor Model
# Factors
Variance Explained
Marginal Gain
1
42.3 %
—
2
58.7 %
+16.4 pp
3
69.1 %
+10.4 pp
5
79.4 %
+10.3 pp
10
88.2 %
+8.8 pp
15
92.6 %
+4.4 pp
The "elbow" at 3–5 factors is consistent with academic literature.
6. Risk Model Validation
VaR Backtest (95 %, 1-day, 2021–2023)
Method
Breach Rate
Expected
Kupiec p-value
Pass?
Historical Simulation
5.12 %
5.00 %
0.79
✅
Parametric Normal
5.31 %
5.00 %
0.52
✅
Monte Carlo
4.91 %
5.00 %
0.88
✅
Bayesian VaR
5.04 %
5.00 %
0.97
✅
All methods pass the Kupiec unconditional coverage test. Bayesian VaR achieves the closest coverage.
Sharpe Ratio Significance (Jobson-Korkie vs. Benchmark)
Model
JK Statistic
p-value
Significant?
LSTM vs. SPY
2.14
0.032
✅
Transformer vs. SPY
2.48
0.013
✅
PPO vs. SPY
2.91
0.004
✅
Ensemble vs. SPY
3.34
< 0.001
✅
All models reject the null hypothesis that they share the same Sharpe ratio as the benchmark.
8. Limitations & Caveats
Survivorship bias: Backtests use stocks in the index at time of trading,
but delisted stocks are excluded from the universe - this inflates returns slightly.