QTJ Editorial · 27 April 2026 · Backtesting & Methodology
Public trader-ranking sites can over-represent currently active records. A transparent synthetic example illustrates the survivorship-bias mechanism without claiming access to a private audit dataset.
A research archive on the quantitative methods used to tell skill from luck in a trading record — the statistics that actually validate trader performance, what a meaningful Sharpe ratio looks like for an independent trader, and where competition and audit data fall short. 35 articles span systematic strategy design, risk management, backtesting methodology, and championship-grade trader analysis.
We apply a Gaussian HMM to classify volatility regimes across equity-index and commodity futures markets and evaluate whether regime-conditional position sizing improves risk-adjusted returns over a naive baseline.
Standard VaR models underestimate joint tail risk in leveraged portfolios. We use vine copulas to model dependence structures across commodity and index futures during 2020–2025 stress events.
A worked example shows how survivorship bias can distort signal-provider directories, using an illustrative cohort rather than a claimed proprietary QTJ dataset.
We update our rolling correlation analysis across equities, bonds, commodities, and FX through Q3 2025, finding that correlation breakdown during the August volatility event was more severe than models predicted.
Full Kelly is theoretically optimal but practically dangerous. We compare half-Kelly, fractional Kelly, and a Bayesian Kelly variant across 10,000 simulated equity curves with realistic transaction costs.
Do on-chain metrics predict short-term price movements? We test 14 commonly cited indicators using Granger causality and out-of-sample forecasting on BTC and ETH daily data from 2020 to 2025.
Walk-forward analysis is widely recommended but poorly implemented. We catalogue seven failure modes observed in practitioner backtests and demonstrate their impact on reported Sharpe ratios.
We examine calibration, liquidity, and arbitrage bounds across three major prediction markets from 2022 to 2025, finding persistent mispricings in low-liquidity contracts that decay with a half-life of approximately 48 hours.
Maximum drawdown alone is insufficient for strategy evaluation. We compare five drawdown metrics — including Calmar, Ulcer Index, and conditional drawdown-at-risk — and assess their discriminative power across strategy types.
We develop a regime-aware mean reversion framework for G10 currency pairs using Ornstein-Uhlenbeck parameter estimation and evaluate signal generation quality across trending and range-bound environments.
Block bootstrap is the default resampling method in strategy evaluation, but it destroys serial dependence. We compare block bootstrap, stationary bootstrap, and parametric Monte Carlo on 12 systematic strategies.
Portfolio diversification relies on stable correlations, which fail precisely when they matter most. We quantify correlation regime shifts across seven major stress events and evaluate dynamic hedging responses.
Momentum premia vary significantly by holding period and look-back window. We decompose returns across 1-day to 12-month horizons in equity indices, FX, and commodities, isolating the contribution of each timescale.
How much data is needed to distinguish skill from luck in trading? We compute the power of standard hypothesis tests at various sample sizes and show that most track records shorter than three years are statistically uninformative.
We analyse publicly available return data from major independently audited trading competitions including the World Cup Trading Championships, examining risk-adjusted performance, consistency across years, and survivorship in the winner cohort.
Constant-volatility targeting is a simple rule that materially improves risk-adjusted returns across asset classes. We examine its interaction with momentum signals and evaluate optimal target levels for different risk budgets.
Carry returns can be decomposed into interest rate differential, spot return, and roll yield components. We examine which components dominate across rate cycles from 2015 to 2024 and test conditional carry strategies.
We analyse the publicly available monthly return series of the 2023 Trading World Champion, examining Sharpe ratio, maximum drawdown, drawdown recovery characteristics, and what the risk-adjusted profile suggests about the underlying strategy.
Point estimates of strategy parameters ignore uncertainty. We demonstrate Bayesian estimation of mean return, volatility, and Sharpe ratio using MCMC, producing credible intervals that give a more honest picture of expected performance.
We document statistically significant time-of-day effects in ES, NQ, CL, and GC futures using five years of tick data. Certain 30-minute windows show persistent directional biases that survive transaction cost adjustment.
Pairs trading strategies based on cointegration have shown declining profitability since the mid-2010s. We test whether structural breaks in cointegration relationships explain the deterioration and evaluate adaptive re-estimation methods.
We demonstrate that Gaussian-based position sizing systematically underestimates tail risk exposure. Using Student-t and stable Paretian fits, we derive adjusted position sizes that account for empirical fat tails in futures returns.
Retail traders face different microstructure frictions than institutional participants. We quantify the impact of spread widening, partial fills, and queue priority on strategy performance for account sizes under $500k.
Combining multiple weak signals can produce a stronger composite. We compare equal weighting, inverse-volatility weighting, and machine learning ensembles on a universe of 20 systematic FX signals.
Value-at-Risk remains the industry standard despite well-known deficiencies. We compare VaR and Expected Shortfall in terms of backtest performance, regulatory capital requirements, and practical risk management utility for leveraged futures portfolios.
Testing many strategy variants on the same data inflates the probability of finding a spuriously profitable result. We review White's Reality Check, Hansen's SPA test, and the Bonferroni correction in the context of strategy selection.
A stylised Level 2 futures-data scenario illustrates order-flow imbalance in ES and NQ across short horizons, including the importance of execution costs.
A worked example shows how spread, slippage, and market impact can be parameterised from published market-microstructure ranges instead of a single fixed-cost assumption.
Trend following strategies are often marketed as providing "crisis alpha." We revisit this claim using CTA index data through 2023, finding that crisis alpha is concentrated in specific sub-strategies and is less reliable than commonly assumed.
Maximum drawdown is the most psychologically salient risk metric for traders, yet its distribution is rarely estimated. We derive confidence intervals for maximum drawdown using the Bai-Perron framework and Monte Carlo simulation.
The implied volatility surface encodes market expectations of future price distributions. We examine whether risk reversal skew and butterfly spreads contain exploitable information for systematic FX strategies.
Feature selection is the most consequential modelling choice in ML-based trading. We compare SHAP values, permutation importance, and Boruta selection on a gradient-boosted model trained on 200+ technical and fundamental features.
The post-2022 rate environment has changed the return dynamics of many systematic strategies. We decompose the rate sensitivity of momentum, carry, and mean reversion strategies across asset classes and assess implications for portfolio construction.