Definition
What is the deflated Sharpe ratio?
In short
The deflated Sharpe ratio is the probability that a strategy's true Sharpe ratio is above zero after allowing for how many strategy variants were tried before this one was chosen, for how long the track record is, and for skewed or fat-tailed returns. It was introduced by David Bailey and Marcos López de Prado in 2014. The idea is simple: when a researcher tests many variants and reports the best, the best will look good even if none of them has any skill. The deflated Sharpe ratio raises the bar from zero to the Sharpe ratio that luck alone would be expected to produce, and asks how confidently the reported figure clears it.
Why an ordinary Sharpe ratio overstates a backtest
A Sharpe ratio measures return per unit of volatility. It says nothing about how the strategy was found. A team that tries one hundred parameter sets, signals or holding periods and publishes the best has bought one hundred lottery tickets, and the winning ticket's Sharpe ratio reflects the search as much as the strategy.
Bailey and López de Prado give the expected maximum Sharpe ratio of N independent strategies with zero true skill. For annualised figures from ten-year backtests, the best of 10 zero-skill variants shows about 0.50 on average, the best of 100 about 0.80 and the best of 1,000 about 1.03. Shorter records make the problem worse: the best of 1,000 five-year backtests shows about 1.46.
None of those strategies can predict anything. The numbers are the height that pure selection reaches.
How the deflated Sharpe ratio is calculated
Step one is the benchmark. Instead of testing the observed Sharpe ratio against zero, the test uses the expected maximum Sharpe ratio across the N trials, which depends on N and on how much the Sharpe ratios of those trials vary.
Step two is the probabilistic Sharpe ratio against that benchmark. The difference between the observed Sharpe ratio and the benchmark is scaled by the square root of the number of observations minus one, and divided by a term that grows with negative skew and with excess kurtosis. Fat left tails therefore lower the result.
The output is a probability between 0 and 1. A common reading is that a value above 0.95 indicates the strategy is unlikely to be a product of selection alone.
A worked example
Take a backtest with an annualised Sharpe ratio of 1.0 over ten years of monthly returns, with normally distributed returns. Tested against zero as if it were the only strategy ever tried, the probability that its true Sharpe ratio is positive is above 0.99.
Now suppose it was the best of 10 independent variants. The benchmark rises to about 0.50 and the deflated Sharpe ratio falls to about 0.94. As the best of 100, the benchmark is about 0.80 and the deflated Sharpe ratio is about 0.73. As the best of 1,000, the benchmark is about 1.03, above the observed figure, and the deflated Sharpe ratio is about 0.46: no better than a coin toss.
With a skew of minus one and a kurtosis of six, the best-of-100 case falls further, to about 0.70. The same backtest can be convincing or meaningless depending on facts that rarely appear in a marketing document.
These figures assume independent trials and use the authors' published approximation. Correlated variants count as fewer effective trials, which is why the number of trials has to be estimated rather than simply counted.
What to ask a strategy provider
How many variants, parameter sets and signals were tested before this one was selected, and is there a record of them? Without that number the deflated Sharpe ratio cannot be calculated and the reported Sharpe ratio cannot be judged.
How long is the record, and what are its skew and kurtosis? A Sharpe ratio of 0.5 needs about eleven years of monthly data before it can be told apart from zero at 95% confidence.
Is any of the record out of sample or live? Selection bias affects every figure computed on data that was used to choose the strategy. Returns earned after the rules were frozen are the evidence it cannot touch.
What it does not tell you
A high deflated Sharpe ratio does not show that a strategy will make money in future. Markets change, costs rise with capacity, and published signals lose much of their return after publication.
It also depends on an honest trial count. A provider that understates how many variants it tried will report a deflated Sharpe ratio that is too high, and an outsider has no way to check.
Treat it as a filter that removes backtests luck can explain, not as a measure of skill.
Related
Put the question to every backtest.
The due-diligence checklist turns this into questions for any provider, including the one we introduce.