The measures, what each treats as risk, and what each cannot see
MeasureWhat it treats as riskWhat it cannot see
Sharpe ratioTotal volatility of the return above a cash benchmark.Skew, fat tails, the shape of the losses, and the size of the search that produced the number.
Sortino ratioOnly the volatility below a target return you choose.The same blind spots, minus the upside penalty. It is also estimated from fewer observations, so it is noisier.
Calmar ratioThe single worst peak-to-trough drawdown in the period.Everything except that one episode. It also looks worse the longer the sample runs, by construction.
Information ratioDeviation from a chosen benchmark, rather than from cash.Whether the benchmark was the right thing to be measured against in the first place.
Deflated Sharpe ratioThe number of variants tried before this result was found.Leaked data and unmodeled costs. It corrects for selection, not for mistakes.

How to compare two strategies on a risk-adjusted basis

  1. Put both on the same frequency and the same period

    Daily against daily, and the same start and end dates. A ratio computed on monthly returns is not comparable to one computed on daily returns, even after both are annualized, and overlapping-but-different periods quietly compare two different market regimes.

  2. Apply the same costs to both

    Spread, commission, slippage and financing, on identical assumptions. A high-turnover strategy and a low-turnover one are not being compared at all until this is done, and whichever was given kinder assumptions wins for free.

  3. Choose the measure that matches the risk you care about

    If large losses are the concern, downside-based measures say more than the Sharpe ratio. If sitting through a long loss is the concern, look at the drawdown profile rather than any ratio.

  4. Compare the trial counts, not just the ratios

    Ask how many variants were tested to arrive at each figure. A strategy that survived one test and a strategy that survived fifty thousand are not comparable on the raw number, and this step is the one almost nobody performs.

  5. Check the sample is long enough to tell them apart

    Two ratios that differ by 0.2 over eighteen months are not distinguishable from each other. Before concluding one strategy beats another, establish that the gap is larger than the noise in the estimate.

Why not just compare returns?

Because return alone can be bought with risk, and buying it is trivial. Double the position size and the return roughly doubles; so does the loss when it comes. Any two strategies can be made to show the same return by adjusting leverage, which means the raw figure carries almost no information about which one is better.

A risk-adjusted return asks the question that survives that adjustment: for each unit of whatever we are calling risk, how much did this earn? That figure is roughly invariant to leverage, which is what makes it comparable at all.

The catch is in "whatever we are calling risk". Every measure below encodes a different answer, and the answer is a modeling choice rather than a fact.

Which risk-adjusted return measure should you use?

Match the measure to the failure you actually care about, rather than reaching for whichever one the platform shows by default.

  • Comparing broadly, or reporting to someone who will compare you to others: the Sharpe ratio, because it is the one everybody computes. Its ubiquity is a real advantage even where its assumptions are poor.
  • The strategy has occasional large gains: the Sortino ratio, which stops treating those gains as risk. The gap between the two numbers is itself informative.
  • You are worried about how deep the hole gets: the Calmar ratio, or better, look at the drawdown distribution directly rather than compressing it into one figure.
  • You are measured against an index rather than against cash: the information ratio.
  • The number came out of a search over many variants: the deflated Sharpe ratio, which is the only one on this list that addresses that at all.

In practice, reporting two or three of these together is more honest than picking the flattering one. A strategy that looks good on every measure is a different proposition from one that looks good on exactly the measure its author chose to report.

What do all of these measures have in common?

They are all backward-looking statistics computed on a return series, and none of them knows how that series was chosen.

This is the blind spot that matters, and it is shared by every measure in the table. Each one is a valid description of what happened. None of them can distinguish between a strategy that was tested once and worked, and a strategy that was the best of fifty thousand attempts on the same data — the arithmetic is identical, and so is the output.

That distinction is not a detail. The first is evidence. The second is roughly what you would expect noise to produce given a large enough search, and it will look exactly as good on every ratio in the table.

The deflated Sharpe ratio exists specifically to break that tie, by asking how high the figure would need to be before it is surprising given the number of trials. It is the only measure here that treats the search as part of the evidence.

Why do two sources report different risk-adjusted returns for the same strategy?

Almost always because of an undisclosed convention rather than an error. Four differences account for most of it.

  • Observation frequency. Daily and monthly figures differ even after annualizing, because annualizing assumes returns are independent from one period to the next and they usually are not.
  • The benchmark. One source subtracts a cash rate, another subtracts zero, a third subtracts an index. In a high-rate environment, subtracting zero flatters every result.
  • The cost assumptions. Gross-of-cost and net-of-cost figures for a high-turnover strategy can differ by more than the entire ratio.
  • The period. A window that starts after a crash and ends before the next one produces a better number than the same strategy measured across both.

None of these is dishonest on its own. What makes them a problem is that the convention is usually not stated, so two numbers get compared as if they were measuring the same thing.

Common questions

What is a good risk-adjusted return?

There is no threshold that transfers across measures and contexts. For a live, multi-year Sharpe ratio, convention treats above 1 as good. For any figure taken from a backtest, the bands do not apply, because they assume the number was measured rather than selected from many attempts.

How do you calculate a risk-adjusted return?

Divide the return above a benchmark by a measure of risk. Which measure of risk is the whole question: total volatility gives the Sharpe ratio, downside volatility the Sortino ratio, worst drawdown the Calmar ratio, benchmark tracking error the information ratio.

Does risk-adjusted return account for drawdown?

Only the Calmar ratio does directly. The Sharpe and Sortino ratios use volatility, which is related to drawdown but not the same thing — two strategies with identical Sharpe ratios can have very different worst losses and very different recovery times.

Is a higher risk-adjusted return always better?

No. In a backtest, past a certain point a higher figure is better evidence that something is wrong — unmodeled costs, leaked data, or a large uncounted search — than that something was found. It also says nothing about whether the strategy has capacity at a size worth trading.

What is the difference between risk-adjusted return and alpha?

A risk-adjusted return divides return by a risk measure. Alpha is the return left over after subtracting what a benchmark or a factor model would have produced given the same exposures. They answer different questions and are not substitutes: a strategy can have a good Sharpe ratio and no alpha, if its returns are explained by an index it could have bought.