Quantitative Trading5 min read

The Benchmark I Never Subtracted

S

Suneet Malhotra

Jul 25, 2026

β€’
1 views
The Benchmark I Never Subtracted - Quantitative Trading blog post
πŸ”§PythonπŸ”§BacktestingπŸ”§Statistics

A Sharpe ratio computed against cash answers a question I should not care about. It asks whether my returns beat Treasury bills. For a book that is long equities most of the time, that test passes in any year the tape went up, which is most years. The arithmetic is right and the number is real and it is still the wrong test, because the null hypothesis for a long equity strategy is not zero. It is buy and hold.

I have published a Sharpe number on this blog. I have never once published the same number for the index over the same window, which is the only comparison that would have made mine mean anything.

Beta is not skill

Write any strategy return as a market piece plus a residual:

r = alpha + beta x m + e

Beta is how much market I am holding. Alpha is what is left after the market has been paid for. A momentum book that buys strength in a rising index earns most of its return through beta by construction, and beta is available for free, in one ticker, with no scheduler, no signal gate, and no slippage band.

The uncomfortable part is what happens to the Sharpe ratio under pure exposure. Take a book with a beta of 0.6 and an alpha of exactly zero. Excess return is 0.6 times the market excess return. Volatility is 0.6 times market volatility. The ratio is identical to the index. Scaling exposure moves the numerator and the denominator by the same factor, so leverage cannot buy a better Sharpe and neither can sitting in cash.

That has a consequence I find clarifying. Any Sharpe above the market Sharpe has to come from somewhere other than how much I held. Selection, timing, or luck. Those are the only three doors and two of them are skill.

Cash drag cuts the other way

The lazy version of this critique is to line raw return up against the index and call a partially invested book a failure. That is also wrong. A strategy that sits in cash forty percent of the month will lose a raw return contest almost every time while carrying less risk, and comparing returns at different risk levels is not a comparison at all.

But being flat part of the time is not free either. If my timing carries no information, moving exposure around adds variance to the return series without adding mean. Random beta is noise wearing a strategy costume. It makes the ratio worse than the index rather than equal to it. So the honest reading of an uninformative timing overlay is not neutral, it is slightly negative, and the drag is easy to miss because it shows up as a wider distribution rather than a smaller average.

The benchmark is a free parameter

Here is the part that should bother anyone who has ever quoted an alpha figure. The benchmark is a choice, and almost nobody makes that choice before looking at the returns.

The broad index, the tech heavy index, the equal weighted version, or the single sector I actually traded. Those four pick out visibly different numbers in a year where a handful of large caps carried the tape. In such a year a book concentrated in those names shows enormous alpha against the equal weighted index and close to none against the tech heavy one. Same returns, same window, opposite conclusion, and the thing that decided it was a selection I made afterward.

I have written before about pre-registering the numbers I intend to judge a system by, before the data exists. The benchmark belongs on that list and it was not on mine. A comparison chosen after the outcome is not a test. It is a search for the frame that flatters.

What this does to the error bar

A week ago I wrote that at a true Sharpe of one it takes roughly four years of returns before the estimate separates from zero with any confidence. Alpha is worse. To estimate alpha I have to estimate beta first, and that second estimate carries its own uncertainty which propagates into everything downstream. The four year figure was the optimistic case for a raw ratio against cash. Against a benchmark, with beta fitted from the same short sample, the interval is wider still.

So the correct summary of most of my own record is not that the edge is small. It is that I have been grading against the wrong answer key, and the right one takes longer to read.

Zero is a comfortable null because nobody owns it and nothing publishes it. Buy and hold has a ticker, a daily close, and a number I can look up any morning I want to know how much of my curve I built and how much of it was handed to me.

Share this post

You Might Also Like

Stay in the Loop

Get weekly insights on AI-driven QA, engineering leadership, and automation strategies.

No spam, ever. Unsubscribe anytime.