master backtest // crypto & us stocks

indicators · assets · one rule each

Does any TA-Lib rule beat buy-and-hold across crypto and US stocks?

Every applicable TA-Lib function — moving averages tested at several periods each, plus combinations of them — turned into a single long/flat/short rule and backtested independently on 10 crypto pairs and 20 US mega-caps, at seven timeframes from daily down to one minute. Ranked by information ratio against buy-and-hold on the same asset, and judged against four acceptance gates. Read the methodology at the bottom before trusting any of this.

Top 10 by total PnL

click a card to inspect

Cumulative PnL — top 8 vs. buy & hold

portfolio-level, $10k/asset, summed across the universe

Full leaderboard

click any column to sort · click a row for detail

Indicator detail

Methodology

Why this is ranked by information ratio, not by profit

Every rule here is scored on information ratio against buy-and-hold on the same asset — not on dollars earned, not on raw Sharpe. Those two reward something that looks like skill and is not.

IR = mean(r − rbuy&hold) ÷ std(r − rbuy&hold) × √(periods per year)

Read it as average lead ÷ how much that lead wobbles. Two rules can finish equally far ahead; the one that got there steadily is skilled, the one that got there in a single lucky year is not. Dividing by the wobble separates them. Doing nothing different from the benchmark scores exactly zero.

What profit cannot tell you

This project has already been caught by exactly that trap: an earlier sweep in the same repo produced “winners” that turned out to be long-bias in a rising market, because it ranked on dollar PnL and raw Sharpe. The scoreboard cards still display dollars, but nothing is ordered by them.

The four gates

Ranking is not the same as passing. The leaderboard orders candidates by IR, but a candidate only counts as an edge if it clears all four gates at once — out of sample, at the real fee schedule. Each gate kills a different way of being fooled, which is why one strong number is never enough.

Notice that each gate has a “too good” band. A result can fail by being too strong: an IR above 2, perfect breadth, or a t above 6 on a sample this short is far more likely to be a look-ahead leak, an accounting error or a stale-data artifact than a discovery. Every one of those has happened in this repo. Treat a sudden winner as a bug until the parity harness and the multiplicity correction have both been re-checked.

Why the gates are this hard

The governing constraint is t = IR × √years, and it is unforgiving. On an 11-year sample √11 ≈ 3.3, so even a genuinely good IR of 0.5 reaches only t ≈ 1.7 — below the bar. A system can be real and still not provable on the history that exists.

Two consequences follow, and neither is optional:

Only two levers actually move these gates, and neither is “try more indicators”: more history (t scales with √years) and lower turnover (fee headroom scales inversely with it).

Built by Gnourt · algorithmic trading systems