We publish what didn't work.
Most of what a research desk tries fails. This is the ledger of ours: each hypothesis we pre-registered, the exact threshold it had to clear, and the measured result that buried it — or, in a few cases, the result that passed and still sits dark because we could not prove it twice. Everything here is a backtest or a paper measurement. There is no real-money record.
The two that taught us the most
Dealer gamma (GEX) as a return signal
VIX term-structure stress scaler on the SPY core
The full ledger
Every hypothesis, its pre-registered threshold, and the measured result — killed, validated but held dark, or still open pending data. Where a number was never measured, the cell says so.
| Hypothesis | Basis | Pre-registered threshold | Measured result | Verdict |
|---|---|---|---|---|
Dealer gamma (GEX) as a return signal #49 | BACKTEST | Return spread ≥ 2%/yr after a 50% post-publication haircut AND |t| ≥ 2.0 on non-overlapping windows. | A naive +15.9%/yr at t = 5.45 collapsed to t = 0.36 once the overlapping 21-day windows were removed. The vol signal was real (+73.6% spread, t = 3.56) but is not monetizable at a monthly cadence; the economic overlay lost Sharpe outright. | Killed |
VIX term-structure stress scaler on the SPY core #51 | BACKTEST | Full-period Sharpe ≥ 0.94, maxDD better by ≥ 3.0 pts, ≤ 6 extra trades/yr, and no sub-period materially worse. | Sharpe 0.893 → 0.971 and maxDD −33.8% → −25.3% — the gate passed on every clause. But it only fires in a genuine backwardation/stress episode; the tape is in contango, so it has never fired. Shipped dark. | Validated · dark |
Bond-vol × VIX combined scaler #48 | BACKTEST | Beat the VIX-only scaler by ≥ +0.03 Sharpe OR ≥ +2.0 pt maxDD, worse on neither metric, ≤ 3 extra trades/yr. | Sharpe 1.034, maxDD −19.0% (+0.062 over VIX-only), passed both gates and improved exactly where the hypothesis predicted (2022). But the circular-shift timing null came back at p ≈ 0.06, so we cannot claim the timing is what works. Held dark, forward-soaking. | Validated · dark |
Multi-speed trend sleeve (Quantica / Man Group speed factor) #47 | BACKTEST | Full-period Sharpe ≥ 0.94 (baseline 0.89 + 0.05), maxDD better by ≥ 3.0 pts, stable across all three sub-periods. | The 3–6 month speed band was directionally right and improved every sub-period, but at +0.035–0.041 Sharpe it sat under the +0.05 gate and inside noise. Separately confirmed: no sleeve-only change can move portfolio maxDD (best +0.6 pt vs +3.0 pt required) — the drawdown is core-driven. | Killed |
VRP / put-write income sleeve Jul 10 R&D | BACKTEST | Must diversify the crisis-alpha book — i.e. carry low or negative correlation to the SPY core. | Earned a real 7.4% CAGR (Sharpe 0.73, maxDD −32.7%) but at +0.86 correlation to SPY and −13.4% in the COVID month. Equity-like income, not a diversifier — adding it would amplify the left tail, the opposite of the book's identity. | Killed |
ML / RL direction models (MARL, DQN, CCR) System 1 | PAPER | Must beat chance out-of-sample to be allowed to drive direction. | Anti-predictive out-of-sample — the correctness flag reads inverted and confidence is miscalibrated. 96 feature cells below chance; MARL ~96.7% hold-biased, near-random. Kept as a weight-0 shadow for research; never lets them drive a trade. | Killed |
MARL behavioural consensus as a vote in the decision System 1 · weight sweep | PAPER | Must change the outcome at all. A vote weight that never alters a decision cannot be promoted, because there is nothing to measure. | A 500-configuration weight sweep moved the MARL vote weight across 0.00 / 0.05 / 0.10 while holding every other weight and the entry threshold fixed. Trades taken, wins, losses, P&L and Sharpe came back identical in 125 of 125 comparison groups — the weight changed nothing, anywhere. The same sweep over the ML weight changed the outcome in 164 of 192 groups, so the harness does discriminate. Cause: the four agents vote hold ~96.7% of the time, and a hold vote is rewarded 0.0 while the weighting formula counts only a positive result as a win — so a hold-biased agent decays to the 0.1 weight floor whether or not holding was the right call, and the consensus almost never reaches a directional vote. The agents are not untrained — the deployed DQNs carry 50 epochs and ~22.9M steps — but their recorded per-epoch training reward is negative and drifts further negative across the run, which is consistent with a vote that carries no information. Sample caveat, stated plainly: the sweep only ever took 11–14 trades, far too few to rank configurations. The finding is that the weight is inert, not that a better-trained version would be. | Killed |
A high-volatility regime edge (intraday) System 1 | PAPER | Must persist out-of-sample rather than flip between halves of the sample. | The walk-forward test killed it: the regime edge flips sign between the two time halves, and the high-vs-low-vol per-trade difference is statistically zero. An earlier −$438/60d read was traced to a corrupted P&L column — good that we tested before building on it. | Killed |
Loosen the entry threshold (min_strategies = 1) System 1 | BACKTEST | Must not lose money versus the current threshold. | Bled execution cost. Rejected. | Killed |
Exempt the edge strategies from the regime block System 1 | BACKTEST | Must add net return. | Lost. ny_liquidity actually needs the high-vol regimes the block was removing, so a broad exemption hurt. | Killed |
Split the daily-trade quota to free up more trades System 1 | BACKTEST | Must add net return. | Net negative. Rejected. | Killed |
Multi-timeframe (MTF) confluence engine as a profit lever MTF audit | PAPER | Must actually fire on the live path and add return. | ~70% dead code. The M15/H1/H4/D1/W1 confluence cascade is only reachable from the decommissioned MT5 endpoint — silent no-ops, 0 fires in 30 days versus 1,439 for the crude DB drift filter. It is a risk filter, not a profit lever. | Killed |
Multi-timeframe trend alignment as a live signal (the MTF shadow) #55 | PAPER | Pre-registered before the first observation: at least 200 scored observations, hit rate ≥ 0.55, edge ≥ 2.0 bps versus unaligned signals, and hit rate ≥ 0.50 in BOTH halves of the date range. The scorer refuses to print a verdict below the 200 floor. | The sample completed and the gate resolved on 19 Aug 2026: 671 signals recorded, 630 scored, 371 of them fully aligned across H1/H4/D1. Hit rate 0.469 against a 0.55 gate. Both halves 0.458 and 0.479, each below the 0.50 floor. Only the edge clause passed, at 11.8 bps. VERDICT: FAIL — a fully-aligned multi-timeframe signal did not predict direction better than a coin flip, in either half of the sample. The shadow book touched no order path at any point, so nothing needs unwinding. Per the rule written into the scorer: a negative result gets written up, not retuned and re-read on the same data. | Killed |
ny_liquidity high-volatility gate #49 catalyst | PAPER | Enable only after ≥ 30 fresh post-Jun-22 distinct signals confirm the high-vol / low-vol win-rate split. | In-sample win rate 0.867 in high-vol regimes versus 0.400 in low-vol / ranging, OOS p = 0.014 — the one statistically clean micro-finding. Shipped dark behind NY_LIQ_HIGHVOL_ONLY; the enable decision is a re-runnable validator, not a hunch. | Validated · dark |
Intraday price-bar signals on US equities cost model · 20 Aug 2026 | PAPER | The gross edge per trade must exceed a realistic round-trip cost. Nothing else matters: an edge inside the cost band is not a small edge, it is a negative one. | Charging a per-asset-class cost model against the round-trip ledger settled it. 113 equity round-trips produced $262.25 gross — a gross edge of 1.96 bps of traded notional — against an estimated 4.0 bps round-trip cost ($535.14 charged), for a net of −$272.89 at a 40.7% win rate. The June edge audit had put realistic cost at 2–7 bps; the measured edge sits inside that band, which makes the book negative at any sample size rather than underpowered. The same ledger cut by holding period says it twice more: trades held under an hour lost −$1,164 at a 24.2% win rate, and every intraday bucket is negative, while the 94 trades held beyond a day made +$7,937 at 61.7%. Verdict: killed, and on the do-not-revive list — no data purchase moves a 1.96 bps edge above its cost floor. | Killed |
Futures-tier trend sleeve (back-adjusted continuous futures) Tranche 2 | BACKTEST | Net Sharpe must beat the ETF blend (0.89) by ≥ 0.10 AND be stable across sub-periods. | not published — the harness is built and self-tested, but a decisive run needs paid back-adjusted daily bars (Databento / Norgate). The free-data run is inconclusive by design. | Open · awaiting data |
The intraday agent, frozen
The intraday agent runs 24 price-bar strategies on a paper account. On Jul 2 2026 we froze all alpha research on it — not because it broke, but because the June research program's own conclusion was that price-bar and ML alpha is exhausted: 96 feature cells below chance, the ML models anti-predictive out-of-sample, the MARL policy near-random. The engineering — idempotency keys, circuit breakers, watchdogs, 90+ tests, a three-layer security model — is a genuine asset. The edge is not.
It keeps trading paper on autopilot, capped at a few hours a week of ops. Two catalysts stay armed — the ny_liquidity high-vol gate and a VIX-backwardation trigger — and the freeze has a pre-registered review date of Dec 15 2026. It reopens only on a genuinely new out-of-sample result, never on a hunch. P(total P&L < 0) under realistic costs was ≈ 37%.
How a hypothesis earns a headstone
Publishing failures is only worth anything if the tests were honest before the result was known. Three disciplines make that true here.
Pre-registered gates
Every threshold on this page was written into the test script before we looked at a single line of output — and the scripts carry a comment saying so. A hypothesis clears the bar it set for itself, or it does not.
Non-overlapping windows
Rolling k-day forward windows on autocorrelated data inflate t-statistics enormously. We report the honest non-overlapping t. In the GEX study that alone moved the t-stat from 5.45 to 0.36 — same data, opposite conclusion.
Post-publication haircut
Published anomaly returns decay ~58% after they are published (McLean & Pontiff). We haircut every published Sharpe 40–60% before believing it, and prefer structural, risk-premium stories over mechanical ones that decay fastest.
What survived, and where to look
The graveyard is the point: the strategy we publish is the small residue that cleared its own gates. See exactly how it works, and watch it run week by week on a paper account.
Read this honestly: every figure on this page is a backtest or a paper measurement. Backtested and paper results are hypothetical, benefit from hindsight, and are not a promise of future performance. There is no real-money track record. Uptogain publishes one public model for everyone — it is not personalized advice, and we neither manage your money nor take custody of it.