Signals-Before-Storms research log, in the order it happened India 2016-07-22 to 2023-12-29, n = 1,814

The model worked.
The strategy did not. A Hidden Markov Model regime overlay for tactical asset allocation, run leak-proof over eight years of Indian and US markets.

No book here separates from a static 60/40 or equal weight once the comparison uses a paired difference test. What the overlay does buy is drawdown, cutting the worst loss by roughly two thirds on Indian data. Published as a negative result with a diagnosis, and it retracts one of its own findings.

Four panels. Top left, next-day annualised return rises with the volatility label while volatility also rises, so the return ordering runs the wrong way. Top right, crisis episodes as individual bars showing only fourteen episodes behind two hundred and sixty-one days. Bottom left, paired Sharpe differences against each benchmark, every interval crossing zero. Bottom right, equity curves and drawdowns, where the overlay's worst loss is a fraction of the benchmark's.
Next-day annualised return rises with the volatility label while volatility also rises, so the return ordering runs the wrong way. Crisis episodes as individual bars, showing only fourteen episodes behind two hundred and sixty-one days, one of which dominates. Paired Sharpe differences against the 60/40 benchmark, every interval crossing zero.
The whole argument in four panels. The states are real and persistent, but they predict variance rather than direction (top left). The evidence behind any regime claim is far thinner than the day count suggests (top right). No book separates from its benchmark under a paired test (bottom left). What the overlay does buy is drawdown (bottom right).

What was built

Three hidden states inferred from causal features, each state routed to its own convex program. The interesting engineering is not the model, it is the two places the pipeline refuses to look at the future.

The pipeline, from prices to deflated statistics Nine stages run top to bottom. Prices from Yahoo Finance feed build_master, which takes log returns and applies a vendor-error guard. build_features derives momentum, realised volatility and VIX features, all causal. An expanding walk-forward fits the scaler on training rows only and refits the HMM every fold. Test labels come from a causal forward filter, never whole-sequence Viterbi. Labels drive a per-regime convex program, and run_book applies a one-day execution lag and 7.5 basis points of cost to every book alike. The final stage computes deflated Sharpe, paired bootstrap intervals and episode counts. Two stages are marked as the leak-proofing guards: the causal forward filter and the execution lag. Yahoo Finance NIFTY 50 / gold / cash / India VIX build_master log returns, vendor-error guard build_features momentum, realised vol, VIX. all right-aligned Expanding walk-forward scaler fit on train rows only, HMM refit per fold Causal forward filter guard: never whole-sequence Viterbi Regime label 0 Bull / 1 Bear / 2 Crisis, ascending risk Per-regime convex program cvxpy, long-only, weight-capped run_book guard: 1-day execution lag, 7.5 bps, every book Deflated Sharpe paired bootstrap CI, episode counts

The stance map

Each regime gets a distinct convex program. Long-only, fully invested, weight-capped.

RegimeObjectiveReasoning
0 BullMaximise SharpeCalm market, take risk
1 BearMinimise varianceStressed, preserve capital
2 CrisisMinimise variance, hard equity capViolent, de-risk hard

The states are not noise. The transition matrix diagonal runs 0.97 to 0.98, so the regimes persist rather than flickering. The corner entries are zero: Bull never jumps straight to Crisis and Crisis never jumps straight to Bull, so the market always passes through the middle state. Nothing in the fit asked for that.

The first result, and the whole problem

Measured at the lag the strategy actually trades: the label is known at the close of day t, the return is earned on day t plus one.

Next-day annualised return and volatility by regime label, both universes.

Regime label India daysIndia returnIndia vol US daysUS returnUS vol
0 Bull821+10.2%11.2%665+10.9%8.7%
1 Bear731+15.0%14.9%841+14.7%15.6%
2 Crisis261+18.4%31.7%379+16.1%32.2%

Volatility orders perfectly with the label. Return orders backwards, on both universes. That single table explains the entire result. De-risking on the Crisis label means selling the highest-returning days, because the rebounds of April 2020 and late 2022 are as violent as the crashes that preceded them.

A state variable has to predict direction before a directional bet on it can pay. Realised volatility is symmetric in sign by construction, so it cannot tell a crash from a rebound. The model was asked for regimes and it delivered regimes. They are the wrong kind.

Bars of next-day annualised return and annualised volatility for each regime label on India. Volatility climbs steeply from 11.2 percent at label 0 to 31.7 percent at label 2. Return also climbs, from 10.2 percent to 18.4 percent, against the expected slope drawn beside it.
The measured slope beside the expected one. The volatility bars climb the way a risk-ordered label should. The return bars climb too, which is exactly backwards from what the stance map is betting on.

Four rescues, each judged against a criterion written down first

The obvious response to a broken result is to fix it. Every attempt was given a success criterion before it was run, so that the outcome could not be reinterpreted afterwards.

Re-rank the states by return
Criterion: a return-ranked ordering should separate direction where a volatility-ranked one does not.
criterion not metUS Sharpe moved 0.542 to 0.620, still under 60/40, and the crisis label came out identical: the same 379 days at the same +16.1%. Reordering cannot add information the state space does not contain.
A structurally different estimator
Criterion: a Statistical Jump Model should find a different partition, and a directional one.
criterion not metIt does find a different partition, agreeing with the HMM only 57.5% of the time on India, and it does what it advertises: US dwell time 27 days to 194, turnover 4.19x to 1.46x. Its US crisis label still carries the highest forward return, +20.0%. Two estimators, same broken ordering.
Volatility targeting at 10%
Criterion: if the book is taking too much risk in the wrong states, a constant risk budget should help.
criterion not metUS 0.542 to 0.562, and bit-identical on India at 0.824. It barely binds, because minimum variance already pins the book at 9.6% volatility on the US and 3.9% on India. The strategy is not taking too much risk. It is taking far too little.
Add a drawdown feature
Criterion, written down first: the crisis label's forward return must turn negative.
criterion not metIt went the other way on both universes: US +16.1% to +17.6%, India +18.4% to +29.7%. It also happens to top the India Sharpe table at 0.877, and it is still not counted as a win and not adopted as the default. Promoting it on a metric other than its stated one is precisely the selection bias the pre-registration exists to prevent.
Mean portfolio weight per asset in each regime. Equity never exceeds roughly a quarter of the book in any regime, with the remainder in cash and gold, and the crisis regime holds the least equity of all.
Why volatility targeting had nothing to bind on. The stance map does exactly what it was told, and this is the cost: equity never exceeds a quarter of the book in any regime. The overlay is not too aggressive, it is permanently defensive.

The retraction

One result did look like a genuine directional state. It was written up. Then it was counted properly, and withdrawn the same day.

The Jump Model's India crisis label reads -17.1% annualised over 94 days, which would have made it the only negative-return state anywhere in this project. It is two episodes: 64 days across COVID at -10.70%, and 30 days in late 2018 at +4.41%. Excluding COVID the label runs +30.17% annualised, the same backwards ordering as everything else.

Worse, a cross-tabulation put all 94 of those days inside the HMM's own crisis label, so it was never a different state space to begin with. The effective sample was one event. The finding is retracted here rather than quietly dropped.

Days are not a sample size

That retraction came out of a general check: a claim about a regime is supported by how many times the regime occurred, not how many rows it spanned. Drop each label's single longest episode and see what survives.

India, out-of-sample. Effective sample size and what remains without the largest episode.

Regime labelEpisodesAnnualised returnExcluding largest episode
0 Bull27+10.2%+2.7%
1 Bear30+15.0%+13.1%
2 Crisis14+18.4%+53.6%

The crisis label spans 261 days, which sounds like evidence, across just 14 episodes, only 3 of which lost money. Drop the longest one and the label runs +53.6% annualised. The backwards return ordering is not an artifact of one event: it survives episode counting on both universes.

Each crisis episode drawn as its own bar, ordered by length. One long COVID episode dominates the chart and the remaining thirteen episodes are much shorter, several of them positive.
Fourteen bars, not 261 days. One episode dominates the record. Any claim resting on this label rests on a sample of fourteen, and mostly on one of them.

The comparison was wrong too

An earlier draft compared strategies by checking whether their separate confidence intervals overlapped. That is not a comparison, and correcting it changed what could honestly be claimed.

Each marginal interval asks whether one book's Sharpe differs from zero. The question that matters is whether the gap between two books differs from zero, which needs the interval of the difference, resampled on the same dates for both books. Since every book holds the same assets on the same days, differencing cancels the shared market move and the paired interval comes out roughly three times tighter.

India. Sharpe difference against each benchmark, with its paired bootstrap interval.

Bookvs 60/40vs equal weight
HMM plus drawdown feature+0.273 (-0.36, +0.86)+0.063 (-0.19, +0.28)
Volatility-threshold rule+0.244 (-0.39, +0.79)+0.033 (-0.23, +0.28)
HMM, regime-conditional+0.220 (-0.47, +0.83)+0.009 (-0.28, +0.26)

Every interval spans zero, on both universes. The "it is all noise" reading survives, and it is now earned rather than assumed. It also sharpens one claim that used to be overstated: on the US the volatility rule's marginal interval excludes zero, which says only that its Sharpe beats zero. Paired against 60/40 it is +0.268 (-0.029, +0.559), so even the ablation that wins the US table cannot be said to beat the benchmarks.

What held up

A negative result is only worth reading if the machinery that produced it is trustworthy. These are the checks that survived their own scrutiny.

heldLeak-proofing, asserted by tests

Lookahead bias is the easiest way to produce a beautiful, worthless backtest. Five defences, each pinned by a unit test that fails if the defence is removed: causal right-aligned features, standardisation fit on training rows alone, a per-fold model refit, causal forward-filter decoding instead of whole-sequence Viterbi, and a one-day execution lag enforced in a single shared code path so every strategy and benchmark inherits it.

heldDeflation at an honest trial count

A Sharpe reported once, from one strategy out of several tried, is a biased number. The deflated Sharpe here charges for 7 trials, stated openly. The volatility rule read 0.928 at 4 trials, and searching three more ideas lowered every number in that column. That is what deflation is for. Intervals use a Politis-Romano stationary bootstrap with the block length read off the data rather than guessed.

heldSensitivity reported as a surface, adopting nothing

Sharpe ratio plotted against each swept parameter in turn, with the shipped default marked on every panel. The cost panel is nearly flat across the whole range.
Costs are not the story. Sharpe moves only 0.841 to 0.785 across free trading to 25 basis points. At zero cost the strategy still does not beat its benchmarks, which is a cleaner statement of the result than any net number.

The sweep reports the whole surface and adopts nothing, because picking the best cell would be a search and would owe the deflated Sharpe another trial. Two knobs, weight_cap and rebalance_confirm_days, are not flat, and the shipped defaults sit at local Sharpe optima. Both were fixed before any result was computed and neither is re-chosen, but a reader is entitled to be told.

heldVendor data guarded, not trusted

GOLDBEES.NS on Yahoo prints a 100x round trip over 2019-12-19 to 2019-12-23: log returns of -4.61 then +4.61. Two bad prints in 2,193 rows. Left alone they inflate gold's return standard deviation from 0.011 to 0.139 and poison every Indian covariance, regime fit and Sharpe. The guard rejects any daily absolute log return above 0.5 with a loud warning, and a test pins that a genuine -13% crash day still survives it.

heldA no-model ablation that can beat the model

The strongest control here is a two-line volatility-threshold rule: same optimiser, same costs, same walk-forward, but the regimes come from a trailing-volatility quantile instead of the HMM. If the HMM cannot beat that, the HMM is decoration. On the US the rule wins outright, 0.958 against 0.542.

The NIFTY equity curve with out-of-sample regime labels shaded behind it in a light to dark ramp. Darker crisis shading lines up with February 2018, the fourth quarter of 2018, the COVID crash and 2022.
Out-of-sample labels, no future information. The causal filter independently flags February 2018, Q4 2018, COVID and 2022. The states are real. That was never the part that failed.

The scorecards

Every book, strategy and benchmark alike, runs through the same cost engine at 7.5 basis points per unit of turnover with a one-day execution lag. A benchmark costed on different terms is not a benchmark.

India, the primary universe. Out-of-sample 2016-07-22 to 2023-12-29, n = 1,814. Sharpe is measured against the 3.79% risk-free rate the cash sleeve actually paid, because scoring against zero would hand every defensive book a free Sharpe for holding cash. Benchmark rows are set in italic.

BookSharpe (net)Sharpe (gross)SortinoMax drawdownCalmarTurnover/yr
HMM plus drawdown feature0.8770.8961.291-6.4%1.1690.99x
Volatility-threshold rule (ablation)0.8480.8631.244-6.5%1.1650.88x
Jump Model regimes0.8290.8331.204-7.3%1.0380.24x
HMM, regime-conditional (shipped default)0.8240.8411.211-6.2%1.1610.87x
HMM plus volatility targeting0.8240.8411.211-6.2%1.1610.87x
Equal weight0.8150.8191.167-15.2%0.6250.41x
HMM, unconditional0.7480.7581.097-6.2%1.1030.50x
Static 60/40 (equity/cash)0.6040.6070.827-23.7%0.4130.36x

Read the Sharpe column and the strategy is unremarkable. The overlay's point estimate sits above both benchmarks, and entry 05 is the reason that means nothing: no book here is statistically distinguishable from either one. Read max drawdown and Calmar, the other two metrics the brief names, and the picture inverts: -6.2% worst drawdown against -15.2% and -23.7%, with Calmar roughly 1.9x and 2.8x the benchmarks. That gap is not noise, it is the mechanical consequence of routing to minimum variance whenever the label is not calm. The overlay buys drawdown protection and pays for it in return. Whether that trade is worth making is a mandate question, not a statistical one.

Equity curves for every book on top and their drawdowns underneath. The 60 40 book climbs highest and falls furthest, dropping about 24 percent. The HMM books track lower and flatter with a worst drawdown near 6 percent.
The trade, drawn. The benchmark ends higher. The overlay never falls far. Those are the same fact seen twice.

US, the robustness universe. Out-of-sample 2016-07-05 to 2023-12-29, n = 1,886, risk-free rate zero, with a real bond sleeve. Benchmark rows are set in italic.

BookSharpe (net)Max drawdownCalmarTurnover/yr
Volatility-threshold rule (ablation)0.958-21.9%0.4463.61x
Static 60/400.690-27.6%0.2730.41x
Jump Model regimes0.682-25.1%0.2621.46x
Equal weight0.650-23.0%0.2580.42x
HMM plus drawdown feature0.590-28.0%0.1913.88x
HMM plus volatility targeting0.562-27.3%0.1824.10x
HMM, regime-conditional0.542-28.7%0.1704.19x
HMM, unconditional0.535-27.4%0.1742.74x

The US does not reproduce India, and saying so is the point of running it. No HMM book improves on either benchmark on Sharpe or Calmar, and the drawdown protection that vindicated the overlay on India does not reappear: the best HMM drawdown here is level with 60/40 and well behind equal weight at -23.0%. Minimum variance had nowhere to hide in 2022, when bonds fell alongside equities. A result that held on one market and was quietly assumed to hold on the other would be the more comfortable story. It is not the one the data tells.

One more thing the sweep decided rather than assumed. Three regimes were chosen to match the economic states being allocated against, then checked with a BIC sweep. It supports three on the primary Indian universe, where the step from three states to four buys only 392 of fit against 6,256 for the step from two to three. It does not support three on the US. Both are reported.

The verdict

Hidden Markov states on returns and volatility are volatility states, and volatility carries no directional information. That is the finding, and it outlives this codebase.

Realised volatility and VIX are symmetric in sign by construction. Any allocator built on them alone will de-risk into rebounds as reliably as into crashes, because it cannot tell the two apart. Finding a directional state needs a signed state variable: credit spreads, market breadth, earnings revisions, positioning, realised skew. Not another estimator on the same features.

The one thing the overlay does buy is drawdown, and only on one of the two universes tested. Two of the three metrics the brief names favour it there. That is reported as exactly what it is, on one universe, rather than promoted into a headline.