FAQ for serious quantitative investors
AZTMM Methodology · Last updated 2026-05-08
An external review of AZTMM raised nine questions that any sophisticated quant investor would ask. We answer each honestly below — what we have, what we don’t, and our roadmap. If you’re evaluating AZTMM as a research source and these answers don’t satisfy you, that’s a useful signal that AZTMM isn’t yet at the maturity level you need. We’d rather earn your trust slowly than oversell.
Top-line framing before the questions: AZTMM is currently an analytical framework, not validated alpha infrastructure. The Market Pulse Index (MPI) and the regime classifications are state-descriptions of public market data. They are not strategies. They do not have a track record in the strict, audited, walk-forward sense that an allocator would require. Where we have partial answers we share them; where we don’t, we say so and give a target date.
Q1. What survives after transaction costs and slippage?
The honest answer is that this question doesn’t map cleanly onto what AZTMM publishes today, because we don’t publish strategy-level returns. The MPI score and the regime label are descriptive measurements of current market state — they are not entries, exits, or position sizes. There is therefore no return stream from which to subtract frictions.
That’s the technically correct answer, but it’s also a dodge if we leave it there, because it points to a deeper gap. The question really being asked is: when AZTMM eventually publishes a strategy backtest, will it model real-world frictions or will it claim frictionless alpha? Here’s the policy.
What we will model
When we publish a strategy that uses MPI or regime as an input — e.g., “long SPY when MPI > 65 and regime is Bull” — the backtest will assume retail-tier execution: roughly 50 bps round-trip on liquid options (commission, exchange fees, plus modeled bid-ask), and a more conservative 5–8 bps round-trip on liquid ETFs like SPY/QQQ. Slippage will be modeled as a function of bar volume, not assumed away. Re-balance dates that fall on illiquid sessions (FOMC, half-days) will carry a penalty.
What we will not do
We will not publish a backtest that claims frictionless fills, perfect-VWAP execution, or zero spread. Most retail backtests fall apart the moment realistic 30 bps + bid-ask + fill-quality assumptions are bolted on. We’d rather not publish a strategy than publish one under under-realistic frictions. A backtest that survives only because the friction model is generous is a marketing artifact, not research.
Roadmap
The first AZTMM strategy backtest with realistic friction modeling is targeted for Q4 2026, alongside the live-Sharpe work in Q2 below. Until then, “what survives after costs” is genuinely TBD and we won’t pretend otherwise.
Q2. What is the live Sharpe ratio?
We don’t have one. There is no live AZTMM strategy, so there is no live return series, so there is no Sharpe ratio — live, ex-post, walk-forward, or otherwise.
This is the cleanest answer to give a sophisticated investor, because the alternative answers are all worse. We could compute a Sharpe of “long SPY when MPI > X” on the in-sample data we’ve already seen, but that number would be cherry-picked by construction — the threshold and the lookback would have been optimized in hindsight. We could also compute a Sharpe of MPI itself as if it were a return stream, but the MPI is a 0–100 score, not a P&L; that calculation would be nonsense.
What we have
A Performance Archive of dated MPI and regime calls with documented hits and misses. It is a track record in the “here’s what we said and here’s what happened” sense, not in the “here’s the audited monthly P&L of a strategy” sense.
Roadmap to a real Sharpe
When we publish a clearly-defined strategy — e.g., “rotate between SPY and short-duration Treasuries based on MPI threshold + regime” — the protocol will be: pick the rule from in-sample data ending 2024-12-31, freeze it, and report walk-forward out-of-sample results from 2025-01-01 forward with realistic frictions and confidence intervals on the Sharpe. Target publication: Q4 2026. By then we’ll have ~12 months of true OOS data on a frozen rule, which is the smallest sample size at which the Sharpe number stops being a coin-flip artifact.
Q3. How stable are hidden states over time?
Detailed methodology lives at Regime Classifier Methodology; the recap and stability discussion are below.
The regime classifier fits three states (Bull / Neutral / Crisis) on a long historical window of broad-market pricing data. It is refit every Sunday and the state labels are reassigned post-fit by sorting on mean return so “Bull” is always the highest-mean state and “Crisis” the lowest-mean state. The label sort is what gives the states semantic continuity across refits — without it, the same statistical state could be called “Bull” one week and “Neutral” the next purely because the fitting algorithm permuted the indices.
What we observe about stability
State means and variances drift slowly week-over-week as new market data shifts the underlying distribution — this is by design. The drift is typically small enough that the regime label assigned to last Sunday rarely flips on this Sunday’s refit. Material drifts do happen at structural breaks: post-2020 the Bear-state variance widened materially as the COVID and 2022 vol regimes entered the training window. We flag this kind of drift in the Regime Classifier Changelog.
What we don’t publish yet
We don’t publish a quantitative state-stability metric — e.g., Frobenius distance between consecutive regime persistence frameworks, or KL divergence on emission distributions. That’s on the roadmap (Q4 2026). For now stability is reported qualitatively in the changelog.
Caveat that matters. The emission assumption breaks during fat-tail events. A −7σ week sits so far in the left tail of the fitted Bull-state distribution that the model essentially refuses to assign it to Bull and snaps the model confidence to Crisis. That’s usually the right call, but it means the “states” are partly artifacts of the probabilistic distribution assumption, not pure features of the data. A heavy-tailed emission would handle fat tails better; that is a candidate change tracked in the changelog.
Q4. What percentage of regime transitions are false positives?
We don’t publish a precise false-positive rate yet, and we want to be transparent about why: there is no settled, model-independent ground truth for “regime transition” in equity markets. Candidate ground-truth sources — NBER recession dates, NBER cycle peaks, formal drawdown thresholds (e.g., −10% from 252-day high), credit-spread spikes — all disagree with each other, and all are themselves models. Picking one and calling it ground truth would let us publish a number, but the number would mostly be a measurement of how well the regime classifier matches our chosen reference model, not how often it’s “wrong.”
What we know qualitatively
The regime classifier lags at fast regime changes — typically 1–2 weeks behind a clean break. Two documented examples in the Performance Archive:
- Aug 2024 yen-carry unwind — the regime classifier flipped to Crisis roughly one week after the initial 5-Aug shock. The MPI’s VIX percentile and credit-spread sub-indicators moved same-day; the regime classifier caught up after the next weekly refit.
- Jan 2025 whipsaw — three regime flips in 14 days. This is the failure mode of a sticky model meeting an unstable tape. Most of those flips were almost certainly false signals in the “the underlying regime didn’t actually change three times” sense.
Slow regime drifts (e.g., Bull-to-Sideways transitions over 4–8 weeks) are handled more gracefully because the long historical window has time to absorb the new distribution before the regime classification flips.
Roadmap
Q4 2026 deliverable: a published FP-rate metric using two explicit ground-truth definitions — (a) a 252-day drawdown rule (≥10% from prior peak = Crisis, ≥0% from prior peak = Bull, in-between = Neutral), and (b) NBER recession dates lagged by the announcement delay. We’ll report FP and FN rates against both, and disclose how sensitive they are to the threshold choice.
Q5. How does the system behave during structural breaks?
Honestly: the regime classifier degrades during structural breaks, because the underlying return-generating distribution is changing faster than a rolling long historical window can absorb. That’s a known weakness of the architecture, not a bug we’ve patched.
Documented behavior
- Mar 2020 COVID. The classifier caught the Crisis regime within ~3 weeks — slower than fundamentals, faster than NBER (which dated the recession from Feb 2020 only retroactively in Jun 2020).
- Aug 2024 yen-carry. ~1-week lag from initial shock to regime classifier Crisis label.
- Jan 2025 whipsaw. Three regime flips in 14 days — the model thrashed between Bull and Neutral as the distribution oscillated. Most of those transitions were noise rather than signal.
Mitigation we use today
The MPI’s sub-indicators — especially the VIX percentile and the HYG−LQD credit spread — respond faster than the regime classifier does, because they don’t require a refit. We pair the current regime with these sub-indicators to give forward-looking texture during stress. When VIX percentile spikes above the 95th percentile and credit spreads widen above their historical percentile threshold while the regime classifier is still labelling Bull, we treat that gap as a signal that the regime is in the process of breaking and the regime classifier hasn’t caught up. The Performance Archive documents how we communicate this gap in the Daily and Weekly Pulse during fast moves.
What we won’t pretend
We won’t claim the regime classifier gives an early warning at structural breaks. By construction it lags. Anyone selling you a regime model that catches the top in real time is selling you a story.
Q6. What is the retraining cadence?
Cadence is fixed and published. There is no discretionary refit-when-it-feels-wrong loop, because that loop is how p-hacking gets dressed up as “adapting to new conditions.”
Regime classifier
- Refit every Sunday on long historical window of broad-market pricing data.
- Model confidences updated daily using the past week’s data — the model parameters are frozen between refits, only the inference is re-run.
- State labels reassigned post-fit by mean-return sort to preserve semantic continuity.
MPI
- Rebuilt twice daily — 9:15 AM ET (pre-market) and 4:15 PM ET (post-close). Mon–Fri only.
- Rolling historical baselines for each sub-indicator update with each session, so percentiles always reflect the most recent quarter.
- Sub-indicator weights are static; they have not been changed since the v6 methodology was frozen. Any future weight change will be dated and logged in the MPI Changelog.
Source-data refresh
- FRED — daily for credit spreads, DXY, TIPS breakeven.
- AAII bull/bear survey — weekly, released Wednesday afternoon ET.
- CBOE — end-of-day for VIX term structure, SKEW, put/call.
- Yahoo / Stooq — intraday delayed for SPY, sector ETFs, FX crosses.
Full cadence and source documentation lives at /methodology/mpi-methodology/ and Regime Classifier Methodology.
Q7. How is feature leakage prevented?
Feature leakage — using information that wouldn’t actually have been available at decision time — is the single most common way honest-looking backtests turn out to be fiction. Our defenses are structural rather than after-the-fact audited.
Regime classifier
Fit on weekly closing data only. There is no intraday data in the training set, so there is no path for intraweek look-ahead to enter the model. The Sunday refit uses data through the previous Friday close; Monday’s model confidence uses the model fit on data up to last Friday plus Monday’s observation.
MPI sub-indicators
Every sub-indicator uses end-of-session or prior-session data — never forward-looking. The multi-factor composite is computed only after all sub-indicator inputs have finalized for the session. The 4:15 PM ET rebuild uses 4:00 PM ET closes; the 9:15 AM ET rebuild uses the prior session’s closes plus any overnight FRED updates that posted before 9:00 AM ET.
Specific guard: AAII weekly survey
The AAII bull-bear spread releases Wednesday afternoon ET. It is incorporated into MPI calculations Wednesday afternoon and forward only. Monday and Tuesday MPI values do not get the new survey retroactively patched in — this is a common leakage error in retail dashboards, and we explicitly guard against it in the rebuild pipeline.
What we don’t publish yet
We don’t publish a feature drift report, point-in-time data audit, or vintage-aligned re-computation log. A reviewer who wanted to verify the no-leakage claim would have to take our word for it today. That’s on the roadmap (a vintage audit of MPI for selected dates is targeted for Q1 2027), and we acknowledge it’s a real gap until then.
Q8. Does complexity materially outperform simpler trend-following systems?
We don’t know yet, because we haven’t published a head-to-head benchmark study. The reviewer who raised this question is right to flag it: in finance, complexity often reduces robustness. Every parameter you add is another knob the optimizer can over-fit. The base rate of complex models out-of-sample beating a 50-day / 200-day SMA crossover on SPY is not 100%, and the burden of proof is on the complex model.
Our actual framing
AZTMM publishes a research composite — the MPI — and a regime label — the regime classifier. Neither is a strategy. The right benchmark question for a composite is therefore “does this composite track market state in a way that’s informationally additive over a simple price-trend signal?” That is a different question from “does this strategy beat trend-following.” Conflating the two would let us claim a benchmark win that we haven’t earned.
Roadmap: the benchmark study we owe
Q4 2026 we’ll publish a paired study covering 2024-01 through 2026-09:
- Benchmark A: SPY 50d/200d SMA crossover (long when 50 > 200, flat otherwise).
- Benchmark B: SPY 200d SMA filter (long when above, flat below).
- Illustrative benchmark specification (for the planned study only — not a strategy we publish or endorse): long SPY when MPI > 55 and regime classifier ≠ Crisis, flat otherwise.
Frictions modeled per Q1 above. We’ll report Sharpe, Sortino, max drawdown, and turnover for all three. If the AZTMM rule doesn’t materially beat both benchmarks net of friction, we will say so on this page. If it ties — which is the most likely outcome given how good 200d-SMA already is — we will also say so. A tie is not a failure for a state-description tool, but it would mean the complexity isn’t paying for itself as a directional signal, and you’d be right to weight a simpler system accordingly.
Until that study ships, the honest answer is: TBD, and the prior should lean toward “no, complexity does not outperform.” The historical base rate says so.
Q9. How does the platform validate causal relationships versus correlations?
Most of what AZTMM publishes is correlative, not causal, and we try to be careful about not letting the language drift in the other direction.
Where we explicitly avoid causal claims
- The MPI score doesn’t predict — it describes. A high MPI is consistent with a risk-on tape; it does not assert that risk-on conditions cause future returns.
- The regime classifier identifies regimes — it doesn’t explain why they happen. The model has no concept of monetary policy, earnings, or geopolitics; it only sees returns and vol.
- Sub-indicators are observed proxies, not first-principle drivers. Credit spreads correlate with regime state. We do not claim credit spreads cause SPY to fall.
Where we rely on weak-causal narratives
In the editorial commentary — the Daily and Weekly Pulse — we sometimes use causal-flavored language (“credit spreads are widening as growth fears resurface”). That is shorthand for a co-movement story we believe is plausible based on prior cycles, not a model claim. We try to flag it as commentary rather than methodology output, and we explicitly avoid putting causal language into the MPI score readouts or the regime labels.
Where causal models would help — and we don’t have them
- Factor decomposition. Decomposing MPI moves into orthogonal risk factors (rates, credit, vol, momentum) would let us say which factor was actually driving a given week’s reading.
- Lead-lag analysis. Cross-correlation of sub-indicators with future SPY returns at multiple lags, with appropriate multiple-testing correction.
- Granger-causality tests across sub-indicators. Useful to identify which inputs are actually informationally additive vs. which are redundant with VIX.
None of these exist in the AZTMM pipeline today. They are realistic but non-trivial work. Roadmap target: 2027, with the lead-lag work likely first because it’s the cheapest to do honestly.
Closing
If you found a question we should add, email nikhil.kothari17@gmail.com. If you found an answer that’s vague or wrong, please flag it — we’ll update the page with a visible changelog note rather than silently editing.
The point of this page is not to convince a serious quantitative investor that AZTMM is ready for them. It is to give them enough to decide, accurately, whether it is. If the answers above tell you it isn’t yet, you’ve gotten the value we intended to deliver.
