Quantitative Research

Gueta Research #002: Does the 13→0 result generalize? Cross-market validation in 5 markets

Same frozen methodology, 5 new markets (EUR/USD, Gold, S&P 500, Bitcoin, Ethereum), 65 pre-registered trials: 0 survivors in every market. The zero generalizes.

Quantitative ResearchActualizado: 1 de septiembre de 2026Tiempo de lectura: ~8 min

Mahdi GoodarziPor Mahdi Goodarzi · Fundador & Product Builder de Gueta Quant

The question

⚠️ Erratum (2026-08-16): addendum published — correction of the execution model (loader) and of the PBO convention (see ADDENDUM). Corrected counts: EURUSD 13→8→7→0→0 · Gold 13→9→8→0→0 · S&P 500 13→10→8→0→0 · BTC 13→5→4→0→0 · ETH 13→7→4→0→00 survivors in all 5 markets holds. Corrected PBOs: EURUSD 0.5639 · Gold 0.1737 · S&P 0.5695 · BTC 0.2139 · ETH 0.1557 (convention sensitivity in the addendum). Gold’s best DSR drops from 0.9444 to 0.8818 (SMA Cross) — still rejected by the 0.95 threshold.

In Gueta Research #001 we pre-registered and ran a validation funnel of 13 simple strategies on daily EURUSD. Result: 13 → 0. None passed the four gates (OOS, walk-forward, DSR, PBO).

That left an obvious question: is that zero an artifact of the market, or does it generalize?

The answer, published today with the same frozen methodology and no change to manufacture better results: the zero generalizes. 5 markets × 13 strategies → 0 survivors everywhere.

Protocol: identical to #001, frozen before the data

The #002 pre-registration was fixed before downloading any data. Contract:

  • The same 13 strategies, the same locked parameters, the same pipeline (bar-close signals, close fills, position from the next bar).
  • The same 4 gates with the same pre-committed thresholds: G1 OOS net return > 0 and Sharpe > 0 · G2 WFO cumulative > 0 · G3 DSR ≥ 0.95 (N=14) · G4 PBO < 0.50 (CSCV, S=16).
  • Continuity check (kill criterion): running the M1 market (EUR/USD) had to reproduce the #001 result exactly (13→8→7→0→0, PBO 0.5593). It reproduced identically — the engine is the same engine.

The only difference is the input: new markets, new data sources. Per-market costs were also fixed before execution (next section), picking the conservative value where there was ambiguity: higher costs = harder to survive, which is the direction that favours honesty, not advertising.

The 5 markets and their pre-registered costs

Market Ticker Bars (2020–2025) Cost per side
EUR/USD EURUSD=X 1,561 2 pips (0.0002) — #001 continuity
Gold GC=F 1,509 2 bps (0.0002) ≈ $0.50/oz @ $2,500
S&P 500 ^GSPC 1,507 2 bps (0.0002)
Bitcoin BTC-USD 2,191 10 bps (0.0010) — realistic taker fee
Ethereum ETH-USD 2,191 10 bps (0.0010)

The data (Yahoo Finance) are not redistributed (terms of service): the engine accepts any Date,Open,High,Low,Close CSV, and reproduction uses user-supplied data with checksum verification — the same standard as #001.

Results: 5 × 13 → 0

Market G1 G2 G3 (DSR ≥ 0.95) G4 (PBO < 0.50) PBO
EUR/USD 8 7 0 0 0.5639
Gold 9 8 0 0 0.1737
S&P 500 10 8 0 0 0.5695
Bitcoin 5 4 0 0 0.2139
Ethereum 7 4 0 0 0.1557
Total (65 trials) 39 31 0 0

No strategy passed the full funnel in any market. The strategies most consistently passing G1 (SMA Cross, Momentum, Bollinger Reversion, Ichimoku — 5 of 5 markets) all failed at G3/G4: the statistical deflation for multiple testing absorbs their apparent edge.

The gold case: the closest miss

2024–2025 was a strong bull market for gold, and the trend strategies reflected it out-of-sample:

Strategy OOS return OOS Sharpe WFO DSR
Bollinger Reversion +5.58% +1.29 +7.05% 0.7894
SMA Cross +111.69% +2.19 +152.1% 0.8818
EMA Cross +105.38% +2.10 +113.9% 0.8594
SuperTrend +97.01% +1.98 +131.7% 0.8245

The best reached DSR 0.8818 — 0.068 from the 0.95 threshold. And it was rejected. The threshold was not moved to let gold through: a pre-registered 0.95 is a pre-registered 0.95. This is the funnel working exactly as it should: raw returns are not evidence; deflated significance is. We present this case as the closest miss of a strict gate — not as an opportunity.

Low PBO ≠ survival

An important point for reading this result correctly: the PBO (the set’s overfitting probability) was low in gold (0.10), S&P 500 (0.20) and Bitcoin (0.11) — the selection-overfitting risk was small. Yet there were zero DSR survivors.

PBO and DSR answer different questions: PBO says “how likely is it that the best backtest is overfitted”; DSR says “how much of the apparent edge remains significant after deflating for 14 trials”. Both must pass; in 5 markets, neither produced a strategy that passed both. The lesson for anyone selling backtests: a market with little overfitting is not a market with a demonstrated edge.

Negative results generalize

  1. 5 × 13 → 0. Within this cohort, these datasets, these costs and this protocol, no strategy passed the final criterion in any market.
  2. Raw OOS strength did not survive deflation — gold proves it with +111.7% OOS rejected at DSR 0.88.
  3. Low PBO does not imply survival — three markets with PBO < 0.21 and zero survivors.
  4. The breadth is still modest (5 markets, 2 years OOS, daily bars): this strengthens Gueta’s falsification capability; it does not establish — nor claim to — discovery capability of profitable strategies.

At Gueta our identity is falsification-first: we prefer to publish five reproducible zeros over an imaginary alpha. Evidence, negative or positive, is the product.

Reproducibility

  • Public repository: github.com/guetaquant-byte/guetaquant-researchfunnel-002/: pre-registration, engine, per-market JSON/CSV results, checksums, certificate CERTIFICATE_002.md.
  • The continuity check (EURUSD identical to #001) is the engine-identity gate for the whole run.
  • User-supplied data (Yahoo ToS); procedure in REPRODUCIBILITY.md (< 15 minutes).
  • Evidence level: E3 (WFO/CPCV) + the E4 reproduction standard; E5 (execution parity) and E6 (live fidelity) are not claimed — this is still research on daily OHLC.

As in #001, every statistic ships with its D8 checklist (formula → primary paper → edge cases → sensitivity), so each published number can be independently traced to its source.

Frequently asked questions

Does this mean trading gold or the S&P 500 doesn’t work? No. It means these 13 simple rules, with these costs and this protocol, showed no deflated edge in 2024–2025. Gold was the closest miss (corrected DSR 0.8818 vs 0.95) and was rejected by pre-registered rule. The research doesn’t prove no edge exists: it proves that claiming one requires exactly this level of evidence.

Why repeat the same experiment in other markets? Because a single-market sample is no empirical basis for any general conclusion. Breadth across markets (and later timeframes, execution models and datasets) is the path to a credible falsification capability.

Are the strategies that passed G1 in several markets “close to being validated”? No. Passing G1 (positive OOS return) is the first of four filters. None passed G3 (DSR ≥ 0.95) in any market, and none would ever be “validated” in that language: the most that could be said of a future survivor is “survived Gueta Research Level E3”, pending execution parity (E5) and live fidelity (E6).

Could the gold negative result be a missed opportunity? It is the opposite: it is the funnel protecting the researcher from himself. The OOS returns of a historical gold bull trend look enormous — and still fail the deflated significance for multiple testing. If a signals seller shows you a similar gold backtest, this research explains exactly why it is not enough.

Is this an investment recommendation? No. Gueta Quant is an educational platform. This research is academic and demonstrative in simulation; it does not constitute financial advice, operational signals or promises of profitability (Decreto 2555 de 2010).


Reproducibility and sources: protocol pre-registered 2026-08-16 (funnel-002-pre-registration.md), results 2026-08-16 (funnel-002-results.md), certificate CERTIFICATE_002.md + JSON, per-market sha256 checksums. Thresholds per Bailey & López de Prado (2014) and Bailey et al. (2017). OHLC data from Yahoo Finance, not redistributed (terms of service); the engine accepts any CSV of the same schema.

Mahdi Goodarzi — Fundador y Product Builder de Gueta Quant

Sobre el Autor: Mahdi Goodarzi

Fundador de Gueta Quant y profesional de marketing y growth con experiencia en fintech y mercados financieros. MBA en Marketing Digital de la Universidad de Teherán y formación en ingeniería. Actualmente desarrolla herramientas educativas para trading cuantitativo: programación en Pine Script v6 (TradingView), gestión de riesgo basada en datos (MQL5) y automatización.