Skip to the investigation

CASE 10

Why a Strategy Works Only on One Start Date

Start on January 1 and the strategy compounds to a 22% CAGR. Start one month later and it falls to 4%. The code did not change; the path did.

One failure mode. One validation verdict.Focused analysis · educationally constructed educational figuresstart date biasbacktest path dependencerolling start datescompounding strategy
BACKTEST DIAGNOSTIC PANELCase INCEPT. Educational illustrative values, not observed market data.BACKTEST DIAGNOSTIC PANELCase INCEPT · educational illustrative valuesJanuary start+22%February start+4%12-start median+10%failure boundaryPoint estimateDependenceTail stressExecutionSelectionReproductionA composite score summarizes evidence; it does not prove robustness.

01

Validation verdict for start-date dependence

A strategy whose verdict depends on one privileged start date has not demonstrated stable performance across entry into the historical path. Start-date sweeps expose timing luck that a single full-period curve conceals.

All figures in this article are educationally constructed examples created to explain the failure mode. They are not real strategy results or recommended thresholds.
02

What the headline metric obscures about start-date dependence

The earliest available date feels objective, so the resulting curve is treated as the strategy’s natural history. But data availability, indicator warm-up, first position, early sequence and compounding can make that date a powerful hidden parameter.

The problem is not that every investor begins on a different date. It is that a robust process should not require one favorable early sequence to create enough capital, margin or exposure for later wins.

03

How start-date dependence enters the backtest

Early trades set the capital base

Percent-of-equity sizing lets early wins expand every later position and early losses permanently shrink them.

Open positions are cut differently

A shifted start can omit an initial losing leg or enter after a trend has already begun.

Rebalancing calendars change

Monthly or quarterly rules can select different bars when the anchor date moves.

Warm-up and indicator state differ

Insufficient prehistory or stateful calculations can generate different first signals across start dates.

04

Compact reconstruction of start-date dependence

CASE 10 · backtest start date dependenceFocused analysis · educationally constructed educational figures
Start offset CAGR Max drawdown Net profit Rank among 12 starts
January anchor 22% −11% +118% 1st
February anchor 4% −31% +19% 11th
March anchor 13% −20% +62% 6th
Median of 12 monthly starts 10% −22% +47%

The January curve is real for that path, but it is the best of twelve entry points. Publishing only it turns an accidental anchor into an unstated model choice. The median and dispersion are more honest summaries of start-date robustness.

05

The test that can overturn the start-date dependence verdict

Run staggered starts while preserving the same end date, then stagger both starts and ends. Separate warm-up effects from capital-path and rebalancing effects.

Test monthly, weekly or signal-cycle start offsets appropriate to the holding period.
Provide enough pre-start history for indicators so only trading begins later, not indicator calculation.
Normalize starting capital and position-sizing rules for each run.
Report median, worst, best and dispersion of return, drawdown and recovery across starts.
Identify which first trades explain the spread and whether those events are unique.
06

What trade-list analysis can and cannot identify about start-date dependence

Export-level red flags for start-date dependence

  • The published start is the best among many available anchors
  • Removing the first winning trade collapses later position size
  • Monthly rebalancing changes sharply with a one-day offset
  • Warm-up length is shorter than the longest indicator memory
  • CAGR dispersion across starts is larger than the median CAGR

What the export reveals about start-date dependence

  • Side-by-side exports for staggered starts, normalized to common capital and end date
  • Distribution of CAGR, net profit, drawdown and recovery across anchor dates
  • Contribution of the first trades and early regimes to later compounded exposure
  • Rolling-window and walk-forward degradation that confirms the anchor dependency

What start-date dependence still requires from settings, code, or market data

  • One trade export contains only one start. The Lab needs multiple exported runs or a custom analysis workflow to compare anchors.
  • Start-date robustness does not guarantee future performance; it only removes one hidden source of historical luck.
ACADEMIC VALIDATION DOSSIER

Turn inception-date dependence into a falsifiable backtest diagnosis.

Case file 10/20 · INCEPT · one failure mechanism, one falsifiable protocol

01

Research abstract: start-date dependence

Case file 10/20 · INCEPT · one failure mechanism, one falsifiable protocol

This article tests one central proposition: a single inception date can mistake compounding of a favorable early sequence for universal performance. The question is not merely whether the displayed net profit or win rate was arithmetically calculated. The deeper identification problem is whether we know what constitutes one observation, what information was available at the decision time, which assumptions are necessary for the profit to exist, and how much of the conclusion survives when those assumptions are perturbed. The research object is therefore not one performance table; it is the linked data-generation, fill-generation, estimation, selection, and capital-allocation process.

The primary estimand is strategy performance that persists across plausible real-world inception dates. The observation unit is defined as overlapping equity cohorts constructed from shifted inception dates. Without this definition, split fills, duplicated signals, common events, synthetic prices, or timestamp conversions can be double-counted as independent evidence. A larger row count does not necessarily contain more independent information. An academically defensible analysis fixes the relationship between the observation unit and the estimand before it reports sample size, standard error, or statistical confidence.

The principal sensitivity axes are start-date shifts, initial capital, warm-up length, and exclusion of the first k trades. The hidden state is initial regime, first few trades, path dependence of compounding, and warm-up handling. In particular, one large early win increases subsequent position sizes and self-amplifies the full-period gap. Means and medians alone are incapable of describing that mechanism, so the analysis combines central estimates with lower quantiles, expected shortfall, sign stability, boundary-hitting frequency, and contribution concentration. The objective is not to find one pessimistic number, but to map the full region in which the original conclusion changes sign or ceases to be economically usable.

The conclusion does not attempt to prove that a backtest is good. It separates the component that remains after attempted falsification from the component that disappears when assumptions are reconstructed. The governing decision principle is to report the median, lower quantile, and sign-stability rate across an inception ensemble rather than advertising one representative start date. This is not trading advice; it is a research procedure for measuring how much evidentiary weight a TradingView trade export can carry. Liquidity not present in the file, broker-specific rules, future regimes, outages, and gaps require separate evidence, and statistical survival never guarantees future profit.

The numerical values illustrate the method for inception-date dependence; they are not a real strategy, client record, or forecast.

02

Hypotheses and identification target for start-date dependence

strategy performance that persists across plausible real-world inception dates

Null hypothesis / H₀

H₀ for start-date dependence: The reported performance is not materially dependent on the suspected failure mechanism and survives reasonable perturbations.

Alternative hypothesis / H₁

H₁ for start-date dependence: The reported performance depends materially on the suspected failure mechanism and deteriorates after reconstruction, perturbation, or dependence-aware resampling.

Estimand

strategy performance that persists across plausible real-world inception dates

Observation unit

overlapping equity cohorts constructed from shifted inception dates

Latent mechanism

initial regime, first few trades, path dependence of compounding, and warm-up handling

Stress axes

start-date shifts, initial capital, warm-up length, and exclusion of the first k trades

03

Formal estimands for start-date dependence

Definitions precede inference.

Θ(s)=T({r_t:t≥s})Performance estimate generated by the same evaluation rule T for inception date s.
Stab̂=|S|⁻¹Σ_{s∈S}1{signΘ(s)=sign med_sΘ(s)}Empirical sign agreement with the median across inception dates; report it as indeterminate when the median is zero.
Q₀.₁(Θ), Q₀.₅(Θ)Lower-decile and median performance across the inception-date ensemble.
Start-date variants120
Positive share58%
Median return+9.2%
Best+47.1%
Worst−18.4%

The primary estimand is strategy performance that persists across plausible real-world inception dates. The observation unit is defined as overlapping equity cohorts constructed from shifted inception dates. Without this definition, split fills, duplicated signals, common events, synthetic prices, or timestamp conversions can be double-counted as independent evidence. A larger row count does not necessarily contain more independent information. An academically defensible analysis fixes the relationship between the observation unit and the estimand before it reports sample size, standard error, or statistical confidence.

The principal sensitivity axes are start-date shifts, initial capital, warm-up length, and exclusion of the first k trades. The hidden state is initial regime, first few trades, path dependence of compounding, and warm-up handling. In particular, one large early win increases subsequent position sizes and self-amplifies the full-period gap. Means and medians alone are incapable of describing that mechanism, so the analysis combines central estimates with lower quantiles, expected shortfall, sign stability, boundary-hitting frequency, and contribution concentration. The objective is not to find one pessimistic number, but to map the full region in which the original conclusion changes sign or ceases to be economically usable.

04

Illustrative recomputation design for start-date dependence

For the start-date dependence reconstruction, table values are illustrative calculations used to expose a verdict reversal; they are not a user’s observed TradingView result.

ID Recomputation layer Operation Comparison Diagnostic purpose
S0 Reported result Restate the Strategy Tester aggregate Base Apparent conclusion
S1 Unit reconstruction overlapping equity cohorts constructed from shifted inception dates Reassess count and dependence Information correction
S2 Independent recomputation Rebuild price, size, cost, and currency row by row Separate reconciliation error Measurement validity
S3 Local stress start-date shifts, initial capital, warm-up length, and exclusion of the first k trades Perturb one factor only Causal sensitivity
S4 Tail injection one large early win increases subsequent position sizes and self-amplifies the full-period gap Recompute lower quantiles and boundary hits Capital preservation
S5 Dependence-aware resampling Generate paths across several block lengths Intervals and sign stability Estimation uncertainty
S6 Selection adjustment Log search, OOS review, and exclusions Correct maximum-selection bias Generalization
S7 Full gate report the median, lower quantile, and sign-stability rate across an inception ensemble rather than advertising one representative start date Compare with predeclared thresholds Pass / hold / reject

The illustrative recomputation for start-date dependence changes one processing layer at a time, then combines only predeclared layers. S0 is never treated as ground truth; it is the statement to be audited. S1 and S2 ask whether the exported unit and arithmetic are coherent. S3 and S4 identify local sensitivity and tail failure. S5 changes the uncertainty model rather than the trade list. S6 adjusts for the search that preceded publication. S7 applies the same gate to every version. This order prevents an adverse result from being explained away by simultaneously changing several assumptions.

In the start-date dependence figures, color and position encode diagnostic sensitivity only; they do not represent statistical significance or future P&L.

05

Diagnostic figures specific to start-date dependence

Four separate visual tests; no decorative chart reuse.

Inception-date performance surfaceSynthetic experiment; axes and thresholds are diagnostic, not forecasts.Inception-date performance surfaceSynthetic experiment; axes and thresholds are diagnostic, not forecasts.shifted inception dates →Educational normalized display. Read direction, slope, and boundary location—not the absolute level.
Figure 1. Primary diagnostic for inception-date dependence. Values are methodological illustrations, not estimates of a real strategy or future return.
Calendar of terminal outcomes by inception dateFigure 2. Calendar of terminal outcomes by inception date. Shifting inception one day at a time reveals whether the verdict depends on one privileged initial regime. Values are illustrative recomputations, not observed performance or forecasts.Calendar of terminal outcomes by inception dateA topic-specific estimand decomposed into one diagnostic viewinception date (left to right)start weekstable gainloss
Figure 2. Calendar of terminal outcomes by inception date. Shifting inception one day at a time reveals whether the verdict depends on one privileged initial regime. Values are illustrative recomputations, not observed performance or forecasts.
Equity-curve fan across shifted inception datesFigure 3. Equity-curve fan across shifted inception dates. With identical rules, shifting inception changes the interaction of early losses and compounding, producing divergent paths. Values are illustrative recomputations, not observed performance or forecasts.Equity-curve fan across shifted inception datesA topic-specific stress test designed to overturn the headline verdicttime since inceptionnormalized wealthonly inception date changes
Figure 3. Equity-curve fan across shifted inception dates. With identical rules, shifting inception changes the interaction of early losses and compounding, producing divergent paths. Values are illustrative recomputations, not observed performance or forecasts.
Path ensemble generated by alternate inception datesFigure 4. Path ensemble generated by alternate inception dates. Inception changes are branches that propagate through sizing and stopping rules, not merely shifts on the time axis. Values are illustrative recomputations, not observed performance or forecasts.Path ensemble generated by alternate inception datesA causal or processing structure separating observations, assumptions, and decisionsstart 1start 2start 3start 4gainflatlossgainsame rules, different inception positions
Figure 4. Path ensemble generated by alternate inception dates. Inception changes are branches that propagate through sizing and stopping rules, not merely shifts on the time axis. Values are illustrative recomputations, not observed performance or forecasts.
The primary diagnostic decomposes inception-specific expectancy, lower-decile performance, sign stability, first-k-trade contribution, and recovery duration along a causal axis. Read slope, curvature, and the first decision-boundary crossing as “start-date shifts, initial capital, warm-up length, and exclusion of the first k trades” changes, not merely the height of the favorable point.
The two-dimensional surface exposes interaction among “start-date shifts, initial capital, warm-up length, and exclusion of the first k trades.” Color is a normalized margin to a predeclared gate, not an empirical probability. A broad connected pass region is different evidence from a narrow isolated island.
The resampling statistic is distribution across inception dates. Compare an IID benchmark with circular inception shifts and time blocks that preserve market regimes across several block lengths, reporting the 2.5th, 50th, and 97.5th percentiles and verdict-reversal rate. Save seeds and repetitions.
The causal map traces “favorable start → compounding of early wins → smooth full-period curve → illusion of universality → weak result from another inception.” A displayed metric is an intermediate product, not the first cause; perturb the input or assumption, rebuild trades and capital boundaries, and return to the predeclared gate.
06

Multi-layer audit questions for start-date dependence

A result is only as strong as its weakest unresolved layer.

AUDIT LAYER 0101 · Fix the estimand

First, fix the estimand as “strategy performance that persists across plausible real-world inception dates.” Do not substitute net profit, win rate, or a visually smooth curve for that target. Declare the horizon, account currency, included frictions, and operating-stop boundary before calculation. Any post-result change creates a new hypothesis and version, preventing the question from being selected after the answer is known.

AUDIT LAYER 0202 · Reconstruct the observation unit

Reconstruct the observation unit as “overlapping equity cohorts constructed from shifted inception dates” before treating rows as independent evidence. Report raw rows, parent trades, decisions, event clusters, and the denominator used for each average or standard error. Recompute inception-specific expectancy, lower-decile performance, sign stability, first-k-trade contribution, and recovery duration under more than one defensible aggregation rule so that a larger export is not mistaken for a larger information set.

AUDIT LAYER 0303 · Preserve provenance and settings

Preserve the hash of the TradingView export and the symbol, timeframe, session, timezone, order-processing settings, costs, account currency, and Pine version. For inception-date dependence, initial regime, first few trades, path dependence of compounding, and warm-up handling directly affects reproducibility. Keep immutable source, normalized, and analysis layers separate, with every join, deletion, imputation, and conversion recorded in a transformation ledger.

AUDIT LAYER 0404 · Separate identification from assumption

The export identifies only what can be rebuilt from recorded time, price, quantity, and P&L. cash-flow timing, minimum trade size, and historical data required for warm-up requires additional evidence. Mark each causal link as observed, bounded by assumption, or externally unverified. This prevents initial regime, first few trades, path dependence of compounding, and warm-up handling from being presented as a confirmed fact when the available data support only an interval or conditional conclusion.

AUDIT LAYER 0505 · Reconcile row-level arithmetic

Do not adopt the platform summary as ground truth. Independently regenerate the wealth path from every realistic inception date using the same rule, initial capital, and warm-up policy, then compare the overlapping cohorts. Reconcile total and row-level differences by sign, date, symbol, and order type. If discrepancies concentrate in the exact state associated with inception-date dependence, treat that concentration as a primary finding rather than dismissing it as rounding.

AUDIT LAYER 0606 · Quantify finite-sample uncertainty

Report inception-specific expectancy, lower-decile performance, sign stability, first-k-trade contribution, and recovery duration with intervals or resampling distributions, not point estimates alone. Match the uncertainty method to sample size, skewness, heavy tails, censoring, and selection history. If normal, quantile, and dependence-aware methods disagree on the sign, classify the edge as unidentified and show the minimum detectable effect and lower decision bound.

AUDIT LAYER 0707 · Preserve serial and cluster dependence

Do not narrow uncertainty with an IID shuffle alone. Resample circular inception shifts and time blocks that preserve market regimes using several fixed block lengths and stationary bootstrap. Preserve random seed, repetition count, wrap rule, and missing-data treatment. For each block specification, report the distribution of inception-specific expectancy, lower-decile performance, sign stability, first-k-trade contribution, and recovery duration, the rejection-side tail mass, and the rate at which the verdict changes sign.

AUDIT LAYER 0808 · Measure tails and operating boundaries

Interrogate the mechanism “initial regime, first few trades, path dependence of compounding, and warm-up handling” with lower quantiles, expected shortfall, influence, cluster length, and boundary-hitting measures. Historical maximum loss is not a loss cap. Define several absorbing or operating boundaries—capital, margin, mandate drawdown, and recovery time—and record which boundary fails first under each stress.

AUDIT LAYER 0909 · Model execution and market frictions

A flat commission deduction is not an execution model for inception-date dependence. Allocate spread, slippage, financing, borrow, roll, conversion, rounding, and rejected orders to the relevant unit. Recompute inception-specific expectancy, lower-decile performance, sign stability, first-k-trade contribution, and recovery duration under base, upper-quantile, and crisis states while preserving the possibility that costs and losses worsen together.

AUDIT LAYER 1010 · Count the complete search path

Count the complete population of periods, symbols, timeframes, parameters, exits, filters, and metrics that were tried. Do not detach the attractive result for inception-date dependence from rejected candidates, interim changes, or repeated validation reviews. Where appropriate, use PBO, SPA, and a Deflated Sharpe Ratio, and treat an unrecorded trial count as a material audit limitation.

AUDIT LAYER 1111 · Condition on market regimes

Test whether inception-date dependence is concentrated in one trend, volatility, liquidity, rate, or session state. Define regimes prospectively or on training data only. Report statewise inception-specific expectancy, lower-decile performance, sign stability, first-k-trade contribution, and recovery duration, occupancy, transition probabilities, and costs, then reweight the mixture to adverse but realistic future compositions.

AUDIT LAYER 1212 · Separate path, inception, and sizing

For the start-date dependence case, the same trade set can follow different capital paths under another inception date, order, initial balance, rounding rule, or stop condition. Separate fixed quantity, fixed R, and percentage sizing, then use circular shifts and block orderings to recompute drawdown, recovery, and boundary hits. Equal terminal P&L does not imply equal path risk.

AUDIT LAYER 1313 · Design counterfactual stress tests

Perturb “start-date shifts, initial capital, warm-up length, and exclusion of the first k trades” one axis at a time before creating a joint sensitivity surface. Add the negative control “remove the first k trades from every cohort to test whether inception dependence is explained by a few early observations.” Predefine the grid and crisis rule so that neither the most favorable nor the most damaging cell is selected after inspection. Save the slope, curvature, and exact point where the decision boundary is crossed.

AUDIT LAYER 1414 · Verify through an independent implementation

Have a second implementation regenerate the wealth path from every realistic inception date using the same rule, initial capital, and warm-up policy, then compare the overlapping cohorts, then compare critical row-level outputs. Regression fixtures should include empty files, duplicate timestamps, extreme costs, reverse ordering, missing values, and boundary cases. Agreement between implementations is insufficient if they share the same bad input, so separate data construction and review roles where feasible.

AUDIT LAYER 1515 · Use a predeclared decision gate

Predeclare the decision rule. This case passes only if “the lower inception-date quantile clears the threshold and profit is not concentrated in one favorable start.” Near a boundary, disclose interval width and economic materiality rather than a binary badge. If only one favorable block length, cost state, or implementation passes, classify the result as assumption-sensitive rather than robust.

AUDIT LAYER 1616 · Maintain a reproducibility ledger

The evidence ledger must store the input hash, code version, settings, exclusions, “start-date shifts, initial capital, warm-up length, and exclusion of the first k trades,” block lengths, random seed, repetition count, and every scenario output. Keep exploratory and confirmatory results in separate namespaces and retain failed trials. When new TradingView data arrive, create a new version and track lower cohort quantiles, first-ten-trade contribution, sign stability, and longest recovery rather than overwriting the old result.

AUDIT LAYER 1717 · Translate statistics into capital impact

Translate statistical changes into capital consequences. A shift in expectancy, lower quantile, recovery time, or boundary risk caused by inception-date dependence should be mapped to trade count, capital, margin, and continuation. A small per-trade difference can compound under high turnover, while a rare loss can be decisive near an absorbing boundary.

AUDIT LAYER 1818 · Separate roles and enforce stop conditions

Separate hypothesis design, implementation, independent recalculation, and approval where practical. Stop automatically on material reconciliation error, unresolved missing data, non-reproducibility, or a predeclared threshold breach. Audit the chain “favorable start → compounding of early wins → smooth full-period curve → illusion of universality → weak result from another inception,” and monitor lower cohort quantiles, first-ten-trade contribution, sign stability, and longest recovery prospectively without turning a historical pass into a promise of future profit.

07

Falsification protocol for start-date dependence

report the median, lower quantile, and sign-stability rate across an inception ensemble rather than advertising one representative start date

Freeze the TradingView source for the start-date dependence audit

Store the export without alteration and record its hash, export time, strategy, symbol, timeframe, and settings. Preserve every column relevant to inception-date dependence; deletions and imputations belong only in derived tables.

Reconstruct the observation unit for start-date dependence

Aggregate rows into “overlapping equity cohorts constructed from shifted inception dates,” and report raw rows, parent trades, events, and independent clusters. Recompute the critical result under another defensible aggregation.

Independently recompute the displayed start-date dependence result

Independently regenerate the wealth path from every realistic inception date using the same rule, initial capital, and warm-up policy, then compare the overlapping cohorts. Reconcile row-level and aggregate outputs with Strategy Tester and preserve where discrepancies concentrate.

Isolate the start-date dependence mechanism

Treat inception-date dependence as the principal mechanism and move “start-date shifts, initial capital, warm-up length, and exclusion of the first k trades” one axis at a time while holding other settings fixed.

Map the operating boundary for start-date dependence

Combine the primary and interacting axes on a predeclared grid and recompute inception-specific expectancy, lower-decile performance, sign stability, first-k-trade contribution, and recovery duration. Record the width and connectivity of the acceptable region and every boundary crossing.

Resample the dependence structure relevant to start-date dependence

Use circular inception shifts and time blocks that preserve market regimes with several fixed block lengths and stationary bootstrap. Save every random seed, repetition count, and block specification.

Inspect influence points and operating boundaries for start-date dependence

For the start-date dependence influence test, remove the largest contributor, top-k contributors, selected periods, and relevant regimes in sequence; then recompute lower-tail measures and the operating boundary.

Apply negative controls and conservative bounds to start-date dependence

Remove the first k trades from every cohort to test whether inception dependence is explained by a few early observations. Bound cash-flow timing, minimum trade size, and historical data required for warm-up as unobserved factors rather than elevating the optimistic value into the final answer.

Apply the predeclared gate to start-date dependence

Do not move the threshold after seeing results. Compare with “the lower inception-date quantile clears the threshold and profit is not concentrated in one favorable start,” and distinguish pass, hold, and reject. Any unresolved material mismatch causes a hold.

Save a reproducible evidence package for start-date dependence

Bundle the source, transformation ledger, formulas, figures, all scenarios, failure logs, and code version for rerun in another environment. Prospectively monitor lower cohort quantiles, first-ten-trade contribution, sign stability, and longest recovery.

08

Decision gate for start-date dependence

Reject the story before trusting the curve.

How to read the start-date dependence figures and equations

The figures for start-date dependence use illustrative recomputations constructed to expose this specific failure mode. Do not infer statistical significance from line position or color alone; first verify the estimand, units, denominator, censoring rule, and cost sign defined by the equations. A sensitivity surface is not a causal estimate. It shows how a conclusion changes only within the stated assumptions. Resampling should compare an IID shuffle with stationary and block bootstrap procedures across several block lengths so that loss clustering and regime persistence are not silently destroyed. Store the random seed, iteration count, block length, bandwidth, and missing-data treatment, and claim reproducibility only after an independent implementation reproduces the same aggregates.

This case passes only if “the lower inception-date quantile clears the threshold and profit is not concentrated in one favorable start” across reconstructed values, local perturbations, joint sensitivity, dependence-preserving resampling, and the negative control, with no material sign reversal or unresolved reconciliation error. A pass is limited evidence against the stated failure mode, not certification of future profit.

  • The estimand and observation unit were fixed before outcomes were reviewed
  • For start-date dependence, any material disagreement between reported and independently recomputed values must be resolved or explicitly explained.
  • The start-date dependence claim passes this gate only when its acceptable stress region is broad and connected rather than one isolated favorable island.
  • The sign of the start-date dependence estimate must remain stable across defensible block lengths, saved seeds, and reasonable interval methods.
  • For start-date dependence, economic margin remains after deleting the largest and top-five contributors and key regimes
  • For start-date dependence, conservative cost, fill, and capital-boundary scenarios remain inside the stopping mandate
6/6required gates · not a performance forecast
09

Limitations, external validity, and reproducibility of the start-date dependence audit

Every inference has a boundary.

The first limitation is that a trade export does not contain the complete market state. If order-book depth, queue position, network latency, rejected orders, broker liquidity, or realized financing history is absent, strategy performance that persists across plausible real-world inception dates remains model-mediated. Model outputs should be displayed as scenario ranges and must not be formatted as though they were directly observed facts.

A second limitation specific to the start-date dependence analysis is structural change. A long historical sample does not guarantee a common population when market rules, participants, volatility, rates, spreads, data construction, or Pine execution semantics change. Do not increase nominal sample size by indiscriminately pooling old periods. Estimate rolling and regime-conditioned behavior and test parameter stability around detected changes.

A third limitation specific to the start-date dependence analysis is reuse of the diagnostic battery. Applying these tests repeatedly to the same data and editing the strategy until it passes turns the diagnostic process itself into another optimizer. Every post-test edit starts a new model version and requires untouched or prospective evidence. A test chosen after reading the outcome belongs to exploration and cannot be counted as independent confirmation.

A fourth limitation for the start-date dependence analysis is the distinction between statistical survival and operational suitability. Behavioral tolerance, locked capital, tax, regulation, outages, account terms, order-size limits, market-order restrictions, and liquidity discontinuities cannot be resolved from a CSV alone. The lab is a diagnostic for discovering hidden failure risk earlier; it is not investment advice, a performance warranty, or a guarantee of bounded loss. User-specific constraints remain a separate decision layer.

LIMIT 01Identification boundary

The estimand “strategy performance that persists across plausible real-world inception dates” is identified only within the columns present in the TradingView export and the stated assumptions. If cash-flow timing, minimum trade size, and historical data required for warm-up cannot be observed, report bounds rather than a false point estimate.

LIMIT 02Structural change

Past estimates of inception-date dependence need not belong to the same population after changes in rules, participants, volatility, costs, or data specifications. Track lower cohort quantiles, first-ten-trade contribution, sign stability, and longest recovery in rolling and regime-specific windows.

LIMIT 03Reuse of the diagnostic

For start-date dependence, repeatedly applying the same diagnostic battery and editing until it passes turns verification into another optimizer. Every post-audit change therefore creates a new model version and requires untouched evidence.

LIMIT 04Operational suitability

Even if the lower inception-date quantile clears the threshold and profit is not concentrated in one favorable start, the analysis does not establish tax, regulatory, behavioral, liquidity, order-size, or systems suitability. Separate statistical diagnosis from live-operating approval.

LIMIT 05Missing data and anomalies

Deleting observations related to initial regime, first few trades, path dependence of compounding, and warm-up handling may improve the result. Compare no deletion, conservative imputation, and worst-case imputation, and display how inception-specific expectancy, lower-decile performance, sign stability, first-k-trade contribution, and recovery duration changes.

LIMIT 06Negative controls

Run the control “remove the first k trades from every cohort to test whether inception dependence is explained by a few early observations.” If the control performs similarly, suspect processing rules or common market drift before attributing performance to the strategy.

LIMIT 07Prospective monitoring

After a provisional pass, log lower cohort quantiles, first-ten-trade contribution, sign stability, and longest recovery sequentially and stop on persistent departures from the predeclared predictive range. Diagnose implementation drift before reoptimizing history.

LIMIT 08Common-mode failure and reporting

Multiple methods can agree because they share the same bad input or the same mechanism “initial regime, first few trades, path dependence of compounding, and warm-up handling.” Give lower-tail outcomes, failed scenarios, and unresolved mismatches the same visual prominence as favorable results; test count is not proof of correctness.

10A

Independent and adversarial findings for start-date dependence

The start-date dependence case has a separate review line for formulas, chart encodings, data definitions, and falsifiability so agreement on one layer cannot mask failure on another.

The formula audit checks numerator, denominator, sign, unit, domain, and every conditioning assumption as one system. The material caution for this case is: Sign stability across inception dates is an empirical diagnostic, not a population probability. When the median is zero, its sign is indeterminate rather than mechanically classified. Inception cohorts overlap, so they must not be treated as independent observations for standard-error calculations. A correct symbolic expression can still calculate the wrong quantity when a column, currency, time unit, or fee sign is misdefined, so those mappings are part of the mathematical audit.

The figure audit assigns distinct jobs: Figure 1 diagnoses inception-date dependence; Figure 2 maps joint sensitivity; Figure 3 shows the dependence-preserving distribution of distribution across inception dates; Figure 4 traces causal propagation. Color denotes distance to a predeclared gate, not probability or observed performance. Axis units, zero, quantiles, censoring, and bounds must agree with captions and tables. A smooth SVG line is explanatory geometry, not evidence of estimation precision.

The adversarial test does not cherry-pick one hostile scenario. It uses the negative control “remove the first k trades from every cohort to test whether inception dependence is explained by a few early observations,” resamples circular inception shifts and time blocks that preserve market regimes at several block lengths, and bounds cash-flow timing, minimum trade size, and historical data required for warm-up as unobserved factors. Repetitions, seeds, exclusions, block specifications, and plotting range are frozen before results so the implementer cannot tune the audit after seeing the answer.

The independent conclusion is restricted to whether “the lower inception-date quantile clears the threshold and profit is not concentrated in one favorable start.” It does not certify a good strategy or future profit. Any material reconciliation error, formula-domain violation, table-figure contradiction, sign reversal across defensible block lengths, or failure to outperform the negative control produces hold or reject. Prospectively, monitor lower cohort quantiles, first-ten-trade contribution, sign stability, and longest recovery.

10

Methodological references for start-date dependence

Primary methods and official platform documentation.

  1. Politis, D. N. & Romano, J. P. (1994). The Stationary Bootstrap. JASA.
  2. Newey, W. K. & West, K. D. (1987). A Simple, Positive Semi-definite, Heteroskedasticity and Autocorrelation Consistent Covariance Matrix. Econometrica.
  3. TradingView Pine Script® documentation: Strategies.
  4. Efron, B. (1979). Bootstrap Methods: Another Look at the Jackknife. Annals of Statistics.
  5. Lo, A. W. (2002). The Statistics of Sharpe Ratios. Financial Analysts Journal.
  6. White, H. (2000). A Reality Check for Data Snooping. Econometrica.

References for the start-date dependence case provide methodological context; they do not validate the synthetic numbers in this article or certify any backtest result. TradingView documentation is used for platform semantics, while statistical papers motivate uncertainty and selection controls.

08

Frequently asked questions about start-date dependence

Should every possible start date be tested?

Use a grid dense enough for the strategy’s horizon—daily for very short systems, weekly or monthly for slower ones—and explain the choice.

Why keep the same end date?

It isolates the effect of joining the historical path later. A second analysis with rolling ends tests window selection more broadly.

Is this just walk-forward testing?

They overlap, but start-date sweeps isolate sensitivity to the anchor and early path. Walk-forward focuses on repeated train/test behavior across time.

Backtest Analysis

Can a backtest exposed to start-date dependence be trusted?

Do not judge the start-date dependence case from a finished equity curve alone. Use the TradingView trade list to inspect the mechanism-specific concentration, path, cost, timing, and dependence evidence shown on this page.

Important limitations for the start-date dependence analysis

This article provides educational, descriptive analysis of constructed backtest failure examples. It is not investment advice, a buy or sell signal, a forecast or a promise of performance. Backtest results depend on data, code, broker-emulator assumptions, costs, sizing and market structure. TradingView is a trademark of TradingView, Inc.; SG Group is independent and does not claim endorsement or sponsorship by TradingView.

Counterpart: 開始日を変えると崩れる戦略|スタート日依存のバックテスト