Skip to the investigation

CASE 08

Profit Factor Without Drawdown Duration Is Incomplete

Both strategies report PF 1.55 and maximum drawdown 12%. One recovers in 18 days. The other needs 310 days.

One failure mode. One validation verdict.Focused analysis · educationally constructed educational figurestime underwaterdrawdown recovery timebacktest Profit Factorstrategy stagnation
BACKTEST DIAGNOSTIC PANELCase DD-TIME. Educational illustrative values, not observed market data.BACKTEST DIAGNOSTIC PANELCase DD-TIME · educational illustrative valuesProfit factor1.55Maximum drawdown−12%Longest underwater310 daysfailure boundaryPoint estimateDependenceTail stressExecutionSelectionReproductionA composite score summarizes evidence; it does not prove robustness.

01

Validation verdict for drawdown duration

Profit Factor summarizes the amount of gross profit relative to gross loss. It is silent about the calendar path, the length of stagnation and the probability that a user abandons the strategy before recovery.

All figures in this article are educationally constructed examples created to explain the failure mode. They are not real strategy results or recommended thresholds.
02

What the headline metric obscures about drawdown duration

PF 1.55 appears to describe efficiency: 1.55 units of gross profit for each unit of gross loss. When two strategies share that value, they can look comparable even if their losses are arranged very differently through time.

A year-long underwater period can be economically and operationally more severe than a fast 12% drop and recovery. Capital remains tied up, subscriptions and infrastructure continue, and the strategy may miss other opportunities while producing no new high.

03

How drawdown duration enters the backtest

PF discards time

The formula sums gains and losses without retaining when they happened or how long the equity stayed below peak.

Equal depth can hide unequal pain

Two 12% drawdowns can have radically different recovery profiles, number of failed rallies and time under water.

Sparse wins create long stagnation

A strategy can eventually earn enough gross profit for a good PF while waiting months between the few trades that repair the curve.

User behavior becomes part of failure

A result that is mathematically recoverable may be practically abandoned or resized during a prolonged drought.

04

Compact reconstruction of drawdown duration

CASE 08 · Profit Factor drawdown durationFocused analysis · educationally constructed educational figures
Strategy Profit Factor Max drawdown Longest recovery Time underwater
A: fast reset 1.55 −12% 18 days 21% of sample
B: slow bleed 1.55 −12% 310 days 63% of sample
B after 1.5× cost 1.19 −18% Not recovered 78% of sample
B, first half only 1.62 −9% 74 days 41% of sample

PF and maximum depth make A and B look identical in the first two rows. The underwater clock reveals that B spends most of the sample below its prior peak and fails to recover under a modest cost stress.

05

The test that can overturn the drawdown duration verdict

Treat every drawdown as an event with depth, start, trough, recovery date and time below peak. Report the distribution of durations, not only the worst depth.

Build an underwater curve and label the longest completed and current unfinished drawdowns.
Measure median, 90th-percentile and maximum recovery duration.
Calculate the share of calendar time below the previous equity peak.
Recompute duration after cost stress and after removing the top recovery trade.
Compare strategy versions on both depth and duration; do not let a small depth hide persistent stagnation.
06

What trade-list analysis can and cannot identify about drawdown duration

Export-level red flags for drawdown duration

  • PF above 1.5 but more than half the sample is underwater
  • The current drawdown is still unrecovered at the test end
  • One trade ends most long drawdowns
  • Recovery duration expands sharply under small costs
  • Backtest coverage ends soon after a new equity high

What the export reveals about drawdown duration

  • Equity and underwater curves with drawdown start, trough, recovery and unfinished status
  • Maximum and percentile recovery durations plus time-under-water share
  • Cost and outlier sensitivity of recovery, not only depth
  • Rolling periods where PF remains acceptable while equity stagnates

What drawdown duration still requires from settings, code, or market data

  • A historical recovery time is not a deadline for the next drawdown. Future stagnation can last longer or never recover.
  • Trade-only data may approximate balance drawdown but not intratrade equity drawdown unless mark-to-market information is available.
ACADEMIC VALIDATION DOSSIER

Turn drawdown depth and underwater duration into a falsifiable backtest diagnosis.

Case file 08/20 · DD-TIME · one failure mechanism, one falsifiable protocol

01

Research abstract: drawdown duration

Case file 08/20 · DD-TIME · one failure mechanism, one falsifiable protocol

This article tests one central proposition: even with a high profit factor, a recovery lasting hundreds of days creates very different economic and behavioral costs at the same drawdown depth. The question is not merely whether the displayed net profit or win rate was arithmetically calculated. The deeper identification problem is whether we know what constitutes one observation, what information was available at the decision time, which assumptions are necessary for the profit to exist, and how much of the conclusion survives when those assumptions are perturbed. The research object is therefore not one performance table; it is the linked data-generation, fill-generation, estimation, selection, and capital-allocation process.

The primary estimand is investability incorporating time spent underwater, not merely gross profits divided by gross losses. The observation unit is defined as equity high-water-mark intervals and drawdown episodes. Without this definition, split fills, duplicated signals, common events, synthetic prices, or timestamp conversions can be double-counted as independent evidence. A larger row count does not necessarily contain more independent information. An academically defensible analysis fixes the relationship between the observation unit and the estimand before it reports sample size, standard error, or statistical confidence.

The principal sensitivity axes are drawdown depth, duration, recovery hazard, and right-censored observations. The hidden state is longest recovery time, right-censored unrecovered episodes, capital lock-up, and investor abandonment risk. In particular, excluding an unrecovered ending episode removes the most important long stagnation from the statistics. Means and medians alone are incapable of describing that mechanism, so the analysis combines central estimates with lower quantiles, expected shortfall, sign stability, boundary-hitting frequency, and contribution concentration. The objective is not to find one pessimistic number, but to map the full region in which the original conclusion changes sign or ceases to be economically usable.

The conclusion does not attempt to prove that a backtest is good. It separates the component that remains after attempted falsification from the component that disappears when assumptions are reconstructed. The governing decision principle is to add longest duration, duration distribution, Kaplan–Meier recovery curves, and capital-lock-up cost to PF and maximum drawdown gates. This is not trading advice; it is a research procedure for measuring how much evidentiary weight a TradingView trade export can carry. Liquidity not present in the file, broker-specific rules, future regimes, outages, and gaps require separate evidence, and statistical survival never guarantees future profit.

The numerical values illustrate the method for drawdown depth and underwater duration; they are not a real strategy, client record, or forecast.

02

Hypotheses and identification target for drawdown duration

investability incorporating time spent underwater, not merely gross profits divided by gross losses

Null hypothesis / H₀

H₀ for drawdown duration: The reported performance is not materially dependent on the suspected failure mechanism and survives reasonable perturbations.

Alternative hypothesis / H₁

H₁ for drawdown duration: The reported performance depends materially on the suspected failure mechanism and deteriorates after reconstruction, perturbation, or dependence-aware resampling.

Estimand

investability incorporating time spent underwater, not merely gross profits divided by gross losses

Observation unit

equity high-water-mark intervals and drawdown episodes

Latent mechanism

longest recovery time, right-censored unrecovered episodes, capital lock-up, and investor abandonment risk

Stress axes

drawdown depth, duration, recovery hazard, and right-censored observations

03

Formal estimands for drawdown duration

Definitions precede inference.

D_t=1−W_t/max_{s≤t}W_s, with H_t=max_{s≤t}W_s>0. Drawdown depth at time t from a positive running wealth peak. If recovery never occurs, set τ_recovery,j=∞ and right-censor at the observation end.
U_j=min(τ_recovery,j,τ_end)−τ_peak,j, δ_j=1{τ_recovery,j≤τ_end}Underwater duration and recovery indicator with unrecovered episodes retained as right-censored.
Ŝ_KM(u)=∏_{v_j≤u}(1−d_j/n_j)Kaplan-Meier estimate of the probability of remaining unrecovered, using recoveries d_j and the risk set n_j.
Profit Factor1.82
Max drawdown−14.6%
Longest underwater418 days
Time underwater71%
Median recovery96 days

The primary estimand is investability incorporating time spent underwater, not merely gross profits divided by gross losses. The observation unit is defined as equity high-water-mark intervals and drawdown episodes. Without this definition, split fills, duplicated signals, common events, synthetic prices, or timestamp conversions can be double-counted as independent evidence. A larger row count does not necessarily contain more independent information. An academically defensible analysis fixes the relationship between the observation unit and the estimand before it reports sample size, standard error, or statistical confidence.

The principal sensitivity axes are drawdown depth, duration, recovery hazard, and right-censored observations. The hidden state is longest recovery time, right-censored unrecovered episodes, capital lock-up, and investor abandonment risk. In particular, excluding an unrecovered ending episode removes the most important long stagnation from the statistics. Means and medians alone are incapable of describing that mechanism, so the analysis combines central estimates with lower quantiles, expected shortfall, sign stability, boundary-hitting frequency, and contribution concentration. The objective is not to find one pessimistic number, but to map the full region in which the original conclusion changes sign or ceases to be economically usable.

04

Illustrative recomputation design for drawdown duration

For the recovery duration reconstruction, table values are illustrative calculations used to expose a verdict reversal; they are not a user’s observed TradingView result.

ID Recomputation layer Operation Comparison Diagnostic purpose
S0 Reported result Restate the Strategy Tester aggregate Base Apparent conclusion
S1 Unit reconstruction equity high-water-mark intervals and drawdown episodes Reassess count and dependence Information correction
S2 Independent recomputation Rebuild price, size, cost, and currency row by row Separate reconciliation error Measurement validity
S3 Local stress drawdown depth, duration, recovery hazard, and right-censored observations Perturb one factor only Causal sensitivity
S4 Tail injection excluding an unrecovered ending episode removes the most important long stagnation from the statistics Recompute lower quantiles and boundary hits Capital preservation
S5 Dependence-aware resampling Generate paths across several block lengths Intervals and sign stability Estimation uncertainty
S6 Selection adjustment Log search, OOS review, and exclusions Correct maximum-selection bias Generalization
S7 Full gate add longest duration, duration distribution, Kaplan–Meier recovery curves, and capital-lock-up cost to PF and maximum drawdown gates Compare with predeclared thresholds Pass / hold / reject

The illustrative recomputation for drawdown duration changes one processing layer at a time, then combines only predeclared layers. S0 is never treated as ground truth; it is the statement to be audited. S1 and S2 ask whether the exported unit and arithmetic are coherent. S3 and S4 identify local sensitivity and tail failure. S5 changes the uncertainty model rather than the trade list. S6 adjusts for the search that preceded publication. S7 applies the same gate to every version. This order prevents an adverse result from being explained away by simultaneously changing several assumptions.

In the recovery duration figures, color and position encode diagnostic sensitivity only; they do not represent statistical significance or future P&L.

05

Diagnostic figures specific to drawdown duration

Four separate visual tests; no decorative chart reuse.

Underwater depth-duration surfaceSynthetic experiment; axes and thresholds are diagnostic, not forecasts.Underwater depth-duration surfaceSynthetic experiment; axes and thresholds are diagnostic, not forecasts.high-water markEducational normalized display. Read direction, slope, and boundary location—not the absolute level.
Figure 1. Primary diagnostic for drawdown depth and underwater duration. Values are methodological illustrations, not estimates of a real strategy or future return.
Scatter of drawdown depth versus underwater durationFigure 2. Scatter of drawdown depth versus underwater duration. Depth and time under water are separate risk dimensions; shallow drawdowns can still persist intolerably long. Values are illustrative recomputations, not observed performance or forecasts.Scatter of drawdown depth versus underwater durationA topic-specific estimand decomposed into one diagnostic viewunderwater duration (trades)maximum depth
Figure 2. Scatter of drawdown depth versus underwater duration. Depth and time under water are separate risk dimensions; shallow drawdowns can still persist intolerably long. Values are illustrative recomputations, not observed performance or forecasts.
Kaplan-Meier curve for remaining unrecoveredFigure 3. Kaplan-Meier curve for remaining unrecovered. Unrecovered episodes remain as right-censored observations rather than disappearing from the duration estimate. Values are illustrative recomputations, not observed performance or forecasts.Kaplan-Meier curve for remaining unrecoveredA topic-specific stress test designed to overturn the headline verdict+++underwater duration (trades)probability unrecovered
Figure 3. Kaplan-Meier curve for remaining unrecovered. Unrecovered episodes remain as right-censored observations rather than disappearing from the duration estimate. Values are illustrative recomputations, not observed performance or forecasts.
State machine for peaks, drawdowns, censoring, and recoveryFigure 4. State machine for peaks, drawdowns, censoring, and recovery. Unrecovered episodes enter the duration model as censored states rather than being dropped in favor of quick recoveries. Values are illustrative recomputations, not observed performance or forecasts.State machine for peaks, drawdowns, censoring, and recoveryA causal or processing structure separating observations, assumptions, and decisionsnew peakunderwaterrecoveredright-censoredduration model
Figure 4. State machine for peaks, drawdowns, censoring, and recovery. Unrecovered episodes enter the duration model as censored states rather than being dropped in favor of quick recoveries. Values are illustrative recomputations, not observed performance or forecasts.
The primary diagnostic decomposes maximum depth, longest underwater duration, Kaplan-Meier recovery survival, recovery hazard, and capital-lockup days along a causal axis. Read slope, curvature, and the first decision-boundary crossing as “drawdown depth, duration, recovery hazard, and right-censored observations” changes, not merely the height of the favorable point.
The two-dimensional surface exposes interaction among “drawdown depth, duration, recovery hazard, and right-censored observations.” Color is a normalized margin to a predeclared gate, not an empirical probability. A broad connected pass region is different evidence from a narrow isolated island.
The resampling statistic is recovery-time survival probability. Compare an IID benchmark with time blocks that do not break drawdown episodes or regime persistence across several block lengths, reporting the 2.5th, 50th, and 97.5th percentiles and verdict-reversal rate. Save seeds and repetitions.
The causal map traces “aggregate gains and losses → erase time → miss prolonged underwater periods → lock capital → inability to continue.” A displayed metric is an intermediate product, not the first cause; perturb the input or assumption, rebuild trades and capital boundaries, and return to the predeclared gate.
06

Multi-layer audit questions for drawdown duration

A result is only as strong as its weakest unresolved layer.

AUDIT LAYER 0101 · Fix the estimand

Under adversarial review, fix the estimand as “investability incorporating time spent underwater, not merely gross profits divided by gross losses.” Do not substitute net profit, win rate, or a visually smooth curve for that target. Declare the horizon, account currency, included frictions, and operating-stop boundary before calculation. Any post-result change creates a new hypothesis and version, preventing the question from being selected after the answer is known.

AUDIT LAYER 0202 · Reconstruct the observation unit

Reconstruct the observation unit as “equity high-water-mark intervals and drawdown episodes” before treating rows as independent evidence. Report raw rows, parent trades, decisions, event clusters, and the denominator used for each average or standard error. Recompute maximum depth, longest underwater duration, Kaplan-Meier recovery survival, recovery hazard, and capital-lockup days under more than one defensible aggregation rule so that a larger export is not mistaken for a larger information set.

AUDIT LAYER 0303 · Preserve provenance and settings

Preserve the hash of the TradingView export and the symbol, timeframe, session, timezone, order-processing settings, costs, account currency, and Pine version. For drawdown depth and underwater duration, longest recovery time, right-censored unrecovered episodes, capital lock-up, and investor abandonment risk directly affects reproducibility. Keep immutable source, normalized, and analysis layers separate, with every join, deletion, imputation, and conversion recorded in a transformation ledger.

AUDIT LAYER 0404 · Separate identification from assumption

The export identifies only what can be rebuilt from recorded time, price, quantity, and P&L. capital withdrawals, investor redemptions, margin calls, and mandate stop rules requires additional evidence. Mark each causal link as observed, bounded by assumption, or externally unverified. This prevents longest recovery time, right-censored unrecovered episodes, capital lock-up, and investor abandonment risk from being presented as a confirmed fact when the available data support only an interval or conditional conclusion.

AUDIT LAYER 0505 · Reconcile row-level arithmetic

Do not adopt the platform summary as ground truth. Independently convert each peak-to-recovery interval into a drawdown episode and retain the final unrecovered episode as right-censored rather than deleting it. Reconcile total and row-level differences by sign, date, symbol, and order type. If discrepancies concentrate in the exact state associated with drawdown depth and underwater duration, treat that concentration as a primary finding rather than dismissing it as rounding.

AUDIT LAYER 0606 · Quantify finite-sample uncertainty

Report maximum depth, longest underwater duration, Kaplan-Meier recovery survival, recovery hazard, and capital-lockup days with intervals or resampling distributions, not point estimates alone. Match the uncertainty method to sample size, skewness, heavy tails, censoring, and selection history. If normal, quantile, and dependence-aware methods disagree on the sign, classify the edge as unidentified and show the minimum detectable effect and lower decision bound.

AUDIT LAYER 0707 · Preserve serial and cluster dependence

Do not narrow uncertainty with an IID shuffle alone. Resample time blocks that do not break drawdown episodes or regime persistence using several fixed block lengths and stationary bootstrap. Preserve random seed, repetition count, wrap rule, and missing-data treatment. For each block specification, report the distribution of maximum depth, longest underwater duration, Kaplan-Meier recovery survival, recovery hazard, and capital-lockup days, the rejection-side tail mass, and the rate at which the verdict changes sign.

AUDIT LAYER 0808 · Measure tails and operating boundaries

Interrogate the mechanism “longest recovery time, right-censored unrecovered episodes, capital lock-up, and investor abandonment risk” with lower quantiles, expected shortfall, influence, cluster length, and boundary-hitting measures. Historical maximum loss is not a loss cap. Define several absorbing or operating boundaries—capital, margin, mandate drawdown, and recovery time—and record which boundary fails first under each stress.

AUDIT LAYER 0909 · Model execution and market frictions

A flat commission deduction is not an execution model for drawdown depth and underwater duration. Allocate spread, slippage, financing, borrow, roll, conversion, rounding, and rejected orders to the relevant unit. Recompute maximum depth, longest underwater duration, Kaplan-Meier recovery survival, recovery hazard, and capital-lockup days under base, upper-quantile, and crisis states while preserving the possibility that costs and losses worsen together.

AUDIT LAYER 1010 · Count the complete search path

Count the complete population of periods, symbols, timeframes, parameters, exits, filters, and metrics that were tried. Do not detach the attractive result for drawdown depth and underwater duration from rejected candidates, interim changes, or repeated validation reviews. Where appropriate, use PBO, SPA, and a Deflated Sharpe Ratio, and treat an unrecorded trial count as a material audit limitation.

AUDIT LAYER 1111 · Condition on market regimes

Test whether drawdown depth and underwater duration is concentrated in one trend, volatility, liquidity, rate, or session state. Define regimes prospectively or on training data only. Report statewise maximum depth, longest underwater duration, Kaplan-Meier recovery survival, recovery hazard, and capital-lockup days, occupancy, transition probabilities, and costs, then reweight the mixture to adverse but realistic future compositions.

AUDIT LAYER 1212 · Separate path, inception, and sizing

For the drawdown duration case, the same trade set can follow different capital paths under another inception date, order, initial balance, rounding rule, or stop condition. Separate fixed quantity, fixed R, and percentage sizing, then use circular shifts and block orderings to recompute drawdown, recovery, and boundary hits. Equal terminal P&L does not imply equal path risk.

AUDIT LAYER 1313 · Design counterfactual stress tests

Perturb “drawdown depth, duration, recovery hazard, and right-censored observations” one axis at a time before creating a joint sensitivity surface. Add the negative control “permute sequences with the same gross profit and gross loss to show that profit factor can remain fixed while recovery duration changes materially.” Predefine the grid and crisis rule so that neither the most favorable nor the most damaging cell is selected after inspection. Save the slope, curvature, and exact point where the decision boundary is crossed.

AUDIT LAYER 1414 · Verify through an independent implementation

Have a second implementation convert each peak-to-recovery interval into a drawdown episode and retain the final unrecovered episode as right-censored rather than deleting it, then compare critical row-level outputs. Regression fixtures should include empty files, duplicate timestamps, extreme costs, reverse ordering, missing values, and boundary cases. Agreement between implementations is insufficient if they share the same bad input, so separate data construction and review roles where feasible.

AUDIT LAYER 1515 · Use a predeclared decision gate

Predeclare the decision rule. This case passes only if “the strategy meets both depth limits and predeclared recovery-probability and capital-lockup limits.” Near a boundary, disclose interval width and economic materiality rather than a binary badge. If only one favorable block length, cost state, or implementation passes, classify the result as assumption-sensitive rather than robust.

AUDIT LAYER 1616 · Maintain a reproducibility ledger

The evidence ledger must store the input hash, code version, settings, exclusions, “drawdown depth, duration, recovery hazard, and right-censored observations,” block lengths, random seed, repetition count, and every scenario output. Keep exploratory and confirmatory results in separate namespaces and retain failed trials. When new TradingView data arrive, create a new version and track underwater days, unrecovered episodes, recovery probability, and depth-duration area rather than overwriting the old result.

AUDIT LAYER 1717 · Translate statistics into capital impact

Translate statistical changes into capital consequences. A shift in expectancy, lower quantile, recovery time, or boundary risk caused by drawdown depth and underwater duration should be mapped to trade count, capital, margin, and continuation. A small per-trade difference can compound under high turnover, while a rare loss can be decisive near an absorbing boundary.

AUDIT LAYER 1818 · Separate roles and enforce stop conditions

Separate hypothesis design, implementation, independent recalculation, and approval where practical. Stop automatically on material reconciliation error, unresolved missing data, non-reproducibility, or a predeclared threshold breach. Audit the chain “aggregate gains and losses → erase time → miss prolonged underwater periods → lock capital → inability to continue,” and monitor underwater days, unrecovered episodes, recovery probability, and depth-duration area prospectively without turning a historical pass into a promise of future profit.

07

Falsification protocol for drawdown duration

add longest duration, duration distribution, Kaplan–Meier recovery curves, and capital-lock-up cost to PF and maximum drawdown gates

Freeze the TradingView source for the drawdown duration audit

Store the export without alteration and record its hash, export time, strategy, symbol, timeframe, and settings. Preserve every column relevant to drawdown depth and underwater duration; deletions and imputations belong only in derived tables.

Reconstruct the observation unit for drawdown duration

Aggregate rows into “equity high-water-mark intervals and drawdown episodes,” and report raw rows, parent trades, events, and independent clusters. Recompute the critical result under another defensible aggregation.

Independently recompute the displayed drawdown duration result

Independently convert each peak-to-recovery interval into a drawdown episode and retain the final unrecovered episode as right-censored rather than deleting it. Reconcile row-level and aggregate outputs with Strategy Tester and preserve where discrepancies concentrate.

Isolate the drawdown duration mechanism

Treat drawdown depth and underwater duration as the principal mechanism and move “drawdown depth, duration, recovery hazard, and right-censored observations” one axis at a time while holding other settings fixed.

Map the operating boundary for drawdown duration

Combine the primary and interacting axes on a predeclared grid and recompute maximum depth, longest underwater duration, Kaplan-Meier recovery survival, recovery hazard, and capital-lockup days. Record the width and connectivity of the acceptable region and every boundary crossing.

Resample the dependence structure relevant to drawdown duration

Use time blocks that do not break drawdown episodes or regime persistence with several fixed block lengths and stationary bootstrap. Save every random seed, repetition count, and block specification.

Inspect influence points and operating boundaries for drawdown duration

For the recovery duration influence test, remove the largest contributor, top-k contributors, selected periods, and relevant regimes in sequence; then recompute lower-tail measures and the operating boundary.

Apply negative controls and conservative bounds to drawdown duration

Permute sequences with the same gross profit and gross loss to show that profit factor can remain fixed while recovery duration changes materially. Bound capital withdrawals, investor redemptions, margin calls, and mandate stop rules as unobserved factors rather than elevating the optimistic value into the final answer.

Apply the predeclared gate to drawdown duration

Do not move the threshold after seeing results. Compare with “the strategy meets both depth limits and predeclared recovery-probability and capital-lockup limits,” and distinguish pass, hold, and reject. Any unresolved material mismatch causes a hold.

Save a reproducible evidence package for drawdown duration

Bundle the source, transformation ledger, formulas, figures, all scenarios, failure logs, and code version for rerun in another environment. Prospectively monitor underwater days, unrecovered episodes, recovery probability, and depth-duration area.

08

Decision gate for drawdown duration

Reject the story before trusting the curve.

How to read the drawdown duration figures and equations

The figures for drawdown duration use illustrative recomputations constructed to expose this specific failure mode. Do not infer statistical significance from line position or color alone; first verify the estimand, units, denominator, censoring rule, and cost sign defined by the equations. A sensitivity surface is not a causal estimate. It shows how a conclusion changes only within the stated assumptions. Resampling should compare an IID shuffle with stationary and block bootstrap procedures across several block lengths so that loss clustering and regime persistence are not silently destroyed. Store the random seed, iteration count, block length, bandwidth, and missing-data treatment, and claim reproducibility only after an independent implementation reproduces the same aggregates.

This case passes only if “the strategy meets both depth limits and predeclared recovery-probability and capital-lockup limits” across reconstructed values, local perturbations, joint sensitivity, dependence-preserving resampling, and the negative control, with no material sign reversal or unresolved reconciliation error. A pass is limited evidence against the stated failure mode, not certification of future profit.

  • The estimand and observation unit were fixed before outcomes were reviewed
  • For recovery duration, any material disagreement between reported and independently recomputed values must be resolved or explicitly explained.
  • The recovery duration claim passes this gate only when its acceptable stress region is broad and connected rather than one isolated favorable island.
  • The sign of the recovery duration estimate must remain stable across defensible block lengths, saved seeds, and reasonable interval methods.
  • For drawdown duration, economic margin remains after deleting the largest and top-five contributors and key regimes
  • For drawdown duration, conservative cost, fill, and capital-boundary scenarios remain inside the stopping mandate
6/6required gates · not a performance forecast
09

Limitations, external validity, and reproducibility of the drawdown duration audit

Every inference has a boundary.

The first limitation is that a trade export does not contain the complete market state. If order-book depth, queue position, network latency, rejected orders, broker liquidity, or realized financing history is absent, investability incorporating time spent underwater, not merely gross profits divided by gross losses remains model-mediated. Model outputs should be displayed as scenario ranges and must not be formatted as though they were directly observed facts.

A second limitation specific to the drawdown duration analysis is structural change. A long historical sample does not guarantee a common population when market rules, participants, volatility, rates, spreads, data construction, or Pine execution semantics change. Do not increase nominal sample size by indiscriminately pooling old periods. Estimate rolling and regime-conditioned behavior and test parameter stability around detected changes.

A third limitation specific to the drawdown duration analysis is reuse of the diagnostic battery. Applying these tests repeatedly to the same data and editing the strategy until it passes turns the diagnostic process itself into another optimizer. Every post-test edit starts a new model version and requires untouched or prospective evidence. A test chosen after reading the outcome belongs to exploration and cannot be counted as independent confirmation.

A fourth limitation for the drawdown duration analysis is the distinction between statistical survival and operational suitability. Behavioral tolerance, locked capital, tax, regulation, outages, account terms, order-size limits, market-order restrictions, and liquidity discontinuities cannot be resolved from a CSV alone. The lab is a diagnostic for discovering hidden failure risk earlier; it is not investment advice, a performance warranty, or a guarantee of bounded loss. User-specific constraints remain a separate decision layer.

LIMIT 01Identification boundary

The estimand “investability incorporating time spent underwater, not merely gross profits divided by gross losses” is identified only within the columns present in the TradingView export and the stated assumptions. If capital withdrawals, investor redemptions, margin calls, and mandate stop rules cannot be observed, report bounds rather than a false point estimate.

LIMIT 02Structural change

Past estimates of drawdown depth and underwater duration need not belong to the same population after changes in rules, participants, volatility, costs, or data specifications. Track underwater days, unrecovered episodes, recovery probability, and depth-duration area in rolling and regime-specific windows.

LIMIT 03Reuse of the diagnostic

For recovery duration, repeatedly applying the same diagnostic battery and editing until it passes turns verification into another optimizer. Every post-audit change therefore creates a new model version and requires untouched evidence.

LIMIT 04Operational suitability

Even if the strategy meets both depth limits and predeclared recovery-probability and capital-lockup limits, the analysis does not establish tax, regulatory, behavioral, liquidity, order-size, or systems suitability. Separate statistical diagnosis from live-operating approval.

LIMIT 05Missing data and anomalies

Deleting observations related to longest recovery time, right-censored unrecovered episodes, capital lock-up, and investor abandonment risk may improve the result. Compare no deletion, conservative imputation, and worst-case imputation, and display how maximum depth, longest underwater duration, Kaplan-Meier recovery survival, recovery hazard, and capital-lockup days changes.

LIMIT 06Negative controls

Run the control “permute sequences with the same gross profit and gross loss to show that profit factor can remain fixed while recovery duration changes materially.” If the control performs similarly, suspect processing rules or common market drift before attributing performance to the strategy.

LIMIT 07Prospective monitoring

After a provisional pass, log underwater days, unrecovered episodes, recovery probability, and depth-duration area sequentially and stop on persistent departures from the predeclared predictive range. Diagnose implementation drift before reoptimizing history.

LIMIT 08Common-mode failure and reporting

Multiple methods can agree because they share the same bad input or the same mechanism “longest recovery time, right-censored unrecovered episodes, capital lock-up, and investor abandonment risk.” Give lower-tail outcomes, failed scenarios, and unresolved mismatches the same visual prominence as favorable results; test count is not proof of correctness.

10A

Independent and adversarial findings for drawdown duration

The recovery duration case has a separate review line for formulas, chart encodings, data definitions, and falsifiability so agreement on one layer cannot mask failure on another.

The formula audit checks numerator, denominator, sign, unit, domain, and every conditioning assumption as one system. The material caution for this case is: Dropping unrecovered drawdowns biases recovery time downward. Retain right censoring through U_j and δ_j, and use the correct at-risk set n_j at each recovery time in the Kaplan-Meier product. Fix the duration unit—bars, calendar days, or trading days—rather than mixing them. A correct symbolic expression can still calculate the wrong quantity when a column, currency, time unit, or fee sign is misdefined, so those mappings are part of the mathematical audit.

The figure audit assigns distinct jobs: Figure 1 diagnoses drawdown depth and underwater duration; Figure 2 maps joint sensitivity; Figure 3 shows the dependence-preserving distribution of recovery-time survival probability; Figure 4 traces causal propagation. Color denotes distance to a predeclared gate, not probability or observed performance. Axis units, zero, quantiles, censoring, and bounds must agree with captions and tables. A smooth SVG line is explanatory geometry, not evidence of estimation precision.

The adversarial test does not cherry-pick one hostile scenario. It uses the negative control “permute sequences with the same gross profit and gross loss to show that profit factor can remain fixed while recovery duration changes materially,” resamples time blocks that do not break drawdown episodes or regime persistence at several block lengths, and bounds capital withdrawals, investor redemptions, margin calls, and mandate stop rules as unobserved factors. Repetitions, seeds, exclusions, block specifications, and plotting range are frozen before results so the implementer cannot tune the audit after seeing the answer.

The independent conclusion is restricted to whether “the strategy meets both depth limits and predeclared recovery-probability and capital-lockup limits.” It does not certify a good strategy or future profit. Any material reconciliation error, formula-domain violation, table-figure contradiction, sign reversal across defensible block lengths, or failure to outperform the negative control produces hold or reject. Prospectively, monitor underwater days, unrecovered episodes, recovery probability, and depth-duration area.

10

Methodological references for drawdown duration

Primary methods and official platform documentation.

  1. Kaplan, E. L. & Meier, P. (1958). Nonparametric Estimation from Incomplete Observations. JASA.
  2. Lo, A. W. (2002). The Statistics of Sharpe Ratios. Financial Analysts Journal.
  3. TradingView Pine Script® documentation: Strategies.
  4. Efron, B. (1979). Bootstrap Methods: Another Look at the Jackknife. Annals of Statistics.
  5. Politis, D. N. & Romano, J. P. (1994). The Stationary Bootstrap. JASA.
  6. Newey, W. K. & West, K. D. (1987). A Simple, Positive Semi-definite, Heteroskedasticity and Autocorrelation Consistent Covariance Matrix. Econometrica.
  7. White, H. (2000). A Reality Check for Data Snooping. Econometrica.

References for the drawdown duration case provide methodological context; they do not validate the synthetic numbers in this article or certify any backtest result. TradingView documentation is used for platform semantics, while statistical papers motivate uncertainty and selection controls.

08

Frequently asked questions about drawdown duration

Is Profit Factor useless?

No. It is useful for gross gain/loss efficiency, but incomplete without trade count, expectancy, drawdown depth, duration, costs and concentration.

Which duration matters most?

Track longest completed recovery, current unfinished drawdown, percentile durations and total time underwater. Each answers a different operational question.

Why can duration worsen before depth?

Small repeated losses and weak recoveries can keep equity below peak for a long time without producing one dramatic new low.

Backtest Analysis

Can a backtest exposed to drawdown duration be trusted?

Do not judge the recovery duration case from a finished equity curve alone. Use the TradingView trade list to inspect the mechanism-specific concentration, path, cost, timing, and dependence evidence shown on this page.

Important limitations for the drawdown duration analysis

This article provides educational, descriptive analysis of constructed backtest failure examples. It is not investment advice, a buy or sell signal, a forecast or a promise of performance. Backtest results depend on data, code, broker-emulator assumptions, costs, sizing and market structure. TradingView is a trademark of TradingView, Inc.; SG Group is independent and does not claim endorsement or sponsorship by TradingView.

Counterpart: プロフィットファクターだけでは足りない|ドローダウン期間の盲点