CASE 20
Why 500 Trades May Not Be 500 Independent Observations
The report contains 500 trades. Two hundred and eighty opened during just twelve macro events, often across correlated symbols within minutes of each other.
Validation verdict for trade clustering and serial dependence
Sample size is not the number of rows when outcomes share signals, timestamps, symbols or market shocks. Dependence reduces the amount of independent evidence and makes naive confidence, resampling and diversification claims too optimistic.
What the headline metric obscures about trade clustering and serial dependence
“Five hundred trades” sounds like repeated confirmation. But ten currency pairs triggered by one dollar shock are not ten independent experiments. They are ten expressions of the same underlying event, and their losses can arrive together.
Randomly shuffling individual rows then breaks the dependence structure that created the risk. The simulation produces many mixed sequences that could never occur, understating clustered drawdown and overstating how precisely the mean outcome is known.
How trade clustering and serial dependence enters the backtest
Signals fire in bursts
Volatility, news or shared indicators can produce many entries in a narrow window rather than evenly distributed opportunities.
Assets carry common factors
Different tickers can share currency, sector, duration, beta or liquidity exposure and fail during the same shock.
Serial dependence links outcomes
Trend persistence, regime state and cooldown rules can make the next trade conditional on the previous one.
Row-level resampling destroys clusters
Bootstrapping isolated trades treats co-occurring positions as separable and dilutes real tail concentration.
Compact reconstruction of trade clustering and serial dependence
| Counting method | Nominal rows | Clusters | Effective sample size (example) | 95% CI half-width |
|---|---|---|---|---|
| Every row independent | 500 | 500 | 500 | ±0.18R |
| Same-hour grouped | 500 | 164 | 164 | ±0.31R |
| Event + factor grouped | 500 | 91 | 91 | ±0.42R |
| 12 dominant event blocks | 500 | 12 major | ≈38 | ±0.65R |
The exact effective sample size depends on the dependence model; the figures are illustrative. The diagnostic point is structural: confidence becomes much wider once repeated rows are recognized as correlated exposures rather than separate confirmations.
The test that can overturn the trade clustering and serial dependence verdict
Map trades into clusters before estimating confidence or sequence risk. Preserve time blocks and shared-factor groups when resampling so that a stress path can carry an entire burst of related losses.
What trade-list analysis can and cannot identify about trade clustering and serial dependence
Export-level red flags for trade clustering and serial dependence
- Many entries occur within the same minute, hour or news event
- Multiple symbols share one macro or sector factor
- Monte Carlo shuffles every row independently
- Confidence intervals use √N with no dependence check
- Portfolio trade count rises while peak concurrent risk also rises
What the export reveals about trade clustering and serial dependence
- Temporal bursts, overlapping positions and cluster concentration from timestamped trades
- Serial correlation and cross-strategy dependence in returns or P&L
- Nominal versus grouped sample counts under several transparent cluster definitions
- Block-resampled drawdown and sequence distributions that preserve related outcomes
What trade clustering and serial dependence still requires from settings, code, or market data
- Effective sample size is model-dependent; no single cluster rule is universally correct. Results should be shown across plausible definitions.
- A trade export may omit the market factors or event labels needed to identify economic dependence. External market data can improve the grouping.
Turn trade clustering and effective information into a falsifiable backtest diagnosis.
Case file 20/20 · DEP-N · one failure mechanism, one falsifiable protocol
Research abstract: trade clustering and serial dependence
Case file 20/20 · DEP-N · one failure mechanism, one falsifiable protocol
This article tests one central proposition: treating 500 trades as independent understates uncertainty and makes a strategy with roughly 100 independent observations look precisely estimated. The question is not merely whether the displayed net profit or win rate was arithmetically calculated. The deeper identification problem is whether we know what constitutes one observation, what information was available at the decision time, which assumptions are necessary for the profit to exist, and how much of the conclusion survives when those assumptions are perturbed. The research object is therefore not one performance table; it is the linked data-generation, fill-generation, estimation, selection, and capital-allocation process.
The primary estimand is the standard error of mean expectancy and effective sample size after correcting for serial dependence. The observation unit is defined as a trade cluster or time block sharing a common market shock. Without this definition, split fills, duplicated signals, common events, synthetic prices, or timestamp conversions can be double-counted as independent evidence. A larger row count does not necessarily contain more independent information. An academically defensible analysis fixes the relationship between the observation unit and the estimand before it reports sample size, standard error, or statistical confidence.
The principal sensitivity axes are block length, autocorrelation truncation, cluster definition, and HAC bandwidth. The hidden state is autocorrelation, overlapping holdings, simultaneous signals, and within-regime clustering. In particular, consecutive losses from one news event are one common shock, and treating them as separate obscures risk concentration. Means and medians alone are incapable of describing that mechanism, so the analysis combines central estimates with lower quantiles, expected shortfall, sign stability, boundary-hitting frequency, and contribution concentration. The objective is not to find one pessimistic number, but to map the full region in which the original conclusion changes sign or ceases to be economically usable.
The conclusion does not attempt to prove that a backtest is good. It separates the component that remains after attempted falsification from the component that disappears when assumptions are reconstructed. The governing decision principle is to combine ACF, Ljung–Box, HAC, and block bootstrap and report confidence using effective rather than raw sample size. This is not trading advice; it is a research procedure for measuring how much evidentiary weight a TradingView trade export can carry. Liquidity not present in the file, broker-specific rules, future regimes, outages, and gaps require separate evidence, and statistical survival never guarantees future profit.
The numerical values illustrate the method for trade clustering and effective information; they are not a real strategy, client record, or forecast.
Hypotheses and identification target for trade clustering and serial dependence
the standard error of mean expectancy and effective sample size after correcting for serial dependence
H₀ for trade clustering and serial dependence: The reported performance is not materially dependent on the suspected failure mechanism and survives reasonable perturbations.
H₁ for trade clustering and serial dependence: The reported performance depends materially on the suspected failure mechanism and deteriorates after reconstruction, perturbation, or dependence-aware resampling.
the standard error of mean expectancy and effective sample size after correcting for serial dependence
a trade cluster or time block sharing a common market shock
autocorrelation, overlapping holdings, simultaneous signals, and within-regime clustering
block length, autocorrelation truncation, cluster definition, and HAC bandwidth
Formal estimands for trade clustering and serial dependence
Definitions precede inference.
Ω̂=γ̂₀+2Σ_{k=1}^Kw_kγ̂_kLong-run variance estimated from kernel-weighted autocovariances; report the kernel and bandwidth.Var_HAC(R̄)=Ω̂/nHAC variance of the sample mean R̄ under serial correlation and heteroskedasticity.n_eff^(LRV)=n·γ̂₀/Ω̂, Ω̂>0Diagnostic information-equivalent sample size from long-run variance. It can exceed n under negative autocorrelation, so do not call it a literal count of independent trades; report the raw value and bandwidth sensitivity.The primary estimand is the standard error of mean expectancy and effective sample size after correcting for serial dependence. The observation unit is defined as a trade cluster or time block sharing a common market shock. Without this definition, split fills, duplicated signals, common events, synthetic prices, or timestamp conversions can be double-counted as independent evidence. A larger row count does not necessarily contain more independent information. An academically defensible analysis fixes the relationship between the observation unit and the estimand before it reports sample size, standard error, or statistical confidence.
The principal sensitivity axes are block length, autocorrelation truncation, cluster definition, and HAC bandwidth. The hidden state is autocorrelation, overlapping holdings, simultaneous signals, and within-regime clustering. In particular, consecutive losses from one news event are one common shock, and treating them as separate obscures risk concentration. Means and medians alone are incapable of describing that mechanism, so the analysis combines central estimates with lower quantiles, expected shortfall, sign stability, boundary-hitting frequency, and contribution concentration. The objective is not to find one pessimistic number, but to map the full region in which the original conclusion changes sign or ceases to be economically usable.
Illustrative recomputation design for trade clustering and serial dependence
For the serial dependence reconstruction, table values are illustrative calculations used to expose a verdict reversal; they are not a user’s observed TradingView result.
| ID | Recomputation layer | Operation | Comparison | Diagnostic purpose |
|---|---|---|---|---|
| S0 | Reported result | Restate the Strategy Tester aggregate | Base | Apparent conclusion |
| S1 | Unit reconstruction | a trade cluster or time block sharing a common market shock | Reassess count and dependence | Information correction |
| S2 | Independent recomputation | Rebuild price, size, cost, and currency row by row | Separate reconciliation error | Measurement validity |
| S3 | Local stress | block length, autocorrelation truncation, cluster definition, and HAC bandwidth | Perturb one factor only | Causal sensitivity |
| S4 | Tail injection | consecutive losses from one news event are one common shock, and treating them as separate obscures risk concentration | Recompute lower quantiles and boundary hits | Capital preservation |
| S5 | Dependence-aware resampling | Generate paths across several block lengths | Intervals and sign stability | Estimation uncertainty |
| S6 | Selection adjustment | Log search, OOS review, and exclusions | Correct maximum-selection bias | Generalization |
| S7 | Full gate | combine ACF, Ljung–Box, HAC, and block bootstrap and report confidence using effective rather than raw sample size | Compare with predeclared thresholds | Pass / hold / reject |
The illustrative recomputation for trade clustering and serial dependence changes one processing layer at a time, then combines only predeclared layers. S0 is never treated as ground truth; it is the statement to be audited. S1 and S2 ask whether the exported unit and arithmetic are coherent. S3 and S4 identify local sensitivity and tail failure. S5 changes the uncertainty model rather than the trade list. S6 adjusts for the search that preceded publication. S7 applies the same gate to every version. This order prevents an adverse result from being explained away by simultaneously changing several assumptions.
In the serial dependence figures, color and position encode diagnostic sensitivity only; they do not represent statistical significance or future P&L.
Diagnostic figures specific to trade clustering and serial dependence
Four separate visual tests; no decorative chart reuse.
Multi-layer audit questions for trade clustering and serial dependence
A result is only as strong as its weakest unresolved layer.
First, fix the estimand as “the standard error of mean expectancy and effective sample size after correcting for serial dependence.” Do not substitute net profit, win rate, or a visually smooth curve for that target. Declare the horizon, account currency, included frictions, and operating-stop boundary before calculation. Any post-result change creates a new hypothesis and version, preventing the question from being selected after the answer is known.
Reconstruct the observation unit as “a trade cluster or time block sharing a common market shock” before treating rows as independent evidence. Report raw rows, parent trades, decisions, event clusters, and the denominator used for each average or standard error. Recompute autocorrelation, long-run variance, HAC standard error, diagnostic effective sample size, cluster-robust interval, and block-length sensitivity under more than one defensible aggregation rule so that a larger export is not mistaken for a larger information set.
Preserve the hash of the TradingView export and the symbol, timeframe, session, timezone, order-processing settings, costs, account currency, and Pine version. For trade clustering and effective information, autocorrelation, overlapping holdings, simultaneous signals, and within-regime clustering directly affects reproducibility. Keep immutable source, normalized, and analysis layers separate, with every join, deletion, imputation, and conversion recorded in a transformation ledger.
The export identifies only what can be rebuilt from recorded time, price, quantity, and P&L. common factors, unrecorded concurrent positions, cross-asset correlation, and news timing requires additional evidence. Mark each causal link as observed, bounded by assumption, or externally unverified. This prevents autocorrelation, overlapping holdings, simultaneous signals, and within-regime clustering from being presented as a confirmed fact when the available data support only an interval or conditional conclusion.
Do not adopt the platform summary as ground truth. Independently cluster simultaneous signals, overlapping positions, common news, and regimes, then estimate long-run variance and the HAC variance of the mean from autocovariances. Reconcile total and row-level differences by sign, date, symbol, and order type. If discrepancies concentrate in the exact state associated with trade clustering and effective information, treat that concentration as a primary finding rather than dismissing it as rounding.
Report autocorrelation, long-run variance, HAC standard error, diagnostic effective sample size, cluster-robust interval, and block-length sensitivity with intervals or resampling distributions, not point estimates alone. Match the uncertainty method to sample size, skewness, heavy tails, censoring, and selection history. If normal, quantile, and dependence-aware methods disagree on the sign, classify the edge as unidentified and show the minimum detectable effect and lower decision bound.
Do not narrow uncertainty with an IID shuffle alone. Resample several fixed block lengths and stationary bootstrap with geometrically distributed block lengths using several fixed block lengths and stationary bootstrap. Preserve random seed, repetition count, wrap rule, and missing-data treatment. For each block specification, report the distribution of autocorrelation, long-run variance, HAC standard error, diagnostic effective sample size, cluster-robust interval, and block-length sensitivity, the rejection-side tail mass, and the rate at which the verdict changes sign.
Interrogate the mechanism “autocorrelation, overlapping holdings, simultaneous signals, and within-regime clustering” with lower quantiles, expected shortfall, influence, cluster length, and boundary-hitting measures. Historical maximum loss is not a loss cap. Define several absorbing or operating boundaries—capital, margin, mandate drawdown, and recovery time—and record which boundary fails first under each stress.
A flat commission deduction is not an execution model for trade clustering and effective information. Allocate spread, slippage, financing, borrow, roll, conversion, rounding, and rejected orders to the relevant unit. Recompute autocorrelation, long-run variance, HAC standard error, diagnostic effective sample size, cluster-robust interval, and block-length sensitivity under base, upper-quantile, and crisis states while preserving the possibility that costs and losses worsen together.
Count the complete population of periods, symbols, timeframes, parameters, exits, filters, and metrics that were tried. Do not detach the attractive result for trade clustering and effective information from rejected candidates, interim changes, or repeated validation reviews. Where appropriate, use PBO, SPA, and a Deflated Sharpe Ratio, and treat an unrecorded trial count as a material audit limitation.
Test whether trade clustering and effective information is concentrated in one trend, volatility, liquidity, rate, or session state. Define regimes prospectively or on training data only. Report statewise autocorrelation, long-run variance, HAC standard error, diagnostic effective sample size, cluster-robust interval, and block-length sensitivity, occupancy, transition probabilities, and costs, then reweight the mixture to adverse but realistic future compositions.
For the trade clustering and serial dependence case, the same trade set can follow different capital paths under another inception date, order, initial balance, rounding rule, or stop condition. Separate fixed quantity, fixed R, and percentage sizing, then use circular shifts and block orderings to recompute drawdown, recovery, and boundary hits. Equal terminal P&L does not imply equal path risk.
Perturb “block length, autocorrelation truncation, cluster definition, and HAC bandwidth” one axis at a time before creating a joint sensitivity surface. Add the negative control “show a fully randomized IID benchmark and quantify how much artificially narrower the interval becomes when dependence is destroyed.” Predefine the grid and crisis rule so that neither the most favorable nor the most damaging cell is selected after inspection. Save the slope, curvature, and exact point where the decision boundary is crossed.
Have a second implementation cluster simultaneous signals, overlapping positions, common news, and regimes, then estimate long-run variance and the HAC variance of the mean from autocovariances, then compare critical row-level outputs. Regression fixtures should include empty files, duplicate timestamps, extreme costs, reverse ordering, missing values, and boundary cases. Agreement between implementations is insufficient if they share the same bad input, so separate data construction and review roles where feasible.
Predeclare the decision rule. This case passes only if “the interval lower bound clears the threshold across defensible bandwidths, kernels, cluster definitions, and block lengths.” Near a boundary, disclose interval width and economic materiality rather than a binary badge. If only one favorable block length, cost state, or implementation passes, classify the result as assumption-sensitive rather than robust.
The evidence ledger must store the input hash, code version, settings, exclusions, “block length, autocorrelation truncation, cluster definition, and HAC bandwidth,” block lengths, random seed, repetition count, and every scenario output. Keep exploratory and confirmatory results in separate namespaces and retain failed trials. When new TradingView data arrive, create a new version and track autocorrelation, cluster count, long-run variance, bandwidth, intervals by block length, and random seed rather than overwriting the old result.
Translate statistical changes into capital consequences. A shift in expectancy, lower quantile, recovery time, or boundary risk caused by trade clustering and effective information should be mapped to trade count, capital, margin, and continuation. A small per-trade difference can compound under high turnover, while a rare loss can be decisive near an absorbing boundary.
Separate hypothesis design, implementation, independent recalculation, and approval where practical. Stop automatically on material reconciliation error, unresolved missing data, non-reproducibility, or a predeclared threshold breach. Audit the chain “large row count → assumption of independence → understated standard error → excessive confidence → reversal under joint losses,” and monitor autocorrelation, cluster count, long-run variance, bandwidth, intervals by block length, and random seed prospectively without turning a historical pass into a promise of future profit.
Falsification protocol for trade clustering and serial dependence
combine ACF, Ljung–Box, HAC, and block bootstrap and report confidence using effective rather than raw sample size
Freeze the TradingView source for the trade clustering and serial dependence audit
Store the export without alteration and record its hash, export time, strategy, symbol, timeframe, and settings. Preserve every column relevant to trade clustering and effective information; deletions and imputations belong only in derived tables.
Reconstruct the observation unit for trade clustering and serial dependence
Aggregate rows into “a trade cluster or time block sharing a common market shock,” and report raw rows, parent trades, events, and independent clusters. Recompute the critical result under another defensible aggregation.
Independently recompute the displayed trade clustering and serial dependence result
Independently cluster simultaneous signals, overlapping positions, common news, and regimes, then estimate long-run variance and the HAC variance of the mean from autocovariances. Reconcile row-level and aggregate outputs with Strategy Tester and preserve where discrepancies concentrate.
Isolate the trade clustering and serial dependence mechanism
Treat trade clustering and effective information as the principal mechanism and move “block length, autocorrelation truncation, cluster definition, and HAC bandwidth” one axis at a time while holding other settings fixed.
Map the operating boundary for trade clustering and serial dependence
Combine the primary and interacting axes on a predeclared grid and recompute autocorrelation, long-run variance, HAC standard error, diagnostic effective sample size, cluster-robust interval, and block-length sensitivity. Record the width and connectivity of the acceptable region and every boundary crossing.
Resample the dependence structure relevant to trade clustering and serial dependence
Use several fixed block lengths and stationary bootstrap with geometrically distributed block lengths with several fixed block lengths and stationary bootstrap. Save every random seed, repetition count, and block specification.
Inspect influence points and operating boundaries for trade clustering and serial dependence
For the serial dependence influence test, remove the largest contributor, top-k contributors, selected periods, and relevant regimes in sequence; then recompute lower-tail measures and the operating boundary.
Apply negative controls and conservative bounds to trade clustering and serial dependence
Show a fully randomized iid benchmark and quantify how much artificially narrower the interval becomes when dependence is destroyed. Bound common factors, unrecorded concurrent positions, cross-asset correlation, and news timing as unobserved factors rather than elevating the optimistic value into the final answer.
Apply the predeclared gate to trade clustering and serial dependence
Do not move the threshold after seeing results. Compare with “the interval lower bound clears the threshold across defensible bandwidths, kernels, cluster definitions, and block lengths,” and distinguish pass, hold, and reject. Any unresolved material mismatch causes a hold.
Save a reproducible evidence package for trade clustering and serial dependence
Bundle the source, transformation ledger, formulas, figures, all scenarios, failure logs, and code version for rerun in another environment. Prospectively monitor autocorrelation, cluster count, long-run variance, bandwidth, intervals by block length, and random seed.
Decision gate for trade clustering and serial dependence
Reject the story before trusting the curve.
How to read the trade clustering and serial dependence figures and equations
The figures for trade clustering and serial dependence use illustrative recomputations constructed to expose this specific failure mode. Do not infer statistical significance from line position or color alone; first verify the estimand, units, denominator, censoring rule, and cost sign defined by the equations. A sensitivity surface is not a causal estimate. It shows how a conclusion changes only within the stated assumptions. Resampling should compare an IID shuffle with stationary and block bootstrap procedures across several block lengths so that loss clustering and regime persistence are not silently destroyed. Store the random seed, iteration count, block length, bandwidth, and missing-data treatment, and claim reproducibility only after an independent implementation reproduces the same aggregates.
This case passes only if “the interval lower bound clears the threshold across defensible bandwidths, kernels, cluster definitions, and block lengths” across reconstructed values, local perturbations, joint sensitivity, dependence-preserving resampling, and the negative control, with no material sign reversal or unresolved reconciliation error. A pass is limited evidence against the stated failure mode, not certification of future profit.
- The estimand and observation unit were fixed before outcomes were reviewed
- For serial dependence, any material disagreement between reported and independently recomputed values must be resolved or explicitly explained.
- The serial dependence claim passes this gate only when its acceptable stress region is broad and connected rather than one isolated favorable island.
- The sign of the serial dependence estimate must remain stable across defensible block lengths, saved seeds, and reasonable interval methods.
- For trade clustering and serial dependence, economic margin remains after deleting the largest and top-five contributors and key regimes
- For trade clustering and serial dependence, conservative cost, fill, and capital-boundary scenarios remain inside the stopping mandate
Limitations, external validity, and reproducibility of the trade clustering and serial dependence audit
Every inference has a boundary.
The first limitation is that a trade export does not contain the complete market state. If order-book depth, queue position, network latency, rejected orders, broker liquidity, or realized financing history is absent, the standard error of mean expectancy and effective sample size after correcting for serial dependence remains model-mediated. Model outputs should be displayed as scenario ranges and must not be formatted as though they were directly observed facts.
A second limitation specific to the trade clustering and serial dependence analysis is structural change. A long historical sample does not guarantee a common population when market rules, participants, volatility, rates, spreads, data construction, or Pine execution semantics change. Do not increase nominal sample size by indiscriminately pooling old periods. Estimate rolling and regime-conditioned behavior and test parameter stability around detected changes.
A third limitation specific to the trade clustering and serial dependence analysis is reuse of the diagnostic battery. Applying these tests repeatedly to the same data and editing the strategy until it passes turns the diagnostic process itself into another optimizer. Every post-test edit starts a new model version and requires untouched or prospective evidence. A test chosen after reading the outcome belongs to exploration and cannot be counted as independent confirmation.
A fourth limitation for the trade clustering and serial dependence analysis is the distinction between statistical survival and operational suitability. Behavioral tolerance, locked capital, tax, regulation, outages, account terms, order-size limits, market-order restrictions, and liquidity discontinuities cannot be resolved from a CSV alone. The lab is a diagnostic for discovering hidden failure risk earlier; it is not investment advice, a performance warranty, or a guarantee of bounded loss. User-specific constraints remain a separate decision layer.
The estimand “the standard error of mean expectancy and effective sample size after correcting for serial dependence” is identified only within the columns present in the TradingView export and the stated assumptions. If common factors, unrecorded concurrent positions, cross-asset correlation, and news timing cannot be observed, report bounds rather than a false point estimate.
Past estimates of trade clustering and effective information need not belong to the same population after changes in rules, participants, volatility, costs, or data specifications. Track autocorrelation, cluster count, long-run variance, bandwidth, intervals by block length, and random seed in rolling and regime-specific windows.
For serial dependence, repeatedly applying the same diagnostic battery and editing until it passes turns verification into another optimizer. Every post-audit change therefore creates a new model version and requires untouched evidence.
Even if the interval lower bound clears the threshold across defensible bandwidths, kernels, cluster definitions, and block lengths, the analysis does not establish tax, regulatory, behavioral, liquidity, order-size, or systems suitability. Separate statistical diagnosis from live-operating approval.
Deleting observations related to autocorrelation, overlapping holdings, simultaneous signals, and within-regime clustering may improve the result. Compare no deletion, conservative imputation, and worst-case imputation, and display how autocorrelation, long-run variance, HAC standard error, diagnostic effective sample size, cluster-robust interval, and block-length sensitivity changes.
Run the control “show a fully randomized IID benchmark and quantify how much artificially narrower the interval becomes when dependence is destroyed.” If the control performs similarly, suspect processing rules or common market drift before attributing performance to the strategy.
After a provisional pass, log autocorrelation, cluster count, long-run variance, bandwidth, intervals by block length, and random seed sequentially and stop on persistent departures from the predeclared predictive range. Diagnose implementation drift before reoptimizing history.
Multiple methods can agree because they share the same bad input or the same mechanism “autocorrelation, overlapping holdings, simultaneous signals, and within-regime clustering.” Give lower-tail outcomes, failed scenarios, and unresolved mismatches the same visual prominence as favorable results; test count is not proof of correctness.
Independent and adversarial findings for trade clustering and serial dependence
The serial dependence case has a separate review line for formulas, chart encodings, data definitions, and falsifiability so agreement on one layer cannot mask failure on another.
The formula audit checks numerator, denominator, sign, unit, domain, and every conditioning assumption as one system. The material caution for this case is: Long-run variance Ω̂ depends on the kernel and bandwidth and must be checked for a valid positive estimate. n_eff^(LRV) is diagnostic information-equivalence, not a literal count of independent trades, and can exceed n under negative autocorrelation. Report the raw value and sensitivity rather than silently capping it. A correct symbolic expression can still calculate the wrong quantity when a column, currency, time unit, or fee sign is misdefined, so those mappings are part of the mathematical audit.
The figure audit assigns distinct jobs: Figure 1 diagnoses trade clustering and effective information; Figure 2 maps joint sensitivity; Figure 3 shows the dependence-preserving distribution of mean expectancy after HAC correction; Figure 4 traces causal propagation. Color denotes distance to a predeclared gate, not probability or observed performance. Axis units, zero, quantiles, censoring, and bounds must agree with captions and tables. A smooth SVG line is explanatory geometry, not evidence of estimation precision.
The adversarial test does not cherry-pick one hostile scenario. It uses the negative control “show a fully randomized IID benchmark and quantify how much artificially narrower the interval becomes when dependence is destroyed,” resamples several fixed block lengths and stationary bootstrap with geometrically distributed block lengths at several block lengths, and bounds common factors, unrecorded concurrent positions, cross-asset correlation, and news timing as unobserved factors. Repetitions, seeds, exclusions, block specifications, and plotting range are frozen before results so the implementer cannot tune the audit after seeing the answer.
The independent conclusion is restricted to whether “the interval lower bound clears the threshold across defensible bandwidths, kernels, cluster definitions, and block lengths.” It does not certify a good strategy or future profit. Any material reconciliation error, formula-domain violation, table-figure contradiction, sign reversal across defensible block lengths, or failure to outperform the negative control produces hold or reject. Prospectively, monitor autocorrelation, cluster count, long-run variance, bandwidth, intervals by block length, and random seed.
Methodological references for trade clustering and serial dependence
Primary methods and official platform documentation.
- Newey, W. K. & West, K. D. (1987). A Simple, Positive Semi-definite, Heteroskedasticity and Autocorrelation Consistent Covariance Matrix. Econometrica.
- Politis, D. N. & Romano, J. P. (1994). The Stationary Bootstrap. JASA.
- Lo, A. W. (2002). The Statistics of Sharpe Ratios. Financial Analysts Journal.
- Efron, B. (1979). Bootstrap Methods: Another Look at the Jackknife. Annals of Statistics.
- White, H. (2000). A Reality Check for Data Snooping. Econometrica.
- TradingView Pine Script® documentation: Strategies.
References for the trade clustering and serial dependence case provide methodological context; they do not validate the synthetic numbers in this article or certify any backtest result. TradingView documentation is used for platform semantics, while statistical papers motivate uncertainty and selection controls.
Frequently asked questions about trade clustering and serial dependence
Does dependence make the strategy invalid?
No. It changes how much evidence the sample contains and how losses should be simulated. A clustered strategy can still be useful if capital is sized for cluster risk.
Is autocorrelation enough to measure this?
It captures serial structure in one series but may miss simultaneous cross-asset and event dependence. Combine time, overlap and factor views.
Should I count one cluster as one trade?
Not necessarily. Keep the raw rows, but supplement them with cluster-level summaries and confidence estimates that respect dependence.
Can a backtest exposed to trade clustering and serial dependence be trusted?
Do not judge the serial dependence case from a finished equity curve alone. Use the TradingView trade list to inspect the mechanism-specific concentration, path, cost, timing, and dependence evidence shown on this page.
Important limitations for the trade clustering and serial dependence analysis
This article provides educational, descriptive analysis of constructed backtest failure examples. It is not investment advice, a buy or sell signal, a forecast or a promise of performance. Backtest results depend on data, code, broker-emulator assumptions, costs, sizing and market structure. TradingView is a trademark of TradingView, Inc.; SG Group is independent and does not claim endorsement or sponsorship by TradingView.
Counterpart: 500回の取引が500個の独立データとは限らない

