Lot sizing decision guide · 17

Compare strategies on the same pre-trade risk unit

This is an analytics article, not a landing page that chooses the next lot. Convert each recorded trade to units of its pre-trade loss budget, then test whether a higher monetary profit was simply produced by taking more size.

Three points to establish first

  • Raw profit does not separate performance from loss-budget scale.
  • An R comparison needs the budget recorded before each trade and net P&L.
  • Average R omits sequence, distribution, drawdown and selection bias.

Durable reference map

A three-stage method to reuse whenever conditions change

Keep the sequence of inputs, calculation and exception testing stable instead of relying on a market forecast.

  1. Align inputs and unitsGo to equations and definitionsTrade result in R・Sample average R・Maximum drawdown in R
  2. Reconcile the worked exampleGo to table and calculation stepsA hypothetical reversal between dollar and R rankings
  3. Test exceptions and next checksGo to rules and counterexampleUse the loss budget saved before entry as the R denominator.

An analysis of records, not a size recommendation

A lot calculation derives a future quantity from a current budget and stop. This analysis starts with completed records and places results generated at different quantities on a common loss-budget unit.

A favorable comparison does not automatically justify increasing the next order.

Fix one R from information available before entry

One R is the monetary loss budget set before an order. Defining one R later as the realized loss makes the denominator depend on the outcome and defeats the comparison.

When budgets vary by trade, divide each row’s net P&L by that row’s pre-trade budget before aggregating.

  • Timestamp the pre-trade budget.
  • Deduct commission and financing from P&L.
  • Convert records to one currency.

When the monetary ranking reverses

A strategy run at a larger lot can show more dollars even with the same or weaker edge. Dividing by the amount put at risk can rank a smaller deployment above it.

If every strategy used the same budget, costs and observation period, normalization may leave the ranking unchanged.

Inspect what an average omits

Two samples with the same average R can have different dependence on a few large winners, different losing sequences and different drawdowns. Preserve trade order and inspect distribution and holding periods.

A small sample should not be described as durable evidence of edge.

  • Median and quantiles
  • Longest losing sequence and maximum R drawdown
  • Gross-to-net difference
  • Market period and instrument mix

Discount a winner selected from many trials

Testing many strategies or parameter sets and reporting only the best increases the chance of selecting noise. Retain the number of trials, discarded candidates and the separation between development and validation data.

If a Sharpe ratio is added, align return frequency, period, risk-free rate and treatment of non-normal returns.

Report a comparison without turning it into a claim

Show monetary P&L, total R, average R, sample size and maximum R drawdown together. Label hypothetical, backtested and live records accurately. The table below is arithmetic, not performance evidence.

  • Do not mix gross and net.
  • Use the same period and method.
  • Do not imply future performance.

Calculation framework

Trade result in R

Read the role of each equation first, then follow the numerical example to check the decision path.

01

Trade result in R

EquationR_i = PnL_i,net / B_i,pretrade
PnL_i,net: trade P&L after stated costs
B_i,pretrade: loss budget saved before the order

In plain language: Normalize trades of different size by the money exposed under the pre-trade plan.

When this conclusion does not apply: Do not divide when the saved pre-trade budget is missing, zero or negative; record that an exact R value cannot be reconstructed.

02

Sample average R

Equationmean_R = (Σ R_i) / n
n: trades included under one consistent selection rule

In plain language: This is an average per trade, not a description of sequence or distribution.

When this conclusion does not apply: Do not calculate mean R when n = 0, and do not call a small-sample difference durable edge.

03

Maximum drawdown in R

EquationMDD_R = max_t(peak_cumR_t − cumR_t)
cumR_t: cumulative R in original trade order, with the starting baseline cumR_0 = 0
peak_cumR_t: max_{0≤s≤t}(cumR_s), the highest cumulative R through t including the starting baseline of zero

In plain language: Two samples with the same total R can have different paths of decline.

When this conclusion does not apply: It cannot be calculated from an aggregate that has lost trade order.

A checkable example

A hypothetical reversal between dollar and R rankings

Each strategy is assumed to have 100 trades and a constant pre-trade budget. These are teaching numbers, not realized or forecast performance.
StrategyTradestradesPre-trade loss budgetUSDNet P&LUSDTotal resultRAverage resultR/trade
A100100012000120.12
B1005008000160.16

Calculation steps

  1. A: USD 12,000 / USD 1,000 = 12 R; 12 R / 100 = 0.12 R per trade.

  2. B: USD 8,000 / USD 500 = 16 R; 16 R / 100 = 0.16 R per trade.

  3. A has more dollars, while B has more net P&L per unit of planned loss.

Result: The dollar ranking alone cannot separate result from the scale of the loss budget.

When this conclusion does not apply: When budgets, trade counts and cost treatment are identical, dollar P&L and total R have the same ranking.

Frame R normalization as analysis of completed records

A forward lot calculation derives quantity from a current stop and budget. R normalization starts with completed trades and expresses each net result relative to the monetary loss budget recorded before that trade. A favorable historical ranking does not itself authorize a larger next order.

This distinction prevents circularity. If strategy performance determines the budget and the enlarged budget then defines the denominator used to praise performance, the comparison can be engineered. The R unit must come from information available before entry, while later P&L remains the numerator.

The analysis can compare different deployment scales, but it cannot remove selection bias, path dependence, or execution differences. R is a common monetary-risk unit under stated records, not a universal measure of strategy quality or a recommendation of exposure.

The report should label each use of R as planned, open, or realized. The denominator here is the pre-trade budget and the numerator is realized net P&L. Using the same letter for a current stop estimate without a qualifier can cause analysts to combine states that were measured at different times.

Fix one R before the outcome is known

For each trade, one R is the authorized account-currency loss budget saved before entry. When budgets vary, divide each trade’s net P&L by its own pre-trade amount before aggregating. Defining one R later as realized loss makes the denominator depend on the outcome and destroys comparability.

A missing, zero, or negative pre-trade budget means exact R cannot be reconstructed. Substituting the average budget, the final stop loss, or a later policy value creates false precision. The record should mark the trade unavailable for exact normalization while retaining its monetary result.

Stop amendments do not retroactively change one R unless the prospective policy explicitly defines how the budget unit evolves. Original planned risk, current open risk, and realized result can all be reported. Conflating them lets favorable management decisions shrink the denominator after the fact.

A reconstruction flag should distinguish directly saved budgets from values inferred later. Inferred rows can be analyzed separately with uncertainty, but they should not silently join the exact series. This preserves sample size information without giving incomplete records the same evidentiary weight as contemporaneous data.

Use net P&L on a common currency and cost basis

The numerator should include the declared transaction costs and use one account currency. Gross results from one strategy cannot be compared with net results from another. When conversion is required, retain the factor and time applicable to each fill or the documented statement method.

Partial exits and multiple fills should reconcile to the parent trade before R is calculated. Counting each profitable child exit as a separate trade while assigning the residual loss once distorts both sample count and normalized result. Event identity and cost-basis rules belong in the source ledger.

If cost data are incomplete, show gross R separately and do not label it net. A small cost difference can change ranking when averages are close. Methodological transparency is more useful than forcing every row into one apparently precise comparison. The missing-cost count should remain beside the aggregate.

Reproduce the reversal between dollar and R rankings

The hypothetical Strategy A has 100 trades, USD 1,000 pre-trade budget per trade, and USD 12,000 net P&L. Dividing gives 12 total R and 0.12 R per trade. Strategy B uses USD 500 per trade and earns USD 8,000, giving 16 R and 0.16 average R. The arithmetic alone causes the ranking reversal.

A ranks higher in dollars, while B ranks higher per unit of planned loss. The reversal shows that raw profit can reflect deployment scale. It does not prove B has a durable edge, because the table omits distributions, sequence, dependence, holding time, and how the two strategies were selected.

All numbers are teaching assumptions, not live, backtested, or forecast performance. If budgets, trade counts, and cost treatment were identical, dollar P&L and total R would keep the same ranking. That counterexample clarifies when normalization changes information and when it only rescales it.

A third comparison could hold total R fixed while changing trade count, revealing why average R moves. Such a sensitivity is useful only if eligibility and costs remain aligned. It should not be used to manufacture a preferred ranking by selecting whichever denominator makes one strategy appear strongest.

Report total and average R as different summaries

Total R accumulates normalized results across the sample. Average R divides by trade count. A strategy can show more total R simply because it took more trades, while average R can hide capacity and time. Both values need the observation period and definition of one eligible trade.

Trade frequency can change cost, overlap, and capital use. Comparing 0.16 R per trade from many simultaneous positions with 0.12 from sparse positions does not address portfolio exposure or return per unit of time. Additional metrics can be added only with aligned definitions.

An average should not be treated as the next expected outcome or guaranteed edge. A few large winners can dominate it, and finite-sample variation can reverse ranking. The report should expose distribution and contribution concentration rather than letting one decimal summarize the evidence.

Median R, quantiles, and largest contributors can add context, but they must come from the same eligible trade set. Adding measures after inspecting which one favors a strategy creates another selection problem. The reporting set and hierarchy should be specified before the final ranking is known.

Preserve sequence for maximum drawdown in R

Maximum R drawdown is the largest decline from a prior cumulative-R peak to a later trough. It cannot be reconstructed from total R after trade order is discarded. Two strategies with equal average and total R can impose very different interim loss paths. Sequence is therefore a required source field.

The ledger should retain decision and realization order, including overlapping trades under a declared ordering convention. Cash flows do not enter R directly but can affect budgets and deployment. Shuffling trades for presentation invalidates the historical drawdown even when every individual R remains unchanged.

Maximum drawdown is itself a sample statistic, not a future limit. An unobserved losing sequence can be worse. It should be shown beside average R, losing runs, distribution, and count, without becoming an automatic multiplier or promise of account resilience.

Overlapping trades require a declared ordering and aggregation rule. Closing-time order can differ from decision-time exposure, and both can answer useful questions. The report should not switch silently between them, because the cumulative path and its peak-to-trough decline can change.

Inspect dependence on a few large outcomes

A positive average can arise from many modest results or a few extreme winners. Contribution analysis can show what share of total R comes from the largest trades. Removing extremes without a predeclared data-error rule is not neutral; it can erase genuine tail behavior or selectively improve the chosen narrative.

Losses can cluster across instruments or strategies that share a factor. Per-trade R normalization does not remove correlation. Portfolio review should retain timestamps and overlapping exposures so several one-R budgets are not mistaken for independent risks when they can fail together.

Holding period matters as well. Equal R over different durations does not imply equal annualized or capacity-adjusted performance. Any time normalization introduces more assumptions and should be labeled separately rather than folded into the basic per-trade R definition.

Capacity can also make historical deployment nonlinear. A strategy may achieve attractive small-size R but incur different fills at larger notional. The normalized record does not remove that execution effect, so using its ranking to scale capital requires separate venue-specific evidence.

Control multiple testing and winner selection

Testing many strategies, parameters, or filters and reporting only the best increases the chance of selecting noise. The number of trials, discarded candidates, and selection rule belong in the evidence. A high average R from the winner of a large search is not equivalent to one prespecified strategy.

Development and validation data should be separated, with the method frozen before later evaluation. Reoptimizing after each disappointing period consumes the independence of the holdout. The final report should label whether results are hypothetical, backtested, paper, or live and state all relevant limitations.

Deflated performance methods can address some selection and non-normality issues, but they require inputs and assumptions beyond this simple R table. Mentioning an advanced metric does not repair missing trade records or an outcome-defined denominator. The basic ledger must be valid first.

A trial registry can preserve every tested strategy and parameter family before performance is summarized. That record lets later methods estimate selection burden and prevents discarded candidates from vanishing. It also separates research exploration from the smaller set that entered independent validation.

Keep Sharpe and R comparisons on compatible periods

If a Sharpe ratio is added, return frequency, observation period, risk-free rate, and treatment of autocorrelation and non-normal returns need alignment. Comparing a daily strategy’s Sharpe with a per-trade series from another process can create a ranking from inconsistent time units.

R normalizes monetary loss budgets, while Sharpe relates average excess return to return variability under its definition. Neither subsumes the other. A strategy can rank differently because the measures answer different questions, and the report should explain that difference rather than select the preferred score.

Missing or sparse periods, overlapping positions, and smoothed marks can bias variability. The source data and aggregation rule should be preserved. A polished risk-adjusted number is not reliable when the underlying return series cannot be reproduced from transaction and valuation records.

Attack normalization with outcome-defined denominators

An adversarial test replaces the pre-trade budget with each realized loss and checks that the system rejects the R calculation. Another supplies zero budget for a winner, which would otherwise create infinite R. A third mixes gross Strategy A P&L with net Strategy B P&L and expects a cost-basis failure.

A sequence test shuffles trades and confirms that total and average R remain but maximum drawdown changes or becomes invalid as historical evidence. A duplicate-fill test verifies that parent P&L is not counted twice. These tests target lineage, not whether the resulting ranking looks reasonable.

Boundary cases include no trades, missing currency, negative loss budget, unrecognized cash flow, and budget versions created after entry. The safe output is unavailable for the affected metric. Imputing values solely to preserve a complete table makes the comparison more persuasive and less defensible.

Build the analysis from a trade-level immutable ledger

Each trade needs strategy version, parent ID, decision and entry times, pre-trade budget and currency, stop basis, quantity, fills, exits, costs, conversion, and net P&L. The R row links to that evidence and states whether the denominator was directly recorded or reconstructed.

Aggregate output should show sample period, trade count, monetary P&L, total R, average R, maximum R drawdown, contribution concentration, and missing-record count. Development and validation labels, trial counts, and exclusion reasons prevent the summary from appearing more independent than it is.

Corrections append events. Recomputing a historical R series after a cost or fill correction is appropriate only if both original and revised versions remain traceable. Silent overwriting can change ranking and drawdown without leaving evidence of why the report moved.

Prevent a favorable ranking from becoming an automatic lot multiplier

Strategy B’s 0.16 average R does not specify the next stop, current account capacity, execution cost, or portfolio concentration. Increasing its lot because it outranks A would turn an analytics result into capital authority. A future order still needs its own loss budget and product-specific one-unit loss.

A capital committee can change deployment after a broader review, but the new budget should be prospective, capped, and separate from the historical denominator. Backfilling the larger budget into old trades would reduce their reported R and obscure the evidence that motivated the decision.

Likewise, a lower R ranking does not require an automatic reduction if the sample is too small or noncomparable. It can trigger investigation. Monitoring and authorization should be linked but not collapsed, preserving the ability to challenge both the statistic and the capital choice.

Interpret performance-method sources within scope

Sharpe’s work supports the need for stated return and risk definitions and consistent periods. Bailey and López de Prado support concern about selection bias, multiple testing, and non-normal returns. SEC marketing guidance supports fair presentation of methodology, risks, limits, periods, fees, and hypothetical status.

These sources do not validate the two hypothetical strategies or prescribe R as a universal ranking. They do not supply the USD amounts or trade counts. The reversal is derived arithmetic designed to reveal deployment scale, while the cautions limit how far the comparison can be interpreted.

The evidence-backed conclusion is methodological: preserve pre-trade denominators, net and aligned outcomes, sequence, and selection context. It is not a performance claim or a recommendation to allocate capital to the higher average-R row. The hypothetical ranking remains conditional.

State the comparison without claiming durable edge

Under the hypothetical records, Strategy A earns more dollars but produces 12 total R and 0.12 average R, while B produces 16 total R and 0.16 average R. The different ranking shows that dollars alone cannot separate result from the monetary scale placed at risk.

The table omits sequence, maximum R drawdown, contribution concentration, dependence, holding time, and selection history. Those limitations prevent a claim that B is superior or that either result will persist. If budgets and counts were identical, normalization would preserve the dollar ranking.

The appropriate decision consequence is better analysis, not an automatic larger lot. R can make deployment differences visible when its denominator was saved before entry. It cannot replace current sizing, portfolio capacity, or independent validation of the strategy record.

Decision and control rules

  1. Use the loss budget saved before entry as the R denominator.
  2. Compare net P&L in one currency.
  3. Report trade count and observation period.
  4. Show average R and maximum R drawdown separately.
  5. Do not turn the ranking into an automatic multiplier for the next lot.

Common failure modes

  • Defining one R from the realized loss after the event.
  • Mixing gross and net results.
  • Reporting only the best of many backtests.
  • Comparing Sharpe ratios built from different periods or frequencies.

Evidence and specifications

  1. William F. Sharpe — The Sharpe Ratio

    What this source supports: Performance comparisons require a stated return and risk definition; Sharpe also explains limitations and the need for consistent measurement periods.

  2. Bailey and López de Prado — The Deflated Sharpe Ratio

    What this source supports: Selection bias, multiple testing and non-normal returns can inflate reported risk-adjusted performance.

  3. SEC — Investment Adviser Marketing Compliance Guide

    What this source supports: Performance information can mislead when methodology, risks, limitations, time periods, fees or hypothetical status are not presented fairly and consistently.

Questions to resolve

Does the highest total R identify the best strategy?

Not by itself. Inspect sample size, sequence, drawdown, costs and selection bias.

What if pre-trade budgets were not saved?

An exact R comparison is unavailable. Do not construct the denominator with hindsight; begin recording it prospectively.

Can accounts of different size be compared?

They become more comparable when each trade’s net P&L and pre-trade budget share one currency and definition.

Should the lot be increased for the higher-R strategy?

This analysis does not answer that. Derive the next quantity separately from current equity, stop distance and loss budget.

Can this method be used for a backtest?

Yes, if it is labeled hypothetical and reports costs, trial count, development-validation separation and data limitations.

Recalculate from current inputs

Store pre-trade budgets and net trade results, then compare the dollar and R rankings side by side.

Important: This is educational analysis of historical or hypothetical records. It does not establish strategy superiority, forecast performance or recommend a quantity.