Durable reference map
A three-stage method to reuse whenever conditions change
Keep the sequence of inputs, calculation and exception testing stable instead of relying on a market forecast.
- Align inputs and unitsGo to equations and definitionsCalibration gap by probability bin・Brier score・Base size independent of confidence
- Reconcile the worked exampleGo to table and calculation stepsCalibrate 60% and 80% bins, then inspect how the multiplier changes loss
- Test exceptions and next checksGo to rules and counterexampleRecord probability before the outcome is known.
Turn confidence into a testable record
Save a 60%, 70% or 80% assessment before the decision and append the outcome after the evaluation window closes. Rewriting probabilities or retaining only attractive cases destroys the test.
Keep strategy, instrument, horizon and outcome definition stable within a bin. One success in a small high-confidence sample is not evidence that the bin is calibrated.
- Freeze the probability before the order.
- Define success and the evaluation horizon in advance.
- Keep cancellations and exclusions with reason codes.
What observed frequency and Brier score can establish
Observed frequency asks how many cases assigned 80% actually succeeded. Brier score averages the squared difference between each probability and its zero-or-one outcome, permitting comparison across probability forecasts.
Neither measure reports loss magnitude, fees, slippage or the account’s ability to survive a run of losses.
An unverified multiplier changes the loss budget
With a 50-point stop, USD 10 per point per lot and a USD 300 budget, base size is 0.60 lot. Applying 1.5 solely because a case is labeled high confidence makes size 0.90 lot and planned stop loss USD 450.
The multiplier has converted a forecast assessment into a new account loss allowance. If that allowance is ever changed, it requires a separate capital-risk rule and approval.
Audit two probability bins across 100 hypothetical cases
The data below are invented solely to demonstrate the arithmetic; they are not trading results. The 60% bin contains 31 successes in 50 cases and the 80% bin contains 30 successes in 50 cases.
The high bin realizes 60%, a calibration gap of minus 20 percentage points. It offers no evidence for the assumed high-confidence multiplier.
Keep this decision separate from win rate and reward-to-risk
This article tests whether an ex-ante subjective probability corresponds to later frequency. Article 5 starts with observed win rate and asks what loss magnitude and ordering do to results. Article 6 separates a distant target from the stop-loss budget.
A calibrated forecast can still have damaging tail losses. A lower-win-rate process is not judged from this article alone.
Use confidence without changing quantity
Retain confidence to monitor analytical quality, but calculate each order from the same loss-budget rule. Even if calibration improves out of sample, any budget change remains a portfolio-level decision rather than a confidence input.
- Set a minimum sample rule for every probability bin.
- Retain an out-of-sample validation period.
- Do not let a confidence field automatically increase lots.
Calculation framework
Calibration gap by probability bin
Read the role of each equation first, then follow the numerical example to check the decision path.
Calibration gap by probability bin
Gap_k = (W_k / N_k) − p_k; Gap_pp,k = 100 × Gap_k- Gap_k: observed frequency minus stated probability in bin k, in probability units
- Gap_pp,k: calibration gap in percentage points for the worked table
- W_k: successful cases
- N_k: evaluated cases
- p_k: probability recorded before the outcome
In plain language: A negative Gap_k means realized frequency was below stated probability. The table displays Gap_pp,k after multiplying Gap_k by 100.
When this conclusion does not apply: Compare only when N_k > 0, 0 ≤ W_k ≤ N_k, 0 ≤ p_k ≤ 1, and sampling, outcome definition and evaluation horizon are consistent.
Brier score
BS = (1 / N) × Σ(p_i − y_i)^2- p_i: ex-ante probability for case i
- y_i: 1 for success and 0 for failure
- N: number of evaluated cases
In plain language: It is the mean squared error of a binary probability forecast.
When this conclusion does not apply: Calculate only when N > 0, every p_i is between zero and one, and each y_i is binary. It measures probability forecast quality, not loss magnitude or safe quantity.
Base size independent of confidence
Q_base = B / (D × V)- B: account-currency loss budget
- D: stop distance
- V: value per lot per distance unit
In plain language: The stated forecast probability is not multiplied into the base quantity.
When this conclusion does not apply: Calculate only when B is non-negative and D and V are positive. Any change to B belongs to a separate capital-risk policy.
Turn a confidence word into a forecast that can fail
High conviction is not a measurable input until it is translated into a probability, recorded before the outcome, and tied to a defined event. A statement such as 80 percent needs an evaluation horizon, success criterion, instrument population, and strategy version. Without those boundaries, later reviewers can move cases between bins until the label appears accurate.
The record should preserve every eligible forecast, including cases that were not traded if the analytical process issued a probability. Keeping only positions that were opened can condition the sample on a later capital or execution decision. Removing unattractive outcomes or assigning confidence retrospectively produces a score about curation rather than forecasting quality.
A probability forecast can be useful even when quantity never changes. It can reveal whether analysts distinguish cases, whether an 80-percent group succeeds more often than a 60-percent group, and whether accuracy changes across later validation periods. None of those questions requires multiplying lots by the forecast number.
Build bins that retain their original population
Calibration compares stated probabilities with observed frequencies among cases that shared that statement. The population must hold strategy, outcome definition, and horizon sufficiently stable to make the comparison interpretable. Pooling a one-hour directional call with a ten-day profit criterion creates an average whose denominator has no common event meaning.
Bins trade resolution for sample size. Very narrow bands may contain too few observations, while broad bands can hide variation within them. The bin edges and treatment of boundary values should be declared before results are inspected. Changing a 79-percent case from one group to another after its outcome undermines the audit even if total counts remain unchanged.
Sample count must accompany observed frequency. Four wins in five cases produce 80 percent, but the small denominator does not establish that the underlying forecast is calibrated. A display that emphasizes the percentage and suppresses count encourages false precision. The correct response to thin evidence is wider uncertainty or no conclusion, not a stronger lot multiplier.
Read the calibration gap with its sign and unit intact
For a probability bin, observed frequency is wins divided by observations. The calibration gap is observed frequency minus stated probability. A negative value means outcomes occurred less often than forecast; a positive value means more often. Multiplying the decimal gap by one hundred expresses percentage points, which must not be confused with a percent change relative to the forecast.
In the hypothetical 60-percent bin, 31 successes among 50 cases give 62 percent and a positive two-percentage-point gap. The 80-percent bin has 30 successes among 50, so its observed rate is 60 percent and its gap is negative 20 percentage points. The higher label therefore has the worse discrepancy in this constructed sample.
Calibration alone does not rank economic outcomes. A correctly forecast success can still have a small gain, and a failure can carry a large loss. The gap does not include costs, gaps, payoff asymmetry, or dependence between cases. It tests the probability statement under the chosen outcome definition, not the account’s ability to bear a position.
Use the Brier score for forecast error, not loss capacity
The Brier score averages the squared difference between each probability and its binary outcome. A forecast of 0.80 receives a small squared error when the event occurs and a much larger one when it does not. Because every case is scored, the measure discourages extreme probabilities that are unsupported by their realized frequency.
The example yields 0.236 for the 60-percent bin and 0.280 for the 80-percent bin, with 0.258 across all 100 cases. Those values are reproducible from the stated counts because every case in a bin shares one probability. They are teaching arithmetic, not evidence about an actual forecasting process or a threshold for acceptable quality.
A lower Brier score does not reveal the monetary consequence of errors. Two forecasters can have identical scores while one is applied to events with much larger downside. Before the measure is compared across samples, the event definition and case mix must remain consistent. It should not be converted into an account-currency budget through an invented scaling rule.
Calculate base quantity without importing subjective probability
The base position uses a USD 300 loss budget, a 50-point stop, and USD 10 per point per lot. One lot therefore carries USD 500 of price-distance loss, and USD 300 divided by USD 500 gives 0.60 lot. The stated probability does not appear in that equation because it is not a monetary allowance or a distance unit.
Applying a 1.5 multiplier to the 0.60-lot result raises quantity to 0.90 lot. Multiplying through the unchanged 50-point stop and point value gives USD 450. The multiplier has increased planned loss by USD 150, so it is economically equivalent to changing the budget even if the interface still displays USD 300.
The contradiction is useful because it cannot be repaired by a confident label. Either the USD 300 boundary remains active and 0.90 lot is rejected, or a separate capital policy authorizes USD 450 under defined conditions. The latter would need its own approval and portfolio analysis; it cannot be smuggled into the calculation as forecast metadata.
Keep calibration evidence and capital authority on separate records
A probability service can publish forecast, model or analyst version, timestamp, event definition, and later outcome. A sizing service can consume stop distance, monetary value, and an authorized loss budget. Separating the records makes it possible to test forecast quality without allowing a descriptive field to override the account rule.
If an organization later studies conditional budgets, the experiment should be specified prospectively with caps, sample separation, and portfolio exposure controls. A successful historical calibration chart is not sufficient because it says nothing about loss magnitude or simultaneous positions. Testing many multipliers and reporting the most attractive one also introduces selection bias.
The approved budget should remain visible beside the forecast and planned stop loss. A front end that shows only the final lot can conceal a multiplier breach. Replaying rounded quantity into account currency provides a simple invariant: the price-distance amount cannot exceed the active budget before separately described costs and execution additions.
Distinguish confidence from observed win rate and planned payoff
Confidence is an ex-ante probability statement. Win rate is an ex-post count over a defined sample. Planned reward-to-risk is a ratio of target distance to stop distance. They can be analyzed together later, but none is a substitute for the others, and none directly supplies an executable quantity under a fixed monetary loss budget.
A forecast can be well calibrated while the strategy has unfavorable payoffs, because frequent small successes may coexist with rare large losses. A high observed win rate can also arise in a selected sample that contains no recorded probabilities. A distant target can raise the reward multiple while reducing attainment. Mixing the labels prevents each claim from being tested on its own terms.
The operational safeguard is a typed schema. Probability belongs on a zero-to-one scale; outcome is binary under a stated horizon; target and stop are price distances; budget is account currency; quantity is lots or contracts. Rejecting unit-incompatible multiplication catches many conceptual errors before a more sophisticated statistical review is needed.
Design an out-of-sample review that cannot be rewritten
A calibration result measured on the same cases used to define the forecasting method is descriptive. Stronger evidence requires freezing the method and evaluating later cases without changing bin definitions after outcomes arrive. The record should mark development and validation periods and disclose any exclusions or missing outcomes.
Market conditions and strategy versions can shift. A deterioration in a later period may mean the forecast mapping changed, or it may reflect sampling variation. The response is to report counts and uncertainty and investigate the population, not to increase quantity on a small favorable run. Version changes should start a new series or maintain an explicit bridge analysis.
Repeated testing creates another hazard. Trying many binning schemes, probability transformations, and multipliers makes some chart look calibrated by chance. Retain the number of trials and the rule used to select the reported design. A probability score chosen after inspection should not be described as independent support for more capital at risk.
Model tail loss outside the probability score
Binary success may ignore how far price moved or how the stop executed. A failed 80-percent forecast can lose more than planned because of a gap, while a successful case can realize less than its target after partial exits and costs. The Brier score treats both outcomes as zero or one and intentionally omits their monetary magnitude.
A separate loss distribution should therefore use realized account-currency or normalized R outcomes, preserving sequence and execution details. Even then, a historical tail does not guarantee the next bound. Portfolio concentration and correlated failures can make several individually modest budgets arrive together, which a per-case confidence score does not address.
This is why good directional calibration does not authorize an automatic lot increase. It may improve one piece of the analytical process, but loss capacity depends on account policy and aggregation. The decision consequence of a confidence error is not limited to being wrong; it is the money exposed when wrong, including execution beyond the planned stop.
Attack the workflow with retrospective and selective labels
An adversarial test can assign 80 percent only after winners are known and verify that the system rejects or flags the late timestamp. Another can delete losing cases and confirm that observation counts no longer reconcile with the issuance log. A third can change the success horizon after a forecast misses, which should create a new outcome definition rather than rewrite the original one.
A sizing test should submit the 0.90-lot multiplier while the active budget remains USD 300. The result must be a breach regardless of the confidence value or calibration score. If the interface accepts it because a high-confidence flag is present, the capital rule has been made subordinate to an unaudited categorical input.
Boundary tests include probabilities below zero or above one, zero observations, wins exceeding observations, missing outcomes, and mixed event definitions. The safe response is no calibration statistic for the affected group. Producing a default zero gap or perfect score from an empty bin would create false evidence exactly where the data are weakest.
Record enough context to reproduce both forecast and size
The forecast ledger needs case ID, issue time, probability, bin rule, strategy and instrument context, horizon, success definition, analyst or model version, and immutable later outcome. The sizing ledger needs risk budget, stop, monetary point value, raw lot, step, executable lot, and checked loss. A reference links them without merging their authority.
Calibration reports should expose eligible case count, missing outcomes, observed frequency, gap in decimal and percentage-point form, Brier score, and evaluation period. They should label whether figures are development, backtest, hypothetical, or later observations. A polished percentage without provenance is not evidence that can support a decision consequence.
Overrides need their own event. If a capital committee changes the loss budget for reasons unrelated to confidence, the record should say so rather than inherit the forecast label. If no override exists, the USD 300 amount remains binding. This lineage prevents later analysts from mistaking unauthorized oversizing for a deliberate experimental policy.
Read the research sources within their proper scope
Brier’s primary work supports scoring probability forecasts against outcomes with a squared error measure. It does not specify an acceptable score for this context or connect a score to lot size. The hypothetical calculations apply the scoring rule only to demonstrate verification and must not be presented as observed SG Group performance.
Gervais and Odean model how biased learning from successes can generate overconfidence. Using that finding to prohibit a direct confidence multiplier is a control inference made by this article, not a literal trading rule in the paper. CME material separately supports deriving position size from stop location and an account risk amount.
Together, the sources justify a separation of claims: probabilities can be tested, confidence can be biased, and quantity can be tied to a monetary stop budget. They do not prove the true success probability of any trade, validate the 60- or 80-percent labels, or guarantee that USD 300 captures every realized loss component.
State what a favorable calibration result would and would not change
Suppose a large later sample showed small calibration gaps and a stable Brier score under a frozen method. That would support the narrow claim that the recorded probabilities corresponded reasonably to observed event frequencies in that population. It would not require any change to the base 0.60-lot calculation.
A separate policy could study whether exposure varies with forecast information, but it would need to model payoff size, tail behavior, dependence, costs, and aggregate account capacity. The comparison should include a fixed-budget baseline and protect the validation sample from the many choices used during development. Until then, forecast quality remains a monitoring input.
If calibration is poor, lowering probability labels may improve honesty but does not automatically dictate a smaller lot either. Quantity still follows the active monetary budget and stop loss per lot. Forecast repair and capital reduction can both be sensible topics, yet their rationales and approvals must remain distinct.
Carry the bounded conclusion into the user-facing decision
For the hypothetical inputs, 0.60 lot is the base quantity and USD 300 is the checked price-distance loss. The 1.5 multiplier produces 0.90 lot and USD 450, breaching that amount by USD 150. The 80-percent bin also realizes only 60 percent in the constructed sample, but the budget breach stands even if calibration had been perfect.
The correct interface consequence is to retain confidence for later evaluation while refusing to let it modify the lot formula implicitly. A budget change requires a separately visible rule and authorization. When those elements are absent, the multiplier is not an advanced use of probability; it is an undocumented increase in money exposed at the stop.
This conclusion is educational, not a recommendation of a probability, budget, stop, or position. Forecasts can fail, samples can shift, and execution can exceed a planned distance. The page can show how to preserve evidence and detect a contradiction. It cannot promise accuracy, profitability, or a realized loss ceiling.
Decision and control rules
- Record probability before the outcome is known.
- Display observed frequency, sample count and Brier score by bin.
- Do not multiply base lots directly by subjective probability.
- Treat a budget change as a separate approval.
- Do not mix strategies or conditions into one calibration population.
Common failure modes
- Assigning high confidence retrospectively only to winners.
- Treating four wins in five cases as a validated 80% bin.
- Using good directional calibration as evidence that tail loss is affordable.
Evidence and specifications
- Brier (1950) — Verification of Forecasts Expressed in Terms of Probability
What this source supports: The primary paper introduces a squared probability score for verifying probabilistic forecasts against outcomes, supporting ex-post testing of stated confidence.
- Gervais and Odean (2001) — Learning to be Overconfident
What this source supports: The finance paper models how biased learning from successes can generate overconfidence. The rule against converting confidence directly into a quantity multiplier is this article’s control inference from that finding.
- CME Group — Proper Position Size
What this source supports: CME states that sizing requires the stop location and the dollar or percentage amount the account is prepared to risk, supporting a base size independent of subjective confidence.
Questions to resolve
Can size rise if confidence is accurate?
Calibration describes forecast probabilities. A larger per-trade loss allowance requires a separate account-risk policy.
How many cases are enough?
There is no universal count. Predefine a minimum by bin and consider dependence, strategy stability and an out-of-sample period.
Does a low Brier score imply profit?
No. Payoff size, costs, execution and trade selection remain separate.
Why record confidence at all?
It allows the analyst or model’s probability statements to be monitored without using them as automatic leverage.
Recalculate from current inputs
Calculate base size from the loss budget and keep confidence as a separate calibration record, not a multiplier.
Important: This is educational material about forecast calibration and sizing. It does not recommend a probability, budget or quantity. Hypothetical outcomes do not indicate future performance, and calibration does not guarantee profit or a loss ceiling.