Durable reference map
A three-stage method to reuse whenever conditions change
Keep the sequence of inputs, calculation and exception testing stable instead of relying on a market forecast.
- Align inputs and unitsGo to equations and definitionsRaw median absolute deviation・Signed robust deviation score
- Reconcile the worked exampleGo to table and calculation stepsA hypothetical 0.32-lot candidate against seven matched records
- Test exceptions and next checksGo to rules and counterexampleUse only matched history in the comparison group.
Build a comparable baseline first
Pooling every lot can misclassify legitimate differences in equity, stop distance, contract specification and strategy. Match at least instrument, account currency, capital band, stop band, side and strategy version.
When a group is too small, avoid a statistical classification and use manual review or an absolute limit.
- Store input timestamp and formula version.
- Retain canceled and corrected orders as separate states.
- Separate records across lot-step changes.
Use median and MAD as an investigation score
The median resists a small number of extreme values. MAD is the median distance of observations from that median. A candidate can then be expressed in raw MAD units from the center.
This is not a probability under a normal distribution or proof of an erroneous order. When repeated values make MAD zero, do not output infinity; select another scale or manual control.
Apply absolute limits before statistical alerts
A loss budget, maximum quantity, maximum notional and permitted-product list can block an order irrespective of history. A statistically ordinary order must not pass when it breaches the loss limit.
Conversely, a planned equity or stop-distance change can justify a larger lot when the recalculation and reason are retained.
- Loss-limit breach: stop.
- Statistical alert: review.
- Document an approved change.
A monitoring design, not claimed product functionality
This is an educational control design. It does not state that the SG Group calculator stores history, computes MAD or rejects orders. An implementation would require separate testing of retention, permissions, alerts, false-positive review and failure behavior.
- Do not connect the score directly to automated execution.
- Record who changed a threshold and when.
- Review false positives and missed cases periodically.
Calculation framework
Raw median absolute deviation
Read the role of each equation first, then follow the numerical example to check the decision path.
Raw median absolute deviation
MAD_raw = median(|Q_i − median(Q)|)- Q_i: historical lots in the comparable group
- MAD_raw: robust dispersion in lot units
In plain language: It describes a typical distance around the center while resisting isolated extremes.
When this conclusion does not apply: Define the median and raw MAD only for at least one comparable non-negative quantity Q_i, with trade direction controlled separately. Do not use the division score when raw MAD is zero.
Signed robust deviation score
S_raw = (Q_candidate − median(Q)) / MAD_raw- Q_candidate: planned lot before entry
- S_raw: dimensionless number of raw MADs from the median
In plain language: A large positive value indicates a larger-than-usual lot; it is not an error probability.
When this conclusion does not apply: Calculate only for a non-empty comparable group, Q_candidate ≥ 0 and MAD_raw > 0, with trade direction controlled separately. Do not adopt a universal threshold such as 3, 5 or 6 without calibration to the group and review costs.
Define the comparable history before calling a lot unusual
A large displayed quantity can be legitimate when equity, stop distance, contract value, strategy, or market condition changes. A monitoring baseline should therefore match instrument, account currency, capital band, stop-distance band, side, strategy version, and market regime before comparing lots.
Pooling unrelated orders can make ordinary structural differences look anomalous. A smaller-stop strategy will naturally use more quantity under the same budget. The alert should not punish correct recalculation merely because another strategy’s lots are smaller.
When the matched group is too small, no stable statistical classification exists. The system can use manual review and hard monetary or notional limits while stating the data gap. Borrowing a universal threshold from an unrelated population creates precision without comparability.
Matching fields should be chosen for their causal relevance to quantity, not because they happened to separate past alerts cleanly. An overly narrow cohort produces constant unavailability, while an overly broad one loses meaning. Coverage and comparability should be reported together so the design tradeoff remains visible.
Use the median as a robust center, not a definition of correctness
The median resists a small number of extreme observations better than the mean. In the invented seven-record baseline, the ordered quantities center at 0.20 lot. That value describes the middle of the matched history; it does not prove that 0.20 is the right lot for the candidate setup.
The candidate still requires an independent stop-loss calculation and hard controls. A value at the median can breach the current budget after a contract or stop change, while a value above the median can be justified by higher equity or a shorter valid distance. Statistical typicality is not authorization.
The baseline should retain every eligible record and correction. Removing a large valid observation merely because it affects the center is selection. Data errors can be excluded only through a declared rule and evidence, with the original and corrected states traceable.
Median calculation should use full-precision executable quantities on one basis, not rounded chart labels from different products. Ties and even sample sizes need a documented convention. Reproduction from the ordered source list is more reliable than accepting a stored center whose membership is unknown.
Calculate raw MAD with its zero boundary visible
Median absolute deviation is the median of absolute distances from the sample median. For [0.16, 0.18, 0.19, 0.20, 0.21, 0.22, 0.24], the deviations are [0.04, 0.02, 0.01, 0, 0.01, 0.02, 0.04], whose median is 0.02 lot. The sorted source values make each step reproducible.
This article uses raw MAD without a normal-consistency scaling factor. The label matters because scaled variants produce different scores. A raw score should not be called a z-score or converted to a normal probability unless the transformation and distributional assumptions are separately justified.
When repeated values make MAD zero, division is undefined. The system should not output infinity and declare every difference an error. It can choose a prospectively defined alternative scale, broaden the matched group if defensible, or route the candidate to manual review.
Any alternative scale must receive its own name and version. Substituting standard deviation only when MAD is zero changes sensitivity and can make scores incomparable across cases. A manual-review state is often more honest than a seamless numeric fallback whose distribution has never been calibrated.
Interpret the signed deviation as a review signal
The score subtracts the 0.20 median from candidate quantity and divides by raw MAD 0.02. A 0.32-lot candidate therefore scores six, while 0.23 scores 1.5 and 0.20 scores zero. Positive values mean larger than the matched center; negative values mean smaller. Direction remains a separate field.
Six raw MADs is not an error probability, a tail probability, or proof of an erroneous order. The example routes the candidate for explanation. A review can then inspect budget, stop, contract, equity, and input provenance rather than automatically block the lot on a statistical label.
The candidate quantities and actions are hypothetical. The row labeled “Confirm reason” does not establish a universal threshold. Alert calibration depends on false-positive costs, missed-error costs, sample quality, and the surrounding hard controls, which differ across systems.
A signed score also permits review of unusually small quantities, which can signal a unit or stop error even when monetary risk is not excessive. The action policy may treat positive and negative tails differently, but that asymmetry must be declared rather than hidden inside an absolute score.
Apply hard monetary and product limits before the alert
A pre-set account-currency loss limit, maximum quantity, notional ceiling, and permitted-product list can block an order regardless of history. A statistically routine lot must not pass if it breaches a current hard boundary. The alert is an additional investigation layer, not the primary loss control.
Conversely, an unusual but correctly recalculated lot can pass hard controls and proceed after its reason is confirmed. The override record should retain the changed equity, stop, specification, or approved strategy state. This keeps the review from becoming a hidden discretionary multiplier.
Hard thresholds are contextual and administratively set. Institutional market-access examples support the control principle but do not provide retail values. The system should not import an exchange or broker threshold and present it as universal safety. A local limit needs its own owner, rationale, effective time, and review history.
Control side and sign without pooling directions blindly
Quantity itself is nonnegative, while trade direction is a separate field. Long and short lots can share a baseline only if the strategy and loss mechanics make them comparable. Pooling sides without review can hide systematic differences in stop placement, borrow, spread, or execution.
The signed robust score describes whether candidate quantity is above or below the center; it does not encode market side. A negative score means smaller quantity, not a short position. Keeping these signs separate avoids an implementation that confuses direction with anomaly magnitude.
A side change may justify a distinct cohort even when the symbol is identical. The record should retain the matching rule and sample counts for each group. If one side has too little history, manual review is more defensible than silently borrowing the other side’s distribution.
Side-specific history can still share hard monetary controls because those controls use current loss arithmetic. Keeping statistical cohorts and absolute limits separate lets the system remain protective when one distribution is unavailable. It also prevents borrowed history from becoming a hidden substitute for direct calculation.
Version the baseline when strategy or specification changes
A new strategy version, provider contract, capital regime, or volume step can shift the normal quantity distribution. Historical lots should remain tied to their original context. A baseline update creates a new version with cohort filters, included IDs, period, median, MAD, and approval time.
Transporting the old median after a contract multiplier changes can label correctly smaller lots anomalous or allow dangerously large ones. The contract-value and stop-loss calculation should detect the economic change first, while the monitoring layer decides whether enough new matched history exists.
A rolling window can adapt, but its length and update schedule should be prospective. Dropping older values after seeing an alert can move the center toward the candidate and erase the signal. Earlier alert decisions remain linked to the version available at their review time.
A regime transition can justify parallel old and new baselines during evaluation. The report should show which one governed the decision and how coverage changed. Selecting whichever baseline produces no alert after seeing the candidate would turn versioning into a discretionary bypass.
Calibrate alerts from review outcomes rather than folklore
A threshold such as three, five, or six has no universal meaning in raw MAD units. Calibration should examine false positives, confirmed input errors, review burden, and missed events in the matched application. Even then, a threshold is an operational choice, not a probability law.
Reviewer decisions need structured reasons: data error, unauthorized budget, correct stop change, equity update, strategy transition, or unresolved. These outcomes can inform later calibration. A simple approved or rejected flag loses the causal information needed to improve the control.
Threshold changes should be versioned and evaluated on later cases. Selecting a cutoff that perfectly separates known historical errors overfits the reviewed sample. The report should state uncertainty and retain a manual path when data are sparse or costs change.
Review capacity belongs in calibration because an alert that cannot be examined promptly may not control anything. Queue time, unresolved rate, and override frequency should be measured beside detection. Tightening a score threshold without operational capacity can create a backlog rather than safer order handling.
Attack the monitor with zero MAD and mixed cohorts
A zero-MAD test supplies identical historical quantities and a different candidate. The system must route to an alternative scale or manual review, not divide by zero. A mixed-cohort test combines different instruments and should fail matching before any median is calculated.
A hard-limit test places a candidate at the median while its current stop loss exceeds budget; the order must still be blocked. An explanation test places 0.32 lot above the center with a verified equity increase and shorter stop, showing that an alert can be resolved without declaring an error.
Boundary cases include an empty group, negative quantity, duplicate records, missing strategy version, stale contract specification, and candidate on a different quantity basis. Clipping or absolute value should not manufacture comparability. The safe state is insufficient evidence with the hard controls still active.
Distinguish monitoring design from implemented product behavior
This article describes an educational control architecture. It does not state that the SG Group calculator stores user history, computes MAD, rejects orders, or sends alerts. Claiming those capabilities without an implemented and tested system would misrepresent both privacy and functionality.
An implementation would require retention rules, permissions, user notice, access control, data minimization, failure behavior, alert delivery, review workflow, and audit logs. Statistical correctness alone does not establish secure or compliant operation. Production claims need independent functional and security evidence.
The page can still provide formulas and a hypothetical example as a design study. It should label the distinction near any workflow language so readers do not mistake a proposed control for a currently available calculator feature. That limitation applies to alerts, storage, rejection, and reviewer routing alike.
A future implementation would need failure-mode tests showing what happens when history is unavailable, storage is delayed, or the review service cannot respond. Default acceptance and default rejection have different costs. The chosen behavior must be explicit rather than inferred from the statistical formula.
Build a review record that preserves both hard and statistical checks
The candidate row needs product, side, strategy version, capital and stop bands, current budget, contract and conversion versions, raw and executable lot, hard-limit results, baseline version, median, raw MAD, score, alert threshold version, and timestamps. Each value needs traceable provenance.
The baseline version lists included unique record IDs, exclusion reasons, period, matching fields, and sample count. Review events retain actor or process, decision, evidence, override reason, and downstream order state. Corrections append rather than rewrite the alert that occurred.
No-score states should be explicit: insufficient sample, zero MAD, unmatched context, stale data, or invalid quantity. They are not equivalent to a zero score. A zero score means the candidate equals the median; unavailable means the comparison itself cannot support an inference.
Exports should retain the matched-record count and identifiers or a reproducible baseline reference. A score without its cohort can be copied into another report and mistaken for a universal property of the order. Provenance keeps the review signal tied to the comparison that gave it meaning.
Evaluate the monitor without using its own decisions as truth
Reviewer confirmation can contain errors, so alert performance should not treat every override label as ground truth without quality checks. Independent sampling, dual review for severe cases, and later reconciliation with actual input corrections can strengthen evaluation while preserving uncertainty.
A low alert rate is not automatically success; the system may be insensitive or the population quiet. A high rate can reflect poor cohort design rather than many bad orders. Metrics should include coverage, unavailable states, review time, reason distribution, and hard-limit interceptions.
Outcome profitability is not the validation target. An erroneous oversized order can profit, while a correct unusual lot can lose. The control evaluates input consistency and authority, not market direction. Using P&L as the primary label would reward lucky errors and punish valid adverse outcomes.
A better audit sample independently rechecks source fields, calculations, approvals, and event timing. Later P&L can be analyzed for other purposes but should not determine whether an input was authorized. This separation preserves the control’s objective even when market outcomes are noisy.
Interpret robust-statistics and market-control sources narrowly
Rousseeuw and Croux support the robustness of median-based scale measures under contamination. SEC and CME institutional materials illustrate pre-set size, price, exposure, and maximum-quantity controls with regular review. They support principles, not a universal retail alert score.
The sources do not supply the seven lots, 0.20 median, 0.02 raw MAD, 0.32 candidate, or score-six action. Those values are invented. They do not establish that the proposed monitor is implemented in the calculator or suitable under any specific legal regime. Local evidence remains necessary.
The evidence-backed conclusion is to compare within context, use robust statistics carefully, and retain hard limits. Threshold selection, privacy design, and workflow implementation remain separate decisions requiring local evidence and testing. No cited source supplies those local choices.
State the anomaly result as a request for explanation
In the hypothetical matched group, median is 0.20 lot and raw MAD is 0.02. The 0.32-lot candidate is six raw MADs above the center. That result routes the order for explanation; it does not assign an error probability or prove the quantity is wrong. Hard limits still decide blocking.
The hard monetary loss check remains prior and binding. A verified equity increase or narrower valid stop can justify the larger quantity if the full recalculation and authority are recorded. A statistically ordinary lot still fails when it breaches the account-currency boundary.
This design is educational and not claimed product functionality. Its decision consequence is a reproducible review signal with explicit unavailable states, contextual cohorts, and versioned thresholds. No universal lot or cutoff follows from the example. Production claims require separate implementation evidence.
Decision and control rules
- Use only matched history in the comparison group.
- Apply the hard monetary loss limit before the statistical alert.
- Use manual review when MAD is zero or the sample is too small.
- Do not describe the score as an error probability.
- Record alert overrides and threshold changes.
Common failure modes
- Pooling lots from every instrument and equity level.
- Allowing extremes to dominate mean and standard deviation.
- Claiming one universal z threshold.
- Presenting the monitoring design as an implemented calculator feature.
Evidence and specifications
- Rousseeuw and Croux — Alternatives to the Median Absolute Deviation
What this source supports: Median-based robust scale estimators resist contamination better than mean and standard deviation in an outlier-prone sample.
- SEC — FAQ on Rule 15c3-5 Risk Management Controls
What this source supports: Institutional market-access controls use pre-set size or price parameters to reject potentially erroneous orders and require regular review; this supports the control principle, not a retail threshold.
- CME Globex Credit Controls
What this source supports: CME provides venue-specific pre-execution exposure and maximum-quantity controls, illustrating that hard limits are contextual and administratively set.
Questions to resolve
How many MADs make a lot abnormal?
There is no universal value. Calibrate against the group, sample size, hard limits and review costs.
What if every historical lot is the same and MAD is zero?
Do not divide. Use a change limit, IQR, hard limit or manual review.
Will this calculator block an order?
This article claims no implementation. Confirm monitoring and rejection features in the actual product specification.
Does a large deviation prove an order is wrong?
No. It is a prompt to verify a changed condition, calculation and reason.
Recalculate from current inputs
Place matched conditions, planned lot, median, raw MAD and the hard loss limit in one pre-entry review record.
Important: This is an educational example of robust statistics and pre-trade control. It provides no universal threshold, rejection feature, error determination or investment outcome.