Cutting a historically weak setup can improve a recomputed equity curve, but the useful question is not “what happens if I delete my worst trades?” It is “what happens if I apply one rule that I could have known before entry?”
A valid setup-removal analysis needs a saved setup definition, one stable version, an exact account and date window, complete membership, recorded costs, and every excluded or unavailable trade still visible. The result describes the selected history. It does not prove that skipped trades would leave capital, timing, behavior, and the remaining executions unchanged in live trading.
This guide keeps the original 240-trade example and its strongest 159% and 212% P&L changes, but rebuilds the ledger so trade counts, wins, gross gains, gross losses, profit factor, and net P&L all reconcile.
Disclosure: Trader’s Second Brain is our product. TSB can preserve a setup version, split matching/control/unavailable records, recompute observed results, audit composition, and open exact evidence. It cannot turn a hindsight winner into a causal forecast.
Decision rule: use historical deletion as a hypothesis generator. Pause one identifiable setup only when its membership can be reproduced, its result is not driven by missing data or one outlier, the comparison survives a later or held-out window, and the operational cost of skipping it is acceptable.
The What-If Question Every Trader Should Ask
Aggregate P&L can hide offsetting processes. One setup may contribute positive net results while another consumes them. A setup-level replay makes that decomposition visible, but only if “setup” refers to a rule-defined category rather than a story invented after the loss.
There are three different questions that are often collapsed into one:
- Decomposition: what did each stored setup version contribute inside this exact history?
- Historical deletion: what do the recorded metrics look like when one reproducible setup group is excluded?
- Prospective decision: should the trader pause that setup for a later, predeclared test?
The first two can be calculated from history. The third needs judgment and new evidence. The strategy report-card framework is the better companion when the decision is whether a setup version has enough rule, edge, and risk evidence to remain active.
Freeze the Counterfactual Contract First
Write the contract before opening the “without this setup” result:
- Scope: account, strategy, date from/to, timezone, display currency or R basis, and evidence cutoff.
- Unit: one logical closed trade; fills are not silently counted as independent trades.
- Subject: one named setup and, where available, one immutable setup version.
- Membership: exact stored IDs for subject, control, and unavailable records.
- Result basis: net P&L after represented costs, or one declared R definition; missing results are not zero.
- Metrics: trade count, outcome counts, gross gains, gross losses, net result, profit factor, expectancy, and chronology-dependent drawdown.
- Decision checkpoint: the later window and observation that would keep, revise, or reverse the pause.
If a setup tag changed meaning halfway through the period, split the versions. If 19 trades lack setup evidence, publish that unavailable count. If several currencies cannot be converted at the cutoff, keep monetary comparison unavailable instead of combining incompatible numbers.
A Reconciled 240-Trade Worked Example
Constructed example—not a customer or TSB aggregate result. Scope: one account, one six-month window, 240 logical closed trades, one display currency, represented costs included, and complete setup membership. The four rows are mutually exclusive.
| Setup ledger | Reconciled evidence |
|---|---|
| Trend continuation | 72 trades · 42 wins · $7,420 gross gains · $2,000 gross losses · +$5,420 net · PF 3.71 |
| Pullback entry | 64 trades · 34 wins · $4,380 gross gains · $2,500 gross losses · +$1,880 net · PF 1.75 |
| Counter-trend reversal | 58 trades · 24 wins · $2,260 gross gains · $3,500 gross losses · -$1,240 net · PF 0.65 |
| FOMO / impulse tag | 46 trades · 15 wins · $1,280 gross gains · $5,000 gross losses · -$3,720 net · PF 0.26 |
The counts close: 72 + 64 + 58 + 46 = 240 trades; 42 + 34 + 24 + 15 = 115 wins, so the displayed portfolio win rate is 115 ÷ 240 = 47.9%. Gross gains close at $15,340 and gross losses at $13,000. Net P&L is therefore $15,340 − $13,000 = +$2,340, and portfolio profit factor is $15,340 ÷ $13,000 = 1.18.
That reconciliation matters. A setup table is not decorative: every row must roll into the same starting population and the same gross/net definitions. If the sum of setup P&L does not equal account P&L, stop and resolve duplicates, missing membership, currency conversion, costs, or open positions before interpreting the “worst” row.
Remove One Setup: The Exact 159% Result
Exclude the 46 trades carrying the pre-existing FOMO / impulse tag. Do not choose individual losers inside that tag.
| Metric | All setups → without FOMO |
|---|---|
| Trades | 240 → 194 · 46 excluded |
| Wins / win rate | 115 / 240 = 47.9% → 100 / 194 = 51.5% |
| Gross gains | $15,340 → $14,060 |
| Gross losses | $13,000 → $8,000 |
| Profit factor | 1.18 → 1.76 |
| Net P&L | $2,340 → $6,060 · +$3,720 · +159.0% |
The headline calculation is ($6,060 − $2,340) ÷ $2,340 = 158.97%, rounded to 159%. The result is large because the removed setup has $5,000 of gross losses but only $1,280 of gross gains in this worked ledger.
What the result establishes is narrow: removing those already recorded rows changes the selected history from +$2,340 to +$6,060. It does not show that every future FOMO-labeled opportunity loses, that the trader could identify the label reliably before entry, or that the remaining 194 executions would have been identical.
Remove Two Setups: The Exact 212% Result
Now exclude both negative rows: 46 FOMO / impulse trades and 58 counter-trend reversals.
| Metric | All setups → two retained setups |
|---|---|
| Trades | 240 → 136 · 104 excluded |
| Wins / win rate | 115 / 240 = 47.9% → 76 / 136 = 55.9% |
| Gross gains | $15,340 → $11,800 |
| Gross losses | $13,000 → $4,500 |
| Profit factor | 1.18 → 2.62 |
| Net P&L | $2,340 → $7,300 · +$4,960 · +212.0% |
The second headline calculation is ($7,300 − $2,340) ÷ $2,340 = 211.97%, rounded to 212%. The retained trade count is 72 + 64 = 136, and retained net P&L is $5,420 + $1,880 = $7,300.
This is a more aggressive hypothesis. It removes 43.3% of the baseline trades and may change activity, opportunity overlap, risk utilization, and behavior. The larger historical number is not automatically the better live decision. Test one change at a time when you need to learn which rule mattered.
Why Max Drawdown Needs the Trade Sequence
Net P&L and gross totals can be recomputed from setup aggregates. Maximum drawdown cannot. It depends on the chronological equity path: which wins and losses occurred first, whether positions overlapped, the capital basis, deposits or withdrawals, and how costs were recorded.
The old article assigned each setup a “max drawdown contribution” and subtracted those values from the account drawdown. That operation is not generally valid. A loss inside one setup can begin a portfolio peak-to-trough interval that includes trades from several setups; removing one row can change both the peak and the trough.
The correct replay orders the retained trades by the declared timestamp, rebuilds cumulative net P&L, and recalculates drawdown from that retained sequence. If chronology, overlap, or money basis is incomplete, report drawdown unavailable. Do not substitute a guessed 30–60% improvement merely because net loss declined.
“Worst Setup” Is Not the Same as “Worst Trades”
Removing the five largest losses will always improve the history by exactly their sum. It is arithmetic, not a decision rule, because “will become one of my five worst losses” is unknowable at entry.
A setup-level test is more useful only when membership can be known without the outcome: a saved pattern definition, allowed context, invalidation, risk rule, and version. “Counter-trend reversal v3” can be prospective. “The counter-trend trades that lost” cannot.
The same distinction applies to grade. Removing every low-grade loser while keeping low-grade winners is hindsight selection. A legitimate quality test applies the pre-trade evidence rule to wins, losses, and breakevens alike. The trade-quality scoring guide shows how to keep process grade separate from P&L.
The Data-Snooping Trap
If you test every setup, session, weekday, instrument, direction, month, grade, and combination, the best-looking deletion is selected from many chances to look good. Halbert White’s Reality Check for Data Snooping formalizes the broader problem: reusing the same history for model selection and inference can make an apparently satisfactory result a product of chance.
For a journal decision, the practical controls are simpler than pretending one magic sample-size threshold solves the problem:
- record every filter tested, not only the winner;
- prefer a setup identity that existed before the result;
- inspect concentration by day, instrument, session, direction, and outlier;
- separate discovery from a later or held-out comparison;
- state what future observation would reverse the pause.
There is no universal “30 trades means valid” rule. Thirty dependent trades from two days or one unusual market event can be weaker than a smaller clean exploratory set; hundreds of selectively tagged trades can still be unusable.
How TSB Runs a Setup Counterfactual
TSB has processed 600K+ imported trades cumulatively. That is meaningful product scale, but it is never the denominator for a 240-trade example or one trader’s setup decision. Each result still needs its own account, period, filters, eligible n, exclusions, unit, money basis, and cutoff.
Retrospective Backtester requires an explicit hypothesis and can bind the subject to a saved setup version. It separates matching trades, other trades, and records whose setup/session/instrument evidence is unavailable. Missing results are not converted to zero. The output retains exact evidence references, inspect links, period context, coverage, recorded costs, and composition by instrument, session, weekday, and direction.
The server can therefore answer “what did these stored rows contribute?” without silently rewriting membership. It cannot answer “would everything else have happened identically if I had skipped them?” The equity curve is a descriptive replay of retained recorded P&L, not a market or behavior simulator.
Test one saved setup without deleting history
Freeze the account, dates, setup version, money basis, and cutoff; then inspect subject, control, unavailable records, composition, exact trades, and the rebuilt sequence.
Open Retrospective Backtester →Pause, Investigate, or Keep the Setup?
| Decision | Evidence pattern |
|---|---|
| Pause prospectively | Membership is reproducible; net result is negative after represented costs; the result is not controlled by one day/outlier; a later test and reversal rule are declared. |
| Investigate execution | Setup identity is stable, but plan adherence, size, stop movement, or missing review evidence differs between the negative rows and compliant executions. |
| Split the version | Entry, invalidation, target, market, or risk rules changed inside the window; the combined label mixes materially different processes. |
| Keep collecting | Membership, P&L, cost, currency, or chronology coverage is incomplete; concentration is extreme; later evidence does not yet exist. |
| Keep active | Later comparable evidence does not reproduce the weakness, or the setup provides a documented portfolio role that the single net-P&L comparison does not capture. |
Before blaming the setup itself, open the losing records. The trade-review workflow separates a rule-defined setup from execution drift, missing evidence, and post-trade narrative.
Turn the Hindsight Result Into a Forward Test
- Freeze the original baseline and its complete subject/control/unavailable membership.
- Choose one setup version to pause; do not remove several categories at once.
- Record every skipped eligible signal so opportunity is not invisible.
- Keep trading and reviewing the remaining setups under their existing rules.
- At the predeclared checkpoint, compare executed results, skipped-signal outcomes where measurable, rule adherence, opportunity count, concentration, and risk use.
- Keep, revise, or reverse the pause without rewriting the baseline.
This is the same discovery-versus-validation separation used in the before/after filtering workflow. A strong in-sample deletion earns a test; it does not earn certainty.
What the 240-Trade Example Does Not Establish
- It does not estimate how common weak setups are across retail traders or TSB accounts.
- It does not show that FOMO tags, counter-trend trades, or any fixed setup class are universally unprofitable.
- It does not prove the remaining trades would have identical size, timing, cost, psychology, or opportunity after a live pause.
- It does not supply a chronology-derived maximum drawdown; that needs the ordered trade path.
- It does not make 159% or 212% a forward target. Those are exact changes inside the constructed ledger.
If the suspected issue is session rather than setup identity, use the session-performance audit and keep the same membership, missingness, composition, and later-window discipline.
Methodology Note
- Preserved worked hooks: 240 trades, four setups, +$2,340 baseline, +$6,060 without one setup, +$7,300 with two retained setups, 159% and 212% changes.
- Corrected arithmetic: setup counts, wins, gross gains, gross losses, net result, win rate, and profit factor now roll into the same portfolio totals.
- Withdrawn claim: no 30–150% population range, Pareto distribution, universal setup threshold, or automatic drawdown reduction is asserted without a defined dataset.
- Drawdown boundary: filtered max drawdown requires exact chronological replay and cannot be subtracted from per-setup aggregate rows.
- Source: White (2000) supports the data-snooping warning; TSB capabilities are verified from the current server code; case values are constructed and labeled.
The evidence and correction policy behind this rebuild follows our published editorial methodology.
Final Verdict: Subtraction Earns a Test, Not a Promise
The worked result is intentionally strong: one reproducible setup deletion changes +$2,340 to +$6,060, or +159%; retaining only the two positive setup rows changes it to +$7,300, or +212%. The calculations now reconcile all the way back to 240 trades.
The decision remains disciplined. Decompose the exact history, distinguish setups from outcomes, keep unavailable evidence visible, rebuild drawdown from chronology, audit every filter tried, and test one reversible pause on later evidence. That is how “do less” becomes a measurable setup decision instead of a hindsight slogan.