Trader's Second Brain Trader's Second Brain

The Equity Curve Hack: Compare All Trades vs Best Setup

The equity-curve comparison hack is one of the fastest ways to see whether one declared part of your trading history behaved differently from the rest. Plot the complete recorded curve beside a fixed setup or context slice, then inspect the gap, sample, drawdown, period stability, and exact trades. The visual shows where contribution concentrated; it does not prove that deleting historical trades would preserve the same account path or future result. This guide keeps the punch of the chart while making the comparison reproducible.

Quick Answer

Freeze one account, period, timezone, money basis, and setup version before reading the chart. Plot all eligible trades and the declared matching slice, but use matching-versus-rest metrics for the decision because the all-trades line already contains the slice. Show both samples, costs, drawdown, concentration, missing evidence, and every filter tried. Treat an attractive historical gap as a hypothesis, then confirm or reject it on a later untouched period.

Performance lab · setup-level evidence
Setup, session and drawdown review

Know which setups actually have an edge.

Find which setups, sessions, and behaviors make or lose money.

Find weak setups →
Trader's Second Brain preview

The equity-curve comparison hack is one of the fastest ways to see whether one declared part of your trading history behaved differently from the rest. Plot the full recorded curve beside a fixed setup or context slice, then inspect the gap, sample, drawdown, and exact trades that created it.

The chart is powerful because it compresses a long ledger into a visible path. But it is not a time machine. A better filtered line describes recorded history under a mathematical exclusion; it does not prove that skipping those trades would have produced the same account path, sizing, attention, or future result. This guide shows how to keep the visual punch while removing the hindsight trap.

Quick answer: keep the title’s “best setup” as a candidate, not a verdict. Define the setup before reading the chart, compare it with the other trades from the same account and period, show both sample sizes, and require an untouched later period before changing a live plan.

What the All-Trades vs Setup Overlay Actually Shows

Build cumulative results in chronological order for one frozen account and period:

  • All-trades curve: every eligible recorded trade in the scope.
  • Matching curve: only trades that meet the predeclared setup or context rule.
  • Other-trades result: the eligible records in the same scope that do not meet that rule.

Because the all-trades curve contains the matching trades, the two plotted lines are not independent portfolios. Their end-point difference is the cumulative recorded contribution of the other trades under the chosen calculation basis. That identity is useful, but calling the gap “money you would have saved” adds a causal claim the chart cannot establish.

For example, suppose a declared setup slice ends at +6R, the other trades end at -2.5R, and the complete eligible history ends at +3.5R. The gap is arithmetically consistent with the recorded other-trade contribution. It does not prove a trader could simply delete those trades and retain the same 6R path: removing activity can change capital, margin, sequence, opportunity, and later decisions.

If the curve definitions themselves are unclear, begin with the equity-curve reading framework: define the y-axis, x-axis, cost basis, open-position policy, and external cash flows before interpreting shape.

Build a Valid Comparison Before Looking for the Winner

  1. Freeze one account. Do not combine unrelated account sizes, rule sets, currencies, or execution environments just to increase the sample.
  2. Freeze the period and timezone. Both groups need the same start, end, and trade-order convention.
  3. Declare the candidate. Use a saved setup version or an explicit context such as instrument, side, session, weekday, or minimum recorded R:R. Write it down before viewing the result.
  4. Declare eligibility. State how open positions, breakevens, partial closes, cancelled records, duplicates, fees, funding, and corrections are handled.
  5. Verify labels. Missing setup or session tags are unavailable evidence, not permission to place a trade in whichever group improves the chart.
  6. Use one money basis. Native currencies cannot be added as though their numeric amounts share a unit. Combined monetary curves need compatible historical conversion evidence.
  7. Preserve the complete result. Show matching trades, other trades, and all eligible trades. Never publish only the flattering line.

The setup name alone is not enough. If entry rules, exits, risk, or instruments changed, split the history by saved setup version. Otherwise the chart can average two different systems into one label. The setup-performance breakdown explains how taxonomy quality and version drift can dominate the apparent ranking.

How to Read the Gap Without Fooling Yourself

Observed patternWhat it supportsWhat to inspect next
Matching rises; other trades fallThe declared slice contributed more favorably in this recorded scopeCounts, costs, outliers, month-by-month stability, exact losing buckets
Both riseBoth groups contributed positively in sampleRisk, drawdown, capacity, overlap, and whether exclusion is even justified
Matching wins on one spikeThe headline depends on concentrated contributionResult with the largest trade identified, not silently deleted
Leadership alternates by monthThe difference is unstable across subperiodsRegime, version, exposure, and whether more untouched data is required
Gap begins after a rule changeA change point deserves investigationSetup version, sizing, execution, data coverage, and contemporaneous notes
Money curve unavailableThe monetary basis is incomplete or incompatibleCurrency, fee, P&L, account mapping, and source coverage before interpretation

Read the chart together with trade count, expectancy, profit factor, win rate, maximum drawdown, and a period breakdown. No single metric crowns a permanent “best” setup. A steep line built from a few concentrated outcomes carries different evidence from a steadier line distributed across many comparable trades.

Use Matching vs Rest for the Decision

The overlay is the attention-grabber; the clean comparison is the declared slice against the rest of the same eligible population. Ask: did the matching group differ, how much evidence supports that difference, and where did it occur? Then open the exact trades. This prevents the all-trades line—which already contains the candidate—from being mistaken for a separate control.

Keep drawdown path-dependent

Removing trades can change the historical peak, so filtered and complete maximum drawdown need not differ by the same amount as final P&L. State the starting basis and ordering convention. A lower filtered drawdown is an observed property of that reconstructed series, not a guarantee about live risk after a rule change.

The Real Filter-Discovery Trap

The old version called this survivorship bias. The sharper diagnosis is selection bias under multiple testing: try enough setup, day, session, direction, instrument, grade, and compound filters on the same history, and one can look exceptional by chance.

Research on backtest overfitting shows why the number of alternatives tried matters: selecting the most attractive result from many historical configurations inflates the apparent winner. The primary paper by Bailey, Borwein, López de Prado, and Zhu describes this mechanism in strategy backtests; the same caution applies at smaller scale when a trader searches many journal slices and reports only the best-looking curve. See the paper and abstract.

Protect the analysis with a two-pass workflow:

  1. Exploration: search the old history, but label every interesting filter as a hypothesis and record how many alternatives you inspected.
  2. Confirmation: freeze the rule, setup version, metrics, account, and evaluation window before a new untouched period begins.
  3. Decision: compare the later evidence with the predeclared expectation. Preserve failures as well as successes.

A filter discovered after seeing the curve is not useless; it is simply exploratory. Its value is the next test it creates, not the historical certainty it appears to provide.

There Is No Universal Trade-Count Threshold

Rules such as “30 trades prove the setup” or “50 trades provide high confidence” are too blunt. Required evidence depends on outcome dispersion, dependence between trades, payoff asymmetry, regime coverage, costs, and how many filters were tested. Thirty nearly identical trades from one market week can contain less independent information than a smaller but broader sample—and neither automatically establishes future edge.

Instead of using one magic number, ask:

  • Are both matching and other groups large enough to show more than one outcome path?
  • Does one trade or one day dominate the difference?
  • Does the direction persist across calendar subperiods and reasonable boundary choices?
  • Were costs, risk units, and setup versions comparable?
  • How many candidate filters were searched?
  • What untouched data will confirm or reject the hypothesis?

If these questions cannot be answered, the correct output is “interesting but not verified,” not an invented confidence label. For a broader decision framework, use the trading-edge evidence guide.

Which Filters Are Worth Comparing?

Prefer categories that existed in the trading plan or source record before the result was known:

  • Saved setup version: one precisely defined playbook against other eligible trades.
  • Session: fixed clock/session labels under one timezone.
  • Instrument and side: stable symbols and long/short classification.
  • Weekday: useful as a descriptive slice, especially when session and timezone are fixed.
  • Recorded quality or rule tag: valid only when the label was applied consistently without knowledge of the eventual result.
  • Minimum recorded R:R: based on a stored planning field, not a value reconstructed from the outcome.

Compound filters can answer precise questions, but each added condition shrinks the sample and increases search freedom. “A-grade version 3 setup during London on Tuesday” may be operationally clear; it may also describe only a handful of trades. Show coverage and let the evidence state constrain the conclusion. The quality-tag analysis covers the special risk of grading trades after their outcomes are known.

Turn the Visual Into a Safe Decision

  1. Name the observation. “The declared setup slice contributed more favorably in this recorded account and period.”
  2. Name the limitation. Include sample balance, concentration, missing labels, searched filters, and any incompatible money evidence.
  3. Inspect the records. Open the trades responsible for the widening and narrowing gaps. Check fills, notes, screenshots, rules, and corrections.
  4. Choose a reversible test. Keep the setup rule fixed, reduce exposure, use simulation, or collect a new period rather than declaring a permanent ban from one chart.
  5. Write an invalidation condition. State what later comparable evidence would make you keep, revise, or reject the filter.
  6. Review on evidence cadence. Re-run when a meaningful new sample or setup version exists—not because a daily line moved.

If the question is explicitly “what did the recorded result look like without this bucket?”, use the impact-analysis workflow. Keep its output counterfactual and separate from the official historical ledger.

How TSB Makes the Hack Reproducible

Ownership disclosure: Trader’s Second Brain is our product. The current TSB Backtester turns the chart idea into an exact, server-calculated hypothesis test over canonical Journal evidence. You choose one account, date range, and either a saved Setup version, an older stored label, or explicit context filters. It separates matching trades from other trades in the same scope rather than inventing a clean control.

The result shows the observed cumulative P&L comparison, matching and other sample counts, expectancy, profit factor, win rate, maximum drawdown, period behavior, composition, and recent supporting trades when those fields are available. Optional trade-order risk changes the order of the same recorded outcomes; it does not simulate new market outcomes. Missing results, setup identity, currency evidence, or context stay unavailable instead of becoming zero.

The strongest part is traceability: the result carries the exact account, period, timezone, evidence cutoff, filters, and Setup version, and it links back to the underlying trades. Backtester never edits the Journal. When evidence is available, Coach can explain the bounded result from the same signed context rather than manufacturing a psychological story or future forecast.

TSB has processed 600K+ imported trades across its import history, and its canonical registry recognizes 330 exact broker, exchange, platform, and prop-export profiles. These figures describe platform-wide imported trades and recognized source routes—not users, Backtester samples, Coach-analyzed trades, complete data, or performance outcomes.

Run the Curve Comparison as a Real Hypothesis

Choose the exact account, freeze the period and Setup version, compare matching trades with the rest, then open the evidence behind the gap.

Try the Backtester demo

Methodology and Evidence Limits

This September 10, 2026 fact cycle checked the guide against the current Backtester V2 page, server contract, client renderer, API ownership and entitlement boundaries, integration tests, canonical import registry, and published research on backtest overfitting. It removed fabricated trader-population percentages, invented dollar examples, universal profit-factor and sample thresholds, causal discipline claims, guaranteed improvement language, and fixed review cadences.

No exact external journal, broker, exchange, firm, or program participates in this general analytical decision, so a provider catalog component is not applicable. Article and BreadcrumbList remain; FAQPage stays tied to visible FAQ content. No artificial Review, Rating, Product, or ItemList schema is added.

Final Verdict: Keep the Visual, Upgrade the Proof

The hack deserves its reputation. An all-trades curve beside one declared setup can make concentrated contribution, drag, instability, and change points visible faster than a page of averages. The chart tells you where to look.

The upgrade is simple: freeze the hypothesis, compare the setup with the rest, show the sample and money basis, inspect the exact trades, and validate the rule on untouched evidence. Then the overlay stops being hindsight theater and becomes a disciplined decision instrument.

Disclosure: Trader’s Second Brain is our product. This guide is educational, does not provide individualized investment or financial advice, and does not guarantee that a filter, setup decision, Backtester result, or Coach explanation will improve trading performance.

Igor Manuilov
Written and reviewed by
Igor Manuilov
Founder of Trader's Second Brain · Trader since 2014
Editorial accountability

Trader since 2014. Built Trader's Second Brain to make execution review more evidence-based and less dependent on memory, scattered spreadsheets, or vague journaling.

Performance lab · setup-level evidence
Setup, session and drawdown review

Turn trading statistics into a review plan.

Find which setups, sessions, and behaviors make or lose money.

Find weak setups →
Trader's Second Brain preview

Frequently Asked Questions

Quick answers to the most common questions about Equity Curve Comparison Hack.

It plots the cumulative result of all eligible trades beside a declared subset such as one Setup version, session, instrument, side, weekday, or recorded quality tag. The end-point gap equals the recorded contribution of the nonmatching trades under the chosen calculation basis. It is descriptive, not proof that skipping those trades would have preserved the same capital, sequence, decisions, or future outcome.

Start with a setup defined in your plan or a saved Setup version, not whichever historical line looks steepest. Compare it with the rest of the same account and period, then inspect sample balance, expectancy, profit factor, drawdown, concentration, period stability, costs, and exact trades. There is no universal profit-factor or trade-count threshold that proves a setup is best; confirmation needs a frozen rule and later untouched evidence.

Retroactive labels can support exploration when contemporaneous notes or screenshots make the classification auditable, but they are vulnerable to outcome-aware judgment. Mark them as reconstructed. For confirmation, define the setup and tagging rule first, apply it consistently going forward, preserve unknown labels as missing evidence, and evaluate only when the later sample represents the intended conditions. No fixed number of days makes weak tagging reliable.

There is no universal gap percentage that justifies an immediate rule change. Investigate any decision-relevant gap, then check sample sizes, outlier concentration, drawdown, costs, missing labels, version consistency, subperiod stability, and how many filters were searched. A large selected gap can still be fragile. Use a reversible test and require later untouched evidence before making the filter permanent.

Possibly. The historical comparison cannot know which future opportunities would occur or how removing trades would change attention, capital, sequence, and execution. Treat the proposed cut as a hypothesis: define it precisely, test it with controlled risk or simulation, record the missed and taken opportunities, and set an invalidation rule. Historical negative contribution is evidence to investigate, not certainty about the next trade.

The main problem is better described as selection bias under multiple testing: search enough setups, sessions, days, instruments, grades, and compound filters on the same history, then report only the most attractive line. Record every candidate tried and label discovered filters exploratory. Freeze one rule before an untouched period begins; that later evidence can confirm, weaken, or reject the hypothesis.

Re-run when a meaningful new comparable sample or Setup version exists, not on a universal calendar. A high-frequency strategy may accumulate evidence faster than a low-frequency one, while a regime or rule change may make older trades less comparable. Keep the evaluation window and hypothesis frozen, avoid reacting to every daily move, and record each review so repeated testing remains visible.