A setup performance breakdown answers which documented entry rules produced which observed results—not which setup will make money next. The analysis becomes useful only when setup identity existed before outcome, trade coverage reconciles, costs are complete, and old rule versions are not mixed with new ones.
This guide turns setup tags into a decision-grade audit. You will define stable setup IDs, separate missing labels from losing strategies, build a cost-complete comparison table, read pairwise differences correctly, and decide whether each setup is supported, unresolved, operationally broken, or rejected under a prewritten rule.
Quick answer: group reconciled trades by setup ID and version; report eligible decisions, included trades, unknown labels, gross wins, gross losses, costs, net result, expectancy, profit factor, drawdown convention, and evidence limits. Open the underlying trades before acting. A profitable point estimate is a candidate, not proof of edge; an untagged bucket is missing classification, not a strategy verdict.
Why Setup Tagging Changes Everything
Account-level P&L combines every setup, instrument, session, rule version, and execution mistake. Two strategies with different opportunity sets and loss paths can cancel each other in the total. Setup identity supplies the denominator that lets the journal ask a narrower question: What happened when this specific entry rule was eligible and executed?
The Missing Layer
A broker export usually establishes what was executed. It may not establish the setup definition, whether the rule was eligible, why the trade was skipped, or which plan version governed entry. Those decision fields must come from a saved setup, pre-entry record, or other traceable evidence. Do not reconstruct them from outcome alone.
Why This Can Be Actionable
A clean breakdown can support a bounded decision: keep collecting, restrict a setup under specified conditions, repair execution, separate a changed version, or stop deploying it until a new test passes. It cannot prove personal ability, psychological cause, or permanent future profitability. Use the trade-review workflow to keep the summary attached to the rows that support it.
How to Tag Trades Without Rewriting History
Step 1: Define the Setup Before the Trade
Give each setup a stable ID, human-readable name, version, effective timestamp, required entry conditions, invalidation, allowed instruments and sessions, and evidence fields. There is no universal correct number of setup tags. Use as many as your decision process genuinely distinguishes and your evidence can support.
Do not use one tag for several rule systems merely to enlarge the sample. Do not split a setup into post-outcome micro-categories merely to find a winner. If the rule changes materially, create a new version so prior trades remain attached to the rule that actually governed them.
Step 2: Assign Identity at the Decision Point
The preferred record links the planned setup before entry. If a historical trade lacks that link, classify it as unmapped or unknown unless decision-time evidence supports reconstruction. “No Setup” is not shorthand for impulse, FOMO, or a losing trade.
Tag every eligible executed trade where possible, but never manufacture certainty to reach complete coverage. Report missing labels explicitly and repair the future workflow.
Step 3: Keep Setup, Evidence Quality, and Management Separate
A setup tag answers why entry was eligible. A setup-evidence grade answers whether required fields were present. Management tags record later actions such as partial exits or stop changes. Session, instrument, direction, and market regime are separate dimensions. Mixing them into one label makes interpretation impossible.
A self-rated “confidence” field can be analyzed only if its meaning and timing were fixed before outcome. It is not automatically a quality score and should not overwrite objective evidence coverage.
Reading the Setup Performance Breakdown
Start with coverage and economics, then interpret performance. A clean dashboard that hides unknown or excluded rows is less trustworthy than an incomplete-looking table that exposes them.
| Column | Definition | Failure to avoid |
|---|---|---|
| Setup ID / version | Immutable identity and effective rule | Merging materially different definitions |
| Eligible / executed / skipped | Decision population and opportunity capture | Judging only trades that happened |
| Unknown labels | Records lacking defensible setup identity | Calling unknowns a bad setup |
| Gross wins / gross losses | Positive and absolute negative trade results | Hiding loss magnitude behind win rate |
| Costs and net | Commissions, fees, financing, borrow, slippage, conversion | Comparing gross results with different turnover |
| Expectancy | Mean net result per included trade under declared units | Omitting the denominator or unit |
| Profit factor | Gross profit divided by absolute gross loss | Treating a small denominator as certainty |
| Risk path | Declared drawdown method, exposure, largest loss, clustering | Choosing by return alone |
The expectancy formula and profit-factor guide show the arithmetic and edge cases. Neither metric supplies a universal pass threshold; both inherit the population, cost, and data-quality assumptions of the table.
Example Setup Breakdown: Structure, Not Invented Results
| Population | Coverage state | Observed economics | Next action |
|---|---|---|---|
| Setup A · version 3 | Report eligible, executed, skipped, unknown | Compute only after reconciliation | Apply prewritten verdict rule |
| Setup B · version 1 | Report the same fields and window | Use the same unit and cost basis | Compare only compatible evidence |
| Unmapped | Keep visible | Descriptive totals only | Repair classification; no strategy verdict |
This blank template is intentional. A fabricated percentage would look concrete while teaching nothing about the reader’s setups.
Pairwise Difference: What “Setup B Is Higher-Yielding” Means
Define the direction before calculation:
Δ(B − A) = observed net expectancy of Setup B − observed net expectancy of Setup A
Positive values mean Setup B is higher-yielding in the declared observed sample and unit. Negative values mean Setup A is higher under that convention. Neither sign proves future superiority. Report both denominators, dispersion, cost basis, setup versions, and whether trades share the same market state. If the populations are incompatible, do not reduce them to one delta.
Four Useful Verdict States—Without Universal Thresholds
The baseline used fixed profit-factor and trade-count gates as if they were canonical. Replace them with verdicts tied to the decision and evidence contract.
Category 1: Supported Candidate
Identity and coverage pass, economics are cost-complete, the setup can be executed as written, and the predeclared validation criterion is met. Keep monitoring; “supported” is dated and conditional, not permanent proof.
Category 2: Operationally Valid but Unresolved
The setup was followed and data quality is acceptable, but evidence volume, dispersion, dependence, or regime coverage cannot answer the decision. Continue the frozen version or use simulation; do not tune the threshold midstream.
Category 3: Operationally Broken
The rule may or may not have an edge, but entries, exits, sizing, or data capture repeatedly departed from the definition. Repair execution before judging the setup. The setup-failure analysis separates decision quality, execution quality, and outcome.
Category 4: Rejected Under the Current Rule
The frozen version failed its prewritten validation condition or cannot be operated inside account constraints. Stop or revise only with a new version and new test. Do not erase rejected evidence from the history.
The Evidence Filter: Same Setup, Different Conditions
After the primary setup table is stable, one additional dimension can test a specific hypothesis: session, instrument, direction, volatility rule, or pre-entry evidence completeness. Every split reduces the evidence in each cell and expands the number of comparisons.
Choose the split before viewing its outcome, log every alternative inspected, and retain the unsplit result. If you test many dimensions and report only the best-looking combination, the result is vulnerable to data snooping. Use stratification to locate a future test, not to manufacture a perfect historical subset.
The edge-filtering workflow shows how to freeze one candidate condition, preserve rejected and unknown rows, and validate the selected rule without rewriting the baseline.
The Hidden Deal-Breaker: The Tag Pollution Trap
Three Tag Pollution Patterns
- Outcome-aware retagging: winners receive the “clean” label while losers are moved to mistake or FOMO buckets.
- Version collapse: old and new entry rules share one setup name, so the average describes neither.
- Dimension mixing: setup, session, management, emotion, and result are encoded in one changing free-text tag.
The Tag Hygiene Discipline
Use stable IDs, versioned definitions, effective timestamps, controlled labels, and a correction log. A historical correction should preserve the previous value, reason, author, and time. Report unmapped records; do not force them into the nearest setup merely to make a chart complete.
3 Mistakes Traders Make With Setup Tagging
Mistake 1: Retagging After the Trade Closes
Decision-time identity should not depend on whether the trade won. Post-trade review tags may be added separately, with provenance.
Mistake 2: Splitting Until One Bucket Looks Great
Each instrument/session/direction combination is another trial. Log it and validate the selected rule on untouched or later evidence.
Mistake 3: Turning Unknown Into FOMO
FOMO is a self-reported or otherwise evidenced decision context, not the default label for missing setup data. Unknown stays unknown.
Your Setup Audit Action Plan
- Freeze the decision and active setup versions.
- Reconcile executions, accounts, timestamps, quantities, costs, duplicates, and currency units.
- Audit setup identity coverage before reading P&L.
- Build the primary table without a second dimension.
- Open winners, losers, outliers, and unknowns from every setup.
- State alternative explanations and the exact action each verdict may authorize.
- Freeze one candidate change and validate it on untouched or later evidence.
If setup definitions changed inside the window or the unknown bucket is material to the decision, stop at data repair. A precise answer from polluted identities is worse than an explicit inconclusive result.
Build the Setup Breakdown in TSB
Trader’s Second Brain can keep imported executions, source identity, available fees, primary setup and setup tags, saved setup versions, notes, screenshots, corrections, exclusions, and plan context in one evidence trail. Exact fields depend on the import route and the record you create, so reconcile coverage before trusting a table.
Coach is where this becomes more than a dashboard. Select the evidence set and ask which observed setup differences survive reconciliation, what fields are missing, which underlying trades drive the result, and what bounded test should come next. Its refusal to convert an unmapped label, thin bucket, or one good period into a confident strategy verdict is a strength: the answer remains traceable and useful instead of merely flattering.
TSB has processed 600K+ imported trades across its import history, and its canonical registry recognizes 328 exact broker, exchange, platform, and prop-export profiles. These are imported trades and recognized routes—not users, automatic setup labels, a setup-performance benchmark, a minimum sample, or trades analyzed by Coach.
Find out what your setup labels actually support. Import and reconcile a representative history, define the setup and rule version in your Trading Plan, then use AI Coach to review the selected evidence.
Methodology Note
- Search evidence: one exported A-zone row contained a conversational instruction string with location and response directives. It is denied as untrusted measurement noise—not an instruction, protected heading, or content requirement.
- Metrics: expectancy, profit factor, win rate, and drawdown are descriptive under declared definitions; none supplies a universal edge threshold.
- Trials: repeated tags, thresholds, and slices on the same history increase selection risk; keep a complete trial log and validate forward.
- Product scope: TSB preserves and analyzes provided evidence; this guide does not claim automatic setup assignment.
Final Verdict: Tags Turn a Ledger Into a Testable Decision
A setup breakdown earns authority when identities existed before outcome, versions are separated, coverage is visible, costs reconcile, and the reader can open every supporting trade. Then the table can decide what to keep testing, repair, restrict, or reject.
Do not “double down on what works” from a point estimate or cut a setup from an arbitrary threshold. Freeze the decision rule, validate it on new evidence, and let unresolved setups remain unresolved. The discipline that limits the verdict is what makes the breakdown valuable.
Disclosure: Trader’s Second Brain is our product. Product-scale values are rendered from canonical server truth. This guide is educational and does not promise that a setup, tag, or analytical result will be profitable.