Trader's Second Brain Trader's Second Brain

Setup Failure Analysis: Why Good Setups Fail

A losing trade does not prove that a good setup failed, and a winning trade does not prove that a bad decision was sound. This evidence-first framework separates the rule, execution, and outcome; organizes reviews into five audit buckets plus Unknown; and raises the standard before any new criterion becomes permanent. Unsupported failure percentages, magic sample counts, and post-hoc chart signals are removed.

Quick Answer

Compare the trade with the rule and evidence that existed before the result. If the rule was followed and no process defect is documented, record a rule-compliant loss—not a guessed cause. Change a criterion only after a pre-written hypothesis is tested across all relevant trades and survives data that were not used to invent it.

Performance lab · setup-level evidence
Setup, session and drawdown review

Know which setups actually have an edge.

Find which setups, sessions, and behaviors make or lose money.

Find weak setups →
Trader's Second Brain preview
Reading map

Three checkpoints in this guide

Follow the full walkthrough in order, or jump directly to one of its main sections.

  1. 01Opening checkpointSeparate the Decision, Execution, and Outcome
  2. 02Middle checkpointUse “Unknown” When the Evidence Is Missing
  3. 03Closing checkpointFinal Verdict: Diagnose the Process Before You Rewrite the Setup

A losing trade does not prove that a good setup failed, and a winning trade does not prove that a bad decision was sound. The outcome is evidence about a distribution; it is not a complete diagnosis of the decision. A useful post-trade review asks what was knowable before entry, whether the documented rule was followed, whether execution matched the plan, and which explanation remains only a hypothesis.

This guide provides five audit buckets, an “unknown” escape valve, and a repeatable diagnosis record. It deliberately removes unsupported percentages for how often each failure occurs, fixed trade-count rules, and chart-reading shortcuts that assign causes after one loss.

Method checked September 9, 2026. Baron and Hershey's outcome-bias experiments and Fischhoff's hindsight study support separating decision quality from known outcomes. The CFA Institute's investment model validation guide supports out-of-sample and time-series validation for strategy-level changes. These are general decision and model-validation sources, not estimates of retail-trading failure-mode frequencies.

Quick answer: compare the trade with the rule and evidence that existed before the result. If the rule was followed and no process defect is documented, record a rule-compliant loss—not a guessed cause. Change a criterion only after a pre-written hypothesis is tested across all relevant trades and survives data that were not used to invent it.

Separate the Decision, Execution, and Outcome

A setup is a selection rule. Execution is what the order and trader actually did. Outcome is the market path after entry. Mixing these three makes every loss look like a setup defect and every win look like validation.

LayerQuestionBest evidence
DecisionWas the trade eligible under the rule version active at entry?Timestamped plan, checklist, setup tag, pre-entry note or image
ExecutionDid the actual order match planned price, size, stop, target, and order type?Orders, fills, modifications, fees, slippage, timestamps
OutcomeWhat path occurred, and was the result inside the strategy's observed range?Market data, realized P&L in R, excursion and exit record

Review the decision using only information available at the time. Then reveal the outcome. This simple order reduces the temptation to rewrite the setup around what the next candles happened to do. The journal-mistakes guide covers missing plans, inconsistent tags, and edited notes that make this audit impossible.

The Five Setup Failure Audit Buckets

These buckets organize evidence; they are not universal causes with known percentages. A trade may contain more than one issue. If the record cannot support a classification, use Unknown instead of forcing a story.

Failure Mode 1: Rule-Compliant Losing Outcome

The setup definition was met, the trade was authorized by the plan, execution stayed within tolerance, and no documented exclusion applied. The trade still lost. The defensible label is rule-compliant loss. Calling it “pure variance” would be stronger than the evidence: one outcome cannot identify the generating mechanism.

Evidence to Record

  • Rule version and each eligibility field as it appeared before entry.
  • Planned versus actual entry, stop, size, target, and risk.
  • Known costs and whether the loss was within the plan's risk tolerance.
  • Any information that arrived after entry, clearly separated from pre-entry facts.

Response: do not add a new criterion from this trade alone. Add the observation to the strategy distribution and continue the pre-committed review plan. Calculate the strategy-level expected value with the expectancy workflow, not with a verdict attached to one row.

Failure Mode 2: Eligibility or Criteria Deviation

The active rule would not have authorized the trade, or the reviewer cannot reproduce the eligibility decision. Examples include a required field left blank, a setup tag applied after the outcome, a minimum condition marked “close enough,” or a rule version that changed without a date.

Diagnostic Indicators

  • A checklist field conflicts with the written plan.
  • The trade lacks evidence for a required criterion.
  • The rule used in review differs from the rule active at entry.
  • Exceptions are remembered but were not logged before execution.

Response: fix compliance and evidence capture before changing the strategy. Tightening the rule does not solve a process that cannot prove which rule it followed.

Failure Mode 3: Execution or Market-Friction Deviation

The setup was eligible, but the filled trade differed materially from the plan. Possible evidence includes a late or partial fill, wrong quantity, wrong contract, missing stop, rejected modification, spread, slippage, commission, or an exit that did not follow the recorded rule.

What Not to Infer

A stop reached soon after entry does not by itself prove the entry was late. A hypothetical better price does not prove that price was obtainable. Diagnose execution from the order and fill trail, not from a cleaner point visible after the chart is complete.

Response: define an execution tolerance, identify the actual deviation, and choose a process fix—order template, instrument check, alert, or checklist. Keep the setup criteria unchanged unless a separate aggregate test supports a setup change.

Failure Mode 4: A Known Context Rule Was Missed

Context belongs in this bucket only when the condition and response were defined before the trade. For example, the plan may prohibit new entries around a specified event, require a particular session, or exclude a volatility state measured by a named rule.

Evidence to Require

  • The context field has a reproducible definition and source.
  • The data were available before entry.
  • The active plan stated what to do when the condition was present.
  • The review does not substitute a new chart narrative for the original rule.

Response: repair the context check or exception workflow. If the condition was not part of the plan, record it as a candidate hypothesis under the next bucket—not as a proven missed rule.

Failure Mode 5: Strategy or Regime Hypothesis

A cluster of rule-compliant losses may justify asking whether the strategy's opportunity set, market state, costs, or implementation has changed. It does not prove a “regime mismatch.” Regimes need operational definitions; otherwise the label merely renames a drawdown.

Build a Testable Hypothesis

  • Name the condition and how it is measured without looking at future data.
  • Explain why it could affect the setup's mechanism or execution.
  • Specify the comparison, metrics, and decision before testing.
  • Include all relevant observations, costs, and failed versions.
  • Reserve later data or use an appropriate time-series validation design.

Response: continue, reduce, pause, or retire the strategy only under the risk and review rules you set in advance. The strategy-abandonment guide separates a tolerable drawdown from a breached premise and a broken process.

Use “Unknown” When the Evidence Is Missing

Unknown is not a failed review. It identifies the field or timestamp the next trade must preserve. Common reasons include missing pre-entry evidence, incomplete fill data, ambiguous timezone, unmatched account or instrument, an edited rule with no version history, and a manual import that omitted fees.

Do not distribute unknown trades among the five buckets to make the dashboard add up. Track the unknown share. If it is high, the priority is data quality—not strategy optimization.

Hidden Deal-Breaker: Outcome and Hindsight Bias

Baron and Hershey found that people rated decision quality differently when they knew whether an uncertain decision ended favorably. Fischhoff found that outcome knowledge changed judgments about what appeared foreseeable. Neither study tested retail trading, but both show why a completed chart can contaminate a process review.

The trading version looks like this: after a loss, a feature appears to have been an obvious warning; after a win, the same feature is ignored or explained away. The reviewer adds rules from memorable losses while failing to record the full comparison set.

Controls That Make the Review Falsifiable

  • Grade eligibility before revealing the next market path when practical.
  • Keep the original screenshot, plan, and checklist immutable.
  • Log candidate criteria from wins, losses, and skipped trades.
  • Record every threshold and filter tested, not only the best one.
  • Make “no change” and “need more data” legitimate outcomes.

Validate a New Criterion Without a Magic Trade Count

There is no universal “30 trades” threshold or win-rate gap that proves a condition predicts losses. Required evidence depends on outcome dispersion, dependence between trades, base rate, costs, search breadth, and the consequence of being wrong.

  1. Write the candidate rule. Define the characteristic, expected direction, metric, comparison, and proposed action.
  2. Audit the data. Confirm account, instrument, timezone, duplicates, fees, and rule version.
  3. Use the complete comparison set. Include trades with and without the characteristic and retain missing states.
  4. Measure more than win rate. Compare expectancy, payoff distribution, drawdown, outliers, cost, and opportunity frequency.
  5. Count every test. Multiple tried definitions make the best historical result less persuasive.
  6. Freeze and validate. Apply the unchanged rule to data that were not used to create it.

For systematic strategies, the backtest-versus-live guide covers leakage, implementation gaps, and why a historical fit is not an execution result.

The Failure Diagnosis Framework

StepCheckOutput
1. IntegrityReconcile account, instrument, timestamps, quantities, fees, and duplicatesUsable record or Unknown
2. EligibilityCompare pre-entry evidence with the active rule versionCompliant or criteria deviation
3. ExecutionCompare plan, orders, fills, modifications, and exitWithin tolerance or execution deviation
4. ContextTest only context exclusions documented before entryNo miss or known context-rule miss
5. OutcomeRecord R result, excursion, costs, and path without assigning causalityRule-compliant loss record
6. HypothesisIf a pattern is suspected, pre-write an aggregate validation planCandidate strategy/regime test

Assign a confidence level and cite the evidence. One trade can establish that a checklist was blank or an order size was wrong. It usually cannot establish that a new market condition predicts failure.

Aggregate the Audit Without Inventing a Normal Distribution

Review on a pre-committed horizon suited to your trade frequency. Report counts and rates for every bucket, including Unknown, but do not compare them with an invented “typical retail” distribution. The relevant baseline is the same strategy, rule version, account scope, and data process.

Investigate changes alongside exposure and opportunity. A higher number of execution deviations may reflect more trades, not a worse process. A cluster of rule-compliant losses may be compatible with the observed payoff distribution, or it may motivate a strategy hypothesis; the audit alone does not choose between them.

Trader's Second Brain is our product. It can preserve imported executions, setup and context tags, notes, screenshots, and review fields so the decision/execution/outcome split remains auditable. Its canonical source registry currently recognizes 330 structured broker, exchange, platform, and prop-export profiles. That is parser coverage, not a guarantee that every source contains every field.

Check your exact route in the live import directory, validate a representative file, and reconcile the imported record before diagnosing. Full Access includes review and rule-tracking workflows and offers a lifetime-access route; verify current terms rather than relying on a copied price.

Review the TSB post-trade audit workflow

Who Should Prioritize Failure Analysis

  • Traders who change rules after individual losses: separate a process defect from an unfavorable result before adding complexity.
  • Teams with inconsistent reviews: use the same fields, evidence hierarchy, and allowed labels.
  • Traders who suspect execution leakage: compare plans with fills and costs rather than judging the completed chart.
  • Strategies in a drawdown: distinguish data, compliance, execution, and hypothesis work before changing risk.
  • Prop-program traders: keep program-rule breaches separate from whether the setup itself had positive expectancy.

Any response must stay inside the account's loss limits and the trader's risk capacity. The risk-management framework covers exposure, stop logic, drawdown capacity, and when investigation should not become another live experiment.

Methodology Note

The five buckets are an editorial audit taxonomy, not a validated universal distribution of trading losses. Outcome- and hindsight-bias research motivates the evidence order, but does not quantify trader behavior. CFA's model-validation guidance applies most directly to formal investment models; this guide borrows its core discipline of independent validation without claiming that discretionary trade review is equivalent to institutional model governance.

No fixed mode percentage, stop-timing signal, sample count, win-rate gap, regime duration, or review interval is presented as universal. Those claims require strategy-specific data and an analysis that accounts for dependence, costs, multiple testing, and changing conditions.

Final Verdict: Diagnose the Process Before You Rewrite the Setup

For one loss, prove what you can: data integrity, eligibility, execution, and known context compliance. Record the outcome without pretending its cause is observable from the chart alone. Use Unknown when the evidence is missing.

For a strategy change, raise the standard. State the hypothesis, use the complete comparison set, log every test, freeze the candidate rule, and validate it on later data. The best failure analysis does not explain every loss; it prevents an unsupported story from becoming a permanent trading rule.

Igor Manuilov
Written and reviewed by
Igor Manuilov
Founder of Trader's Second Brain · Trader since 2014
Editorial accountability

Trader since 2014. Built Trader's Second Brain to make execution review more evidence-based and less dependent on memory, scattered spreadsheets, or vague journaling.

Performance lab · setup-level evidence
Setup, session and drawdown review

Turn trading statistics into a review plan.

Find which setups, sessions, and behaviors make or lose money.

Find weak setups →
Trader's Second Brain preview

Frequently Asked Questions

Quick answers to the most common questions about Setup Failure Analysis.

A valid setup can lose because eligibility does not guarantee an outcome. Audit five evidence buckets: rule-compliant losing outcome, criteria deviation, execution or market-friction deviation, missed pre-defined context rule, and a strategy or regime hypothesis that needs aggregate validation. Use Unknown when the record cannot support a classification; there is no universal percentage for any bucket.

Not from the outcome alone. Record the candidate condition, define how it would be measured before entry, use the complete comparison set, log every tested version, and validate the frozen rule on data not used to invent it. There is no universal 30-trade or win-rate-gap threshold that proves the criterion.

One trade rarely identifies the generating cause. First prove data integrity, rule eligibility, execution, and compliance with context rules that existed before entry. If those pass, label it a rule-compliant loss. A strategy or regime problem is a separate hypothesis that needs an operational definition and aggregate, preferably out-of-sample, testing.

A major error is judging the decision by the known outcome and then adding a rule around a feature that looks obvious on the completed chart. Preserve the original plan and evidence, review eligibility before the later path when practical, include wins and skipped trades in comparisons, and keep a log of every candidate criterion tested.

There is no universal duration. A regime label must have an operational, pre-entry definition and a response specified by the risk plan. A cluster of losses can motivate that hypothesis, but it does not prove the regime caused them. Validate the condition across an appropriate time-series design before making it permanent.

Review each loss consistently if your process can do so without changing labels after seeing the aggregate. Capture evidence, classify what the record supports, and use Unknown when it does not. A per-trade audit can prove a rule or execution deviation; it should not automatically trigger a new strategy criterion.

Use individual trades to audit data, compliance, and execution; use a pre-committed aggregate review to test strategy-level hypotheses. Report every bucket, including Unknown, against the same strategy and rule version. Do not compare your results with invented universal failure-mode percentages.