Three checkpoints in this guide
Follow the full walkthrough in order, or jump directly to one of its main sections.
Performance attribution turns one account-level P/L number into a decision ledger. For a discretionary trader, that means grouping the same net trade results by setup, market regime, plan compliance, sizing decision, and instrument. It can show where gains and losses occurred and which question to test next. It cannot, by itself, prove that a category caused the result, separate skill from luck, or predict the next period.
The control that makes the analysis honest is reconciliation: every single-dimension breakdown must add back to the same total. If the account made +3.6R, the setup rows must total +3.6R, the regime rows must total +3.6R, and the instrument rows must total +3.6R. Never add those three dimension totals together; they are three views of the same trades.
Quick answer: start with closed trades, a stable definition of net P/L, and labels recorded without seeing the final outcome. Report trade count, net R, average R, median R, and uncertainty for every category. Treat the output as descriptive evidence. Promote a pattern into a rule only after it survives a predeclared test on later trades.
Important distinction: formal portfolio attribution usually explains return relative to a benchmark and connects excess return to portfolio decisions. The five-dimension method below is a retail trade-review adaptation. It is useful because it follows the decision process, but it is not a Brinson model and should not borrow institutional certainty it has not earned.
Why Attribution Matters More Than Aggregate P/L
Aggregate P/L is indispensable: it is the result that must reconcile to statements and cash. It is simply too compressed to diagnose a process. Three different problems can hide inside the same total.
Problem 1: One Outlier Can Dominate the Period
A profitable month may contain many small, ordinary trades and one unusually large winner. The total is real, but it does not tell you whether the typical decision had positive results. Attribution keeps the outlier visible by pairing total contribution with count, average, median, and largest-trade share. Removing the outlier is not automatically correct; showing both results is.
Problem 2: Regime Exposure Can Resemble Improvement
A trend strategy can earn most of its period P/L during a trending segment. That is useful evidence of conditional fit, not proof that the trader improved or that the result will reverse on schedule. Define regimes independently of the strategy outcome, then compare the same setup across regimes. The market-regime identification framework explains how to freeze that label before using it as an explanation.
Problem 3: A Good Outcome Can Hide a Broken Decision
A trade can violate the written plan and still win. Outcome-only review rewards it; process attribution records the positive P/L and the violation at the same time. The point is not to pretend the win was a loss. It is to keep outcome evidence and policy evidence in separate columns so the result cannot rewrite the rule after the fact.
The Five Attribution Dimensions
These dimensions are a practical schema, not universal laws. Use only labels that map to real choices in your process. An empty or inconsistently populated field should remain “Not verified,” not be reconstructed from memory to complete a chart.
Dimension 1: Setup Quality
Group trades by the setup identifier and, if the grading rubric existed beforehand, by grade. Compare count, net R, average R, median R, dispersion, and largest-trade share. A-grade producing more total P/L is not enough: it may simply contain more trades or larger risk. The stronger comparison is per-unit performance under the same grade definition, with the outlier view beside it.
Dimension 2: Regime Fit
Group the same trades by a predeclared regime label such as trend, range, expansion, or compression. Record the classifier and timestamp because changing the definition can change the result. Regime attribution describes conditional performance in the observed sample; it does not establish when the regime will change.
Dimension 3: Execution Discipline
Compare trades that followed the documented plan with trades that did not. The compliance rule must be auditable: planned entry zone, stop, target or exit policy, risk ceiling, and any permitted adjustment. Use the trade-plan adherence method when a binary label would hide partial execution drift.
Dimension 4: Sizing Decisions
Separate the trade idea from the amount risked. Show actual net P/L, net R based on planned risk, and a counterfactual using one fixed base risk when the necessary inputs exist. The difference between actual and fixed-risk results is a sizing overlay for this sample—not proof that confidence caused better or worse outcomes. The variable-position-sizing guide covers the policy boundary.
Dimension 5: Instrument Selection
Group by exact instrument and contract where applicable. Normalize point value, currency conversion, fees, and contract changes before comparison. Instrument contribution often reflects session, volatility, setup mix, or a single event as well as selection. Cross-tab instrument × setup or instrument × session before acting on an apparent winner.
Attribution Calculation Methodology
Step 1: Freeze the Data Contract
For each closed trade, preserve a trade ID, account, open and close time, exact instrument, side, quantity, gross realized P/L, commissions and fees, planned risk at entry, setup, regime, plan-compliance evidence, and sizing-policy label. Keep the raw import value beside any normalized value. If planned risk or a category is absent, mark the row missing rather than inventing it.
Net P/L = gross realized P/L − commissions − fees
Net R = net P/L ÷ planned loss at entry
Net R makes trades with different planned risk more comparable, but only when the denominator was known at entry and uses a consistent currency basis. It is not valid to infer planned risk from the final loss after slippage or a moved stop.
Step 2: Reconcile Every Dimension
This small illustrative ledger has 12 trades and +3.6R total. Each dimension independently reconciles to +3.6R:
| Dimension | Category | Trades | Net R | Average R |
|---|---|---|---|---|
| Setup | Breakout | 7 | +4.1R | +0.59R |
| Setup | Pullback | 5 | −0.5R | −0.10R |
| Regime | Trend | 8 | +3.0R | +0.38R |
| Regime | Range | 4 | +0.6R | +0.15R |
| Plan | Followed | 10 | +4.3R | +0.43R |
| Plan | Deviated | 2 | −0.7R | −0.35R |
The setup subtotal is +3.6R. The regime subtotal is also +3.6R. The plan subtotal is also +3.6R. Adding them would report +10.8R and triple-count every trade. This example demonstrates the accounting rule only; its 12 rows are far too few for a durable conclusion.
Step 3: Add Distribution, Not Just Contribution
For every category, report the denominator and shape of results: number of eligible trades, missing-label count, total net R, average, median, standard deviation or interquartile range, largest winner, largest loser, and share of P/L from the largest trade. A category with +4R from one trade is a different evidence state from +4R distributed over 40 trades.
Step 4: Test Interactions
A single pivot can confound variables. Breakout trades may appear strongest only because most occurred in the trend regime. Build a setup × regime table, then inspect the same setup across regimes and the same regime across setups. Add cells only when they answer a written question; slicing every field creates tiny samples and false discoveries.
Step 5: Separate Observation From Decision
Write the observed pattern first, then a competing explanation, then the next test. Example: “Pullbacks were −0.5R across five trades; four occurred in range conditions. Keep the setup unchanged and collect the next 20 eligible signals with the regime label frozen.” This is slower than declaring the setup broken, but it produces a falsifiable decision.
How Much Data Is Enough?
There is no universal “60 trades means confidence” rule. Required evidence depends on effect size, outcome variance, dependence between trades, number of categories, and how costly the decision is. Ten similar trades can expose a data-quality failure; hundreds may still be weak evidence for a small edge split across many regimes.
Use uncertainty rather than a magic cutoff. Show a bootstrap interval for the category mean, but preserve trading sequence when outcomes cluster—for example, resample days or sessions rather than pretending every trade is independent. Andrew Lo’s work on Sharpe ratios shows why serial correlation changes performance inference; the same warning applies whenever adjacent trades share regime and risk.
Also record how many ideas you tested. If 30 tags are searched for the best-looking subset, one can appear exceptional by chance. Backtest-overfitting research makes the general point: more tried configurations raise the chance of selecting a historical mirage. Predeclare the primary split, keep exploratory findings labeled exploratory, and confirm them on later data.
Decision Implications by Attribution Result
Setup Contribution Is Concentrated
Check whether the setup has enough observations, whether one trade dominates, and whether the result survives regime and time splits. If it does, prioritize the setup for a forward observation window; do not automatically increase risk.
Returns Are Concentrated in One Regime
Confirm that the regime label was independent of outcome and compare identical setups across regimes. The immediate action is usually a permission or sizing test, not a prediction that the current regime is about to end. If multiple strategies are involved, the multi-strategy portfolio framework helps inspect overlap rather than assuming diversification.
Off-Plan Trades Look Better
Do not reward the violation or hide its profit. Audit whether the written rule is obsolete, whether the compliance label is wrong, or whether a small sample rewarded uncontrolled risk. Freeze any revised rule prospectively and compare later eligible trades.
Actual Sizing Beat Fixed Sizing
Inspect whether larger positions happened to coincide with easier regimes, stronger setups, or outliers. The sizing overlay is a useful clue, but changing capital allocation requires a separate, predeclared test with hard risk limits.
No Clear Pattern Appears
That is a valid result. Preserve the stable policy, improve missing fields, and wait for more relevant observations. An ambiguous attribution report is safer than a precise story manufactured from noise.
Operational Implementation Framework
- Choose one decision. Example: keep, narrow, or pause a setup—not “analyze everything.”
- Freeze labels and eligibility. Define the setup, regime classifier, account scope, date range, costs, and missing-data rule before viewing the result.
- Reconcile to source statements. Closed-trade count and net P/L must match the selected account and period.
- Run one-dimensional pivots. Confirm that each one totals to the same result.
- Inspect distributions and interactions. Look for outliers, sparse cells, and confounding.
- Write the limitation. State what the data cannot distinguish.
- Set one forward test. Name the next eligible sample, fixed rule, decision threshold, and review date.
Review cadence should follow the decision horizon. A high-frequency process may accumulate a useful weekly sample; a swing strategy may need a quarter or more. Calendar frequency is not evidence quality.
Who Should Prioritize Attribution Analysis
- Traders changing a strategy after one strong or weak period: attribution exposes concentration and missing comparisons before the change becomes permanent.
- Multi-setup traders: it separates setup mix from account-level results.
- Prop traders: it can distinguish strategy P/L from the effect of program constraints, provided exact rule state is preserved.
- Systematic researchers: the same ledger can be applied to out-of-sample results, with trial count and model-selection history retained.
- Coaches and teams: immutable labels make the review about evidence rather than memory or authority.
How Trader's Second Brain Turns Attribution Into a Review Loop
Disclosure: Trader's Second Brain (TSB) is our product. We built it to keep imported execution, review evidence, analytics, and the next rule connected instead of scattering them across a broker export, spreadsheet, screenshots, and chat.
TSB has processed 600K+ imported trades and recognizes 331 exact source profiles. Those figures mean imported trades and recognized import routes—not users, universal compatibility, or trades analyzed by AI Coach. The exact route still has to reconcile before its metrics are trusted.
The powerful part is not another pivot table. Reports can expose the setup, session, instrument, risk, fee, and sequence pattern; Leak Map can turn supported negative comparisons into a ranked review surface; Current Focus can keep one evidence-backed rule in front of the trader; and Retrospective Backtester can test a clearly defined historical filter. AI Coach sits across those surfaces and can interrogate the selected evidence set, explain the strongest supported relationship, name the limitation, and turn it into one precise next review. That is the difference between collecting statistics and running a learning system.
Coach does not need to become timid to stay grounded. When the required fields exist, it can be direct about what the account evidence shows and compare the exact relevant cohorts. When setup, regime, size, costs, or eligible trades are missing, it says what is insufficient and routes the user to the best available lens instead of fabricating a personalized diagnosis. That refusal is evidence integrity—not a weaker product.
Start with one question: which setup, regime, execution rule, sizing choice, or instrument deserves the next controlled test?
Methodology Note
- Formal definition: CFA Institute describes performance measurement, attribution, and appraisal as separate layers, and says an effective attribution process should reconcile to total return or risk and reflect the actual decision process.
- Retail adaptation: the five dimensions in this guide are an editorial trade-ledger schema. They are not presented as a universal institutional model or a causal estimator.
- Statistical boundary: return estimates carry sampling error and serial dependence can change inference. Multiple exploratory cuts also raise selection risk.
- Data boundary: missing contemporaneous labels remain missing. Retrospective labels are allowed only when visibly marked and are not used as equivalent evidence.
- Decision boundary: attribution chooses the next test; it does not authorize a larger position, promise persistence, or establish causality.
Primary references: CFA Institute portfolio performance evaluation, CFA Institute performance attribution history, Andrew Lo, The Statistics of Sharpe Ratios, and Bailey et al. on backtest overfitting.
For our editorial controls, see TSB editorial methodology.
Final Verdict: Attribution Reveals What P/L Hides
A good attribution report does three things: it reconciles exactly to account results, preserves the decision labels that existed before the outcome, and ends with a testable next action. It does not add overlapping dimensions, turn correlation into causation, or certify skill from a small historical slice.
Build the first report with five columns before five dimensions if necessary. Clean trade IDs, net P/L, planned risk, setup, and timestamps are more valuable than a sophisticated chart over reconstructed data. Once the ledger is trustworthy, add regime, compliance, sizing overlay, instrument interactions, and uncertainty in that order.
Turn the P/L Total Into Evidence You Can Act On
Import the trades, reconcile the period, expose the strongest supported pattern, and let AI Coach turn it into a focused review without inventing the missing story.
Build the attribution review