Correction and evidence check: September 7, 2026. The earlier version treated unpublished TSB-user patterns as universal evidence and attached fixed percentages, trade-count thresholds, timing rules, forward P/L predictions, and a 90-day improvement promise to them. Those claims are removed. This revision keeps the protected title and A/B/C search intent, but makes every grade depend on a written, versioned rubric and keeps process quality separate from outcome statistics.

Why P/L Alone Is Misleading

P/L answers what happened; a trade grade records whether the decision and execution matched a plan that existed before the outcome. Neither measure can replace the other. A profitable result does not prove that an unplanned entry was sound, and a loss does not prove that a plan-compliant entry was unsound. Equally, an A-grade does not prove that the underlying strategy has positive expectancy.

Baron and Hershey’s original outcome-bias experiments found that participants rated decisions more favorably when their outcomes were favorable even though the relevant pre-decision information was held constant. The five studies used students judging other people’s medical decisions or monetary gambles—not traders grading their own trades. Applying the finding to trade review is therefore a useful risk model, not direct evidence that every trader will show the same bias.

Is This a Good Trade? Ask Two Different Questions

QuestionEvidenceValid conclusionInvalid shortcut
Was this trade executed well?Plan snapshot, entry trigger, risk, size, stop, management, exit, and rule evidenceA/B/C execution grade under a named rubric versionWinner = good trade
Does this setup have an edge?Comparable cohort, net R or P/L, costs, sample count, dispersion, and time periodDescriptive expectancy with uncertainty and validationA-grade = profitable strategy

An A-grade loss can be a higher-quality execution than a C-grade win. It is not automatically “better for long-term performance”: that stronger claim needs a stable rubric, comparable observations, and evidence that plan-compliant trades have useful expectancy after costs.

The A-B-C Grading Framework

The letters are labels, not a universal standard. Write the conditions before using them, name the version, and apply the same version to every trade in the comparison window. The conservative rubric below grades the most serious observed deviation; a harmless detail cannot average away a risk breach.

A-Grade: Plan-Compliant Execution

  • The setup and invalidation were written or captured before entry.
  • Instrument, session, direction, entry trigger, size, stop, and total open risk stayed inside the active plan.
  • Management and exit followed the declared rule, or a permitted exception was documented by contemporaneous evidence.
  • No material deviation increased risk or changed the trade thesis after entry.

An A does not mean perfect prediction, calm emotion, or profit. It means the evidence satisfies this rubric.

B-Grade: Contained, Predefined Deviation

  • The setup and risk thesis remained valid.
  • One or more deviations crossed a written tolerance, but did not break a hard risk, eligibility, or invalidation rule.
  • The exact deviation is named—for example, entry timing outside the preferred window—rather than softened as “close enough.”

Do not invent percentage bands after seeing the trade. If the plan has no tolerance for a field, define one prospectively or score that field only as pass/fail.

C-Grade: Material Rule Break or Ungradable Trade

  • The trade had no usable pre-entry plan or setup identity.
  • A hard risk, size, stop, session, instrument, daily-loss, or eligibility rule was broken.
  • The invalidation was removed, risk was added without a declared rule, or the post-trade story cannot be reconciled with contemporaneous evidence.

A C can win. That result belongs in the outcome column; it does not repair the process violation.

Tie-break rule: assign the lowest grade triggered by any material field. Keep “not enough evidence” separate when the record cannot support a grade; silently converting missing data to B or C contaminates both the completeness rate and the grade distribution.

How to Grade Every Trade Without Rewriting the Past

1. Freeze the Pre-Trade Evidence

Before entry, preserve the setup version, required conditions, invalidation, intended entry, planned stop, planned size, maximum open risk, management rule, and permitted exceptions. A screenshot or timestamped plan helps distinguish what was known then from what became obvious later.

2. Record Actual Execution Separately

Store actual fills, size, stop changes, partial exits, costs, rule events, and the evidence timestamp. Do not replace the planned values with the final values. The grade is the comparison between those two records.

3. Apply One Rubric Version

Evaluate each field, record the decisive reason, then apply the tie-break rule. Grade as soon as the required evidence is complete, but do not claim that an arbitrary five-minute or same-day deadline is scientifically necessary. The operational safeguard is that the evidence and rubric cannot be rewritten to fit the result.

4. Store Outcome in a Different Column

Save net P/L, net R, MAE/MFE when available, and any rule outcome without letting those fields alter the execution grade. Later analysis can cross-tab grade and result. A single row should always answer both questions: “how was it executed?” and “what happened?”

5. Audit Grade Consistency

Periodically take a small random sample, hide outcome columns where practical, and apply the same rubric again. Report the number regraded, exact agreement, and the A↔B, B↔C, and A↔C disagreements. For example, matching 12 of 15 regraded rows is 80% raw agreement; the three disagreements—not a universal benchmark—show which rubric boundaries need clarification.

How to Analyze Your Quality Data

Start with data quality, then process distribution, then outcomes. Reversing that order makes it easy to choose the interpretation that flatters the latest P/L.

  1. Declare the cohort: setup version, market, session, phase, account, date range, and inclusion/exclusion rules.
  2. Report completeness: eligible trades, graded trades, missing plans, and missing outcomes.
  3. Count each grade: show A/B/C counts and percentages with their denominator.
  4. Measure process: plan-followed rate, deviation reasons, risk breaches, and grade-agreement audit.
  5. Measure outcomes: net R or net P/L, win/loss/breakeven counts, average win, average loss, expectancy, drawdown, and costs by grade.
  6. Validate a change: name one rule or checklist change before the next comparable period, then retain both favorable and unfavorable observations.

A Transparent Grade-Level Example

The table is constructed arithmetic, not TSB customer data or a performance promise. Breakevens are absent from this toy sample; a real report must state how they are handled.

GradeTradesWins / lossesAverage winAverage lossCalculated expectancy
A1810 / 81.20R0.80R+0.31R per trade
B105 / 51.10R0.90R+0.10R per trade
C64 / 20.60R1.40R−0.07R per trade

expectancy = win rate × average win − loss rate × average loss. For A, that is (10 ÷ 18 × 1.20R) − (8 ÷ 18 × 0.80R) = 0.31R. The calculation is exact for the rows supplied; its stability outside this sample is not.

How Many Trades Do You Need?

There is no defensible universal answer such as 50 for “moderate confidence” or 100 for “high confidence.” Required sample size depends on the effect you want to detect, variability, dependence between trades, the number of subgroups inspected, and the error rates you accept. NIST’s sample-size guidance for proportions makes the detectable change, significance level, power, and expected proportions explicit.

Always show the count. NIST also documents Wilson and exact binomial intervals and warns that simple symmetric approximations can be poor for small samples or rare outcomes. In the constructed A-grade row, 10 wins in 18 trades is 55.6%, while a two-sided 95% Wilson interval is approximately 33.7%–75.4%. That width is why “A wins more” should remain descriptive until substantially more comparable evidence accumulates.

What If C-Grade Trades Are More Profitable?

Do not upgrade them. Check four separate hypotheses: the C sample may be tiny; the result may be concentrated in one outlier; the rubric may classify a genuinely useful setup as a violation; or the written plan itself may lack edge. Freeze the current rubric, inspect the full distribution and costs, and test any revised rule on later or otherwise unseen trades. The answer can be “uncertain.”

Patterns to Test—Not Patterns to Assume

ObservationCompeting explanationsNext check
C-grades appear after lossesReaction to loss, specific session conditions, or a changing strategy mixCompare prior-result sequences within the same setup and session
B-grades cluster at certain timesFatigue, volatility regime, spread/cost changes, or vague timing rulesCompare the same setup across time buckets and inspect the named deviations
A-grade outcomes look strongerBetter execution, easier setups receiving A, different risk, or selective missing gradesMatch cohort and risk; report missingness and uncertainty
A-grade share risesReal improvement, rubric drift, easier market conditions, or grade inflationKeep rubric version fixed and blind-regrade a sample

These are investigation paths. The old claims that particular grade frequencies predict P/L 30–90 days ahead, that A-grade win rate should exceed overall win rate by a fixed range, or that a cooling interval converts a fixed share of grades had no cited dataset and are not retained.

How to Increase A-Grade Execution Without Gaming the Grade

Choose One Failure Mode

Rank C-grade reasons by frequency and risk consequence, then choose one. “Late entry after the planned trigger expired” is actionable; “bad discipline” is not. Preserve the other categories so the total cannot improve merely by relabeling them.

Turn It Into a Pre-Trade Check

Make the control observable before the order: setup version selected, invalidation captured, quantity calculated, open risk updated, or session eligibility confirmed. A checklist records whether the gate happened; it cannot guarantee a favorable outcome.

Declare the Next Comparison Window

Before collecting new results, write the cohort, metric, minimum useful observations for your own precision goal, and rollback condition. Keep the old rubric and rows. Do not move a boundary because the first few trades look uncomfortable.

Review Process and Outcome Together

A rising A share with worsening A-grade expectancy can mean the strategy changed, the market changed, the rubric is too easy, or costs rose. A falling C share is not enough to justify larger risk. Process improvement and strategy evidence are two gates.

Mistakes That Break Trade Quality Scoring

  1. Letting outcome set the grade. A winner is upgraded or a loser downgraded even though the plan evidence is unchanged.
  2. Changing the rubric mid-window. The same action receives different grades without a version boundary.
  3. Using vague traits. “Confident,” “disciplined,” or “felt right” cannot be audited like size, stop, trigger, or rule adherence.
  4. Forcing every trade into A/B/C. Missing plans and missing evidence need a visible ungradable state.
  5. Hiding the denominator. A percentage without counts conceals tiny cells and incomplete reviews.
  6. Comparing unlike cohorts. Different setups, market regimes, risk units, or rubric versions can create a false grade effect.
  7. Optimizing the distribution. More A labels are meaningless if the boundary was softened or A-grade expectancy remains unsupported.

Who Should Pause Quality Grading for Now?

  • No written plan: capture setup, risk, invalidation, and management rules before attaching a quality label.
  • Strategy under active rewrite: finish a rubric version and mark a clean boundary before comparing periods.
  • Incomplete execution evidence: repair the data or report “not enough evidence” instead of inferring compliance.
  • Fully systematic execution: prefer explicit rule-conformance events and implementation errors; a subjective letter may add little.

Low frequency is not, by itself, a reason to skip. A grade can still improve a single-trade review; it simply may not support a stable comparison between grade buckets.

Grade and Audit Trades in Trader’s Second Brain

Trader’s Second Brain is our product. In the current local build, Post-Trade Review stores a Setup, manual Execution Grade (A/B/C), Plan Adherence (Yes/Partially/No), a Main Mistake when adherence is partial or broken, and an optional Lesson. Review completeness requires setup, grade, and plan adherence; a mistake tag becomes required for a partial or broken plan.

TSB does not automatically certify that an A-grade was honest. The manual grade records the reviewer’s judgment under their rubric. The product’s setup-level Strategy Grades are a different output calculated from observed win rate, profit factor, expectancy, and sample count; they must not be substituted for the per-trade execution grade. Review Mode separately reports completed versus pending reviews, plan-followed rate, and the most frequent tagged mistake in the active scope.

Complete the next evidence-backed trade review

Open a pending trade, compare plan versus actual execution, assign the manual grade, and record the decisive deviation without rewriting the outcome.

Review the next trade in TSB →

Methodology Note and Sources

The rubric and examples are an editorial measurement framework, not an industry standard, trading advice, or a performance guarantee. No exact prop firm or program participates in this decision, so a firm/program catalog component would be artificial and is not added. The layout retains Article and BreadcrumbList; the page is not a visible product rating and adds no ItemList, Review, Rating, or Product schema.

Final Verdict: Grade the Process, Then Test the Outcome

A useful trade quality score is boringly auditable. Freeze what was planned, capture what actually happened, apply one rubric version, record the decisive reason, and keep P/L in a separate field. An A-grade loss and a C-grade win can coexist without contradiction because execution quality and outcome are different variables.

The grade becomes evidence only at the cohort level—and only with visible counts, missingness, costs, uncertainty, and comparable trades. If the rubric changes, start a new version. If the record is incomplete, say so. If A-grade expectancy is weak, investigate the plan rather than manufacturing more A labels.

Continue with trade quality versus P/L, the trade-review workflow, journal field design, and expectancy and edge measurement.