Three checkpoints in this guide
Follow the full walkthrough in order, or jump directly to one of its main sections.
A trading report card should grade the process against a frozen strategy contract—not rank the trader against invented universal benchmarks. An A+ is useful only when the rubric, evidence window, costs, and uncertainty are visible. Otherwise a letter grade compresses noise into false confidence.
This framework keeps the familiar F-to-A+ summary while making every grade traceable to execution, edge, risk, and evidence quality.
Quick answer: grade four separate dimensions—data integrity, rule execution, net outcome evidence, and risk containment—against predeclared strategy-specific criteria. Show the underlying metrics and sample limitations, let the weakest safety dimension cap the composite, and compare grade trends only within the same rubric version.
What a Report Card Is For
The report card is a review interface. It should help answer: Is the dataset trustworthy? Did the trader execute the declared system? Does the system show net evidence after costs? Did the account stay inside its risk contract? A single P/L value cannot answer those questions.
The grade is not a credential, prediction, or promise that the next period will resemble the last. Its job is prioritization: expose the weakest verified dimension and link it to the records that created the result.
Every grade should therefore be reversible when the source data or calculation is corrected.
Build the Rubric Before Seeing the Grade
- Define the unit. Grade one strategy version, account scope, market set, and evaluation window.
- Name each metric. Specify formulas, currencies, fee treatment, deposits, partial fills, and missing-data handling.
- Set evidence bands. Define A through F criteria from the strategy’s tested expectations and risk contract—not generic retail averages.
- Add hard caps. A serious data-integrity failure or risk breach can cap the composite regardless of profit.
- Freeze the version. Changing a threshold creates a new rubric; do not rewrite prior grades silently.
The trade-quality score guide shows how to keep process quality separate from whether an individual trade won.
The Four-Dimension Trading Strategy Report Card
| Dimension | Evidence | Primary failure |
|---|---|---|
| Data integrity | Complete trades, reconciled fills, stable tags, costs, version IDs | The grade cannot be trusted |
| Execution quality | Eligible signals, plan compliance, risk and exit adherence | Results do not represent the declared system |
| Net edge evidence | Expectancy distribution, costs, stability and untouched evidence | Apparent edge may be noise or leakage |
| Risk containment | Drawdown path, tail loss, concentration, account-limit utilization | A profitable period can still be unacceptable |
Keep the dimensions visible. A composite without components hides whether a good result came from repeatable execution or one concentrated outcome.
Define F to A+ Without Fake Precision
- A+: exceeds the frozen target band, passes integrity and risk gates, and has relevant untouched or forward support.
- A: meets the target band with no material integrity or safety failure.
- B: broadly acceptable but one bounded weakness needs a named corrective test.
- C: inconclusive or marginal; keep exposure constrained while evidence grows.
- D: material process, evidence, or risk weakness; do not scale.
- F: the evidence is unusable, the process is not the declared strategy, or a non-negotiable risk gate failed.
Plus and minus modifiers should have written definitions or be omitted. Do not infer them from small decimal differences. “Not enough evidence” is often more honest than C.
Dimension 1: Data Integrity
Reconcile the journal to broker or account records. Check duplicate and missing trades, time zones, currency conversion, partial fills, deposits and withdrawals, fees, account IDs, setup versions, and edits made after outcomes were known.
A dataset can be profitable and still deserve an F for integrity. Missing losers, merged strategies, or outcome-edited tags invalidate later conclusions. Preserve a coverage ratio and a list of unresolved records beside the grade.
Dimension 2: Rule Execution
Grade whether the trader performed the strategy that was supposedly tested. Measure signal eligibility, entry and exit compliance, planned versus actual risk, unplanned trades, overrides, and whether reasons were recorded contemporaneously.
Do not award execution points because a violation won. A compliant loss is valid system evidence; a profitable violation is a process exception. Mixing them teaches the rubric to reward luck.
Dimension 3: Net Edge Evidence
Evaluate net expectancy, average win and loss, win rate, distribution shape, costs, dependence, regime coverage, and stability across time. Compare the strategy with its frozen hypothesis and control, not with a universal profit factor or win-rate table.
The edge-validation guide covers development versus untouched evidence. Grade evidence strength separately from the current point estimate: an attractive result with a small, dependent, or mined sample can remain C for confidence.
Dimension 4: Risk Containment
Measure maximum and rolling drawdown, loss clusters, gap and slippage paths, concentration by trade/day/market, simultaneous exposure, and utilization of any daily or maximum-loss boundary. The worst credible path matters more than a smooth average.
The drawdown recovery analysis helps separate recovery arithmetic from claims that a particular curve shape is safe. A risk breach should cap the overall grade even when one outlier keeps total P/L positive.
Calculate the Composite With Safety Caps
If numeric weights are used, publish them. A simple approach is equal weights for the four dimensions, followed by explicit caps:
- unreconciled data prevents a deployable composite;
- a hard risk breach prevents an A or B;
- no untouched or forward evidence caps edge confidence;
- material execution drift prevents the outcome grade from representing the strategy.
These are example governance choices, not universal scoring law. Teams can choose another rubric if it is written before results and applied consistently.
Worked Example: One Result, Four Different Grades
Imagine a profitable evaluation window. The journal reconciles every source fill and fee, so data integrity meets its A band. Most eligible signals followed the frozen plan, but several exits were changed without a contemporaneous reason, so execution is B. Net expectancy is positive, yet the untouched window is short and several trades share one market driver, so edge evidence remains C. One oversized position exceeded the written portfolio cap, so risk is F.
The composite must not average that path into B and call the strategy healthy. Under a predeclared safety cap, the risk failure controls the overall verdict. The next action is not to improve win rate; it is to prevent the position-sizing path from recurring, replay the affected sequence, and collect new compliant evidence.
This example is illustrative, not a benchmark. Change the letters only through the rubric that existed before the data was graded.
Make Missing Evidence Visible
A report card should distinguish “failed,” “passed,” and “not assessable.” If planned risk, setup eligibility, or costs are absent, the corresponding component cannot be inferred from P/L. Show field coverage, unresolved mappings, and excluded records beside the grade.
Do not reward incomplete evidence with a neutral midpoint. Depending on the decision, missingness may suppress the composite or cap it until the required records are reconciled. That policy belongs in the rubric.
Test Whether the Grade Is Portable
A strong grade on one account, instrument, or market regime may not transfer. Re-evaluate the frozen strategy under the new execution route, cost structure, leverage, contract rules, and liquidity state before scaling. Preserve the original grade as a historical result rather than rewriting it with the new environment.
Portability is evidence about repeated process under changed conditions. It is not established by copying a position size or assuming that the same symbol creates the same risk.
Grade the Trend Without Moving the Goalposts
Store every rubric version and compare like windows. A rising grade can come from better execution, easier market conditions, missing data, or a relaxed threshold. Show component changes and the underlying metric deltas so the reason is visible.
Use rolling and fixed-period views together. Rolling windows show recent change but overlap heavily; calendar windows are easier to audit but can be noisy. Never stitch the best parts of both into a retrospective success story.
Annotate strategy releases, broker migrations, fee changes, extended outages, and material account-rule changes on the timeline. A discontinuity after one of those events belongs to a new comparison context, even if the dashboard can draw one continuous line.
Turn the Weakest Grade Into One Test
- Open the records behind the weakest verified component.
- Separate data defects, strategy defects, execution defects, and risk-policy defects.
- Choose one controllable cause with direct evidence.
- Write a narrow intervention and the metric it should change.
- Freeze the rest of the system for the evaluation window.
- Re-grade with the same rubric; retain, revise, or roll back the intervention.
The performance-analysis workflow supplies the drill-down sequence. Do not promise one letter per month or quarter; improvement cadence depends on trade frequency, variance, and the intervention.
How TSB Builds an Auditable Report Card
Trader’s Second Brain can keep imported fills, fees, accounts, strategy and setup versions, risk, compliance tags, notes, and outcomes in one journal. Reports and Backtester expose the rows behind a grade instead of presenting the letter as an oracle.
Coach can summarize the selected evidence, find which component changed, and ask whether missing data or a rubric revision explains the result. That is the strong use: rapidly turning a large journal into traceable diagnostic questions. It should qualify or refuse a grade when required fields, reconciliation, or sample support are missing.
TSB recognizes 331 exact import profiles and has normalized 600K+ imported trades. These figures describe import coverage and imported trade volume—not users, the sample behind any grade, proof of edge, or promised improvement.
TSB is our product. We disclose that ownership because this guide recommends its journal, Reports, Backtester, and Coach workflow.
Methodology Note
- Protected intent: the F-to-A+ report-card interface remains, with transparent strategy-relative definitions.
- Removed claims: universal metric cutoffs, percentages of traders at grades, fixed improvement cadence, personality labels, and deterministic fixes were not retained.
- Evidence boundary: grade components summarize current evidence; they are not forecasts or trader credentials.
- Schema boundary: this visible rubric is editorial content, not artificial Review, Rating, or Product structured data.
For our evidence and correction process, see the editorial methodology.
Final Verdict: A Grade Must Be Traceable
Keep the letter only if every reader can open the rubric, source metrics, and exceptions behind it. Grade data, execution, edge evidence, and risk separately; let integrity and safety cap the composite.
The report card succeeds when it points to one bounded next test. It fails when A+ becomes a confidence badge detached from the records.