A trading edge is positive expected net value for a precisely defined, repeatable decision process. Your trade history can estimate that value; it cannot certify it from a magic trade count or one profit-factor threshold. The useful three-metric test is: positive net expectancy, uncertainty around that estimate, and stability outside the sample used to discover the rule.

Win rate, average winner versus average loser, and profit factor describe the same realized outcomes from different angles. They help explain the estimate, but they are not three independent proofs. A credible edge survives costs, outliers, rule changes, clustered trades, and a forward test whose rules were frozen before the result was known.

Quick verdict: you have evidence of an edge—not certainty—when a defined strategy has positive net expectancy, the plausible range is narrow enough for the decision, and comparable unseen trades do not collapse the result. If any one of those gates fails, keep the conclusion “not yet verified.”

The three gates: (1) economics—is average outcome positive after all trading costs? (2) uncertainty—could the observed mean plausibly be noise under a defensible dependence model? (3) robustness—does the result survive outliers, neighboring periods, and data not used to invent the strategy?

What “Edge” Actually Means in Trading

Define the population before the formula: exact setup version, instruments, session, market condition, risk policy, execution rules, costs, and eligible dates. “My breakout setup on liquid index futures during the first two hours, version 3, executed under the written stop rule” is testable. “I trade price action” is not.

The Mathematical Definition

For a simplified two-outcome model with no breakeven trades, expected result in risk units is:

E[R] = p(win) × average net win R − p(loss) × average net loss R

For real trade data, calculate the sample mean directly:

Estimated net expectancy = (R1 + R2 + … + Rn) ÷ n

Each R result should include commissions, fees, financing, and measurable execution costs. Planned risk at entry must be known; inferring it from the final loss after a moved stop changes the denominator.

Observed Edge Is an Estimate, Not a Property Stamp

A positive sample mean says this sample was positive. It does not show that the true expectation is positive, that the rules remained stable, or that the next sample will match it. Edge language becomes useful only when the estimate is paired with uncertainty and a clear validation history.

The Three Metrics That Test Your Edge

Metric 1: Net Expectancy

Net expectancy answers the economic question: what was the average result per eligible trade after costs? Report it in R and account currency. R supports comparison across position sizes; currency shows whether the effect matters after real friction.

Win rateAverage net winAverage net lossExpectancyProfit factor
30%3.00R1.00R+0.20R1.29
50%1.20R1.00R+0.10R1.20
70%0.55R1.00R+0.09R1.28
90%0.10R1.00R−0.01R0.90

The table is arithmetic, not a recommendation. It shows why win rate cannot prove edge and why reward-to-risk must be based on realized net outcomes. See the expectancy formula guide for breakeven geometry and cost treatment.

Metric 2: Uncertainty Around Expectancy

Report more than the mean: eligible trade count, standard deviation or interquartile range, median R, largest winner and loser, share of net P/L from the largest trade, and an interval around the mean. The interval must respect how the trades were generated. Trades from the same session or regime are not necessarily independent.

As a deliberately simplified diagnostic, suppose 100 independent trades show mean +0.12R and sample standard deviation 1.30R. The estimated standard error is 1.30 ÷ √100 = 0.13R. A normal-approximation 95% interval is roughly +0.12R ± 0.25R, or −0.13R to +0.37R. The profitable sample is real, but the sign of the underlying mean is not resolved by that calculation.

That formula assumes independent, similarly distributed observations and ignores strategy-selection history. When trades cluster, use a day/session block bootstrap or another dependence-aware method. The sample-size guide shows why required evidence depends on variance and the precision you need—not a universal 60- or 200-trade badge.

Metric 3: Out-of-Sample Stability

Split discovery from confirmation. Use the first period to form the rule, freeze it, and evaluate later eligible trades without changing setup definitions, exits, filters, or the primary metric. Then inspect rolling periods, outlier removal, costs, and relevant regimes. A result that survives is stronger evidence; it still is not a guarantee.

Record every material variant tried. Selecting the best of many rules inflates apparent performance even when each calculation is correct. Bailey and López de Prado’s deflated Sharpe work addresses selection bias, backtest overfitting, and non-normal returns; the practical lesson is simple: trial history belongs beside the winning result.

Where Win Rate, Reward-to-Risk, and Profit Factor Fit

Win Rate

Win rate is the share of eligible trades with positive net outcomes. It is useful for understanding outcome frequency and losing-streak exposure. It says nothing about payoff magnitude by itself. Include breakeven handling explicitly because moving scratch trades between wins and losses changes the percentage.

Realized Reward-to-Risk

Average net winner divided by average absolute net loser is often called payoff ratio. It differs from the planned target-to-stop ratio: partial exits, slippage, stop changes, and gaps alter realized outcomes. Review both, but do not substitute the plan for execution evidence.

Profit Factor

Profit factor equals gross net winning P/L divided by the absolute gross net losing P/L for the same eligible sample. A value above 1 means wins exceeded losses in that sample. There is no universal level—1.2, 1.5, or 2.0—that independently proves a durable edge. Profit factor is derived from the same winners and losers that determine expectancy and can be dominated by one outlier. Use profit-factor benchmarks as descriptive context, not a certification ladder.

What “Edge Ratio” Measures—and Why It Is Different

Edge ratio is not a standardized synonym for trading edge. A common implementation compares maximum favorable excursion (MFE) with maximum adverse excursion (MAE), often after volatility normalization and over a fixed post-entry horizon. MetaTrader documentation defines MFE as the maximum favorable movement during a position and MAE as the maximum adverse movement.

An MFE/MAE ratio can evaluate entry-signal excursion, but it does not include the actual exit rule, costs, fill quality, position size, or the full outcome distribution. Always publish the exact formula, volatility denominator, horizon, and zero-MAE handling. Never mix an “edge ratio” from one platform with expectancy or profit factor from another without reconciling definitions.

The Sample Size Problem

Sample size has no meaning without dispersion and a target question. “Is there any positive expectancy?” may require much more data than “is the import missing half the commissions?” A small sample can expose operational errors; it usually cannot resolve a noisy mean.

Observed stateWhat it supportsWhat it does not supportNext action
Few trades, missing costsData-quality auditEdge estimateRepair import and reconcile
Positive mean, wide intervalPromising observationConfirmed signCollect frozen-rule sample
Positive result, one dominant tradeRealized gainTypical-trade edgeReport with/without outlier
Stable unseen periodsStronger repeatability evidenceFuture guaranteeMonitor drift and costs

Do not stop at a p-value or confidence label. Ask whether the estimated effect is economically large enough after costs and whether the test matches the live implementation.

The Hidden Deal-Breaker: Researcher Degrees of Freedom

The dangerous version of confirmation bias is mechanical: changing date windows, excluding trades, relabeling setups, moving exits, or searching sessions after seeing results. With enough choices, a flattering subset can emerge by chance.

Keep a research log with the original hypothesis, eligible universe, rule version, costs, primary metric, every variant tried, and the untouched confirmation period. Exploratory findings are valuable; label them “discovered here, not validated here.” The backtest-versus-live guide covers the execution and data differences that remain after statistical validation.

Five Things That Can Erase an Observed Edge

1. Costs and Execution Friction

Commission, spread, slippage, financing, rejected orders, partial fills, and latency belong in the same unit as outcomes. Gross-positive and net-negative are different strategies.

2. Rule-Version Drift

If entry, exit, filter, or risk rules change, do not pool all trades as one stable process. Version the strategy and show pooled results only as a portfolio history.

3. Post-Hoc Selection

The best session or setup found after searching is a new hypothesis. Confirm it later before narrowing the live playbook.

4. Outlier and Tail Dependence

One winner can create most of the profit; clustered losses can arrive together. Show concentration and sequence stress rather than assuming independent, bounded outcomes.

5. Backtest–Execution Mismatch

A backtest may assume fills or signals the trader cannot reproduce. Compare planned and realized entries, stops, exits, and costs on the exact live process before calling the gap psychology or decay.

How to Measure Your Edge Right Now (5 Steps)

  1. Freeze the object. Name the strategy version, eligibility, account, period, instruments, session, regime rule, and risk policy.
  2. Reconcile the data. Match closed trades and net P/L to source records; quantify missing costs, stops, and labels.
  3. Run the three gates. Calculate net expectancy, a dependence-aware uncertainty view, and out-of-sample or forward stability.
  4. Stress the result. Show profit factor, win rate, payoff ratio, median, largest-trade share, neighboring windows, and results with major outliers isolated.
  5. Write one next decision. Keep, narrow, pause, or gather more evidence under a predeclared rule. Do not increase risk merely because the point estimate is positive.

If you need to locate where the result came from before testing it, use the performance-attribution ledger; it prevents setup, regime, sizing, and instrument views from double-counting the same P/L.

How Trader's Second Brain Tests Edge Without Losing the Evidence Trail

Disclosure: Trader's Second Brain (TSB) is our product. We designed it to connect import reconciliation, performance analysis, review evidence, and the next rule instead of leaving the trader to stitch a broker export, spreadsheet, screenshots, and chat together.

TSB has processed 600K+ imported trades and recognizes 328 exact source profiles. Those numbers mean imported trades and exact recognized import routes—not users, universal compatibility, or trades analyzed by AI Coach. Every route still has to reconcile before an edge metric is trusted.

Once the evidence is clean, Reports can calculate the observed pattern; setup, session, instrument, fee, risk, and sequence views can expose concentration; Leak Map can rank supported negative comparisons; Retrospective Backtester can test a defined historical rule; and Current Focus can carry one evidence-backed instruction into the next session.

AI Coach is the intelligence layer that makes those surfaces compound. It can interrogate the selected evidence set, connect the relevant cohorts, challenge the most attractive alternative explanation, state exactly what the sample supports, and convert the result into one decisive next review. When evidence is missing, Coach identifies the missing field or denominator and routes to a valid fallback lens. It does not dilute a strong finding into generic caution, and it does not manufacture certainty the account never supplied.

Methodology Note

  • Economics: expectancy and profit factor are calculated from net eligible outcomes. Win rate and payoff ratio explain the arithmetic but do not independently validate it.
  • Inference: the quick interval example is explicitly IID and approximate. Real decisions should account for clustering, tails, missingness, and strategy selection.
  • Selection: Bailey and López de Prado show why selection bias, multiple trials, and non-normal returns can inflate reported performance.
  • Performance estimation: Andrew Lo demonstrates that serial correlation can materially alter Sharpe-ratio inference; it supports the broader warning against naive independence assumptions.
  • Edge ratio: MFE and MAE are path statistics; any ratio built from them requires an explicit formula and horizon and is not equivalent to net expectancy.

Primary references: Andrew Lo, The Statistics of Sharpe Ratios, Bailey and López de Prado, The Deflated Sharpe Ratio, MetaTrader 5 testing-report definitions for MFE and MAE, and CFA Institute performance-evaluation boundaries.

For our evidence controls, see TSB editorial methodology.

Final Verdict: Measure the Claim, Not the Feeling

A positive month is a result. A positive net mean is an estimate. An edge claim needs uncertainty and validation. The distinction protects good strategies from being abandoned after noise and weak strategies from being scaled after luck.

Use the three gates in order: economics, uncertainty, robustness. Keep win rate, payoff ratio, profit factor, edge ratio, and attribution in their proper roles. If the evidence is mixed, “not yet verified” is not failure—it is the exact state that tells you what to measure next.

EXPECTANCY · UNCERTAINTY · ROBUSTNESS

Put Your Edge Claim Through All Three Gates

Reconcile the trades, expose the real net estimate, and let AI Coach turn the strongest supported finding into one focused forward test.

Test the evidence in TSB