Decision Rules for Promote, Revise, or Reject

Direct Answer

A strategy's fate, whether it gets deployed, revised, or archived, should be governed by criteria written before the results arrive, not negotiated after the fact. Pre-specified decision rules prevent the most pervasive form of bias in strategy lifecycle management: the tendency to find reasons to keep developing a strategy when results are disappointing, or to launch a strategy with insufficient evidence because it looks promising and the researcher wants to move forward.

Three outcomes cover every strategy evaluation: promote (the evidence meets the pre-specified standard for live deployment), revise (the evidence identifies a specific, theoretically motivated improvement worth testing in a new hypothesis), or reject (the strategy idea has been sufficiently tested and does not demonstrate a real edge). Only promote requires the full evidence standard. Revise requires that the new hypothesis be logged before retesting. Reject requires that the strategy be archived, not deleted, so its contribution to the multiple-testing count is preserved.

Key Takeaways

  • Decision criteria must precede results: Writing promote/revise/reject criteria after seeing the results allows the criteria to be fitted to the result, which defeats the purpose of having criteria.
  • Promote requires meeting all pre-specified criteria: Not "most" of them, and not criteria that were adjusted to fit the result. The promote decision is only meaningful if the criteria were strict enough to exclude borderline results.
  • Revise is not a way to keep testing the same idea indefinitely: Each revision is a new hypothesis, and the total number of hypothesis tests in a research line contributes to the multiple-testing count. Revising without discipline collapses into data mining.
  • Reject means archive, not delete: Archived rejected strategies document the search breadth and prevent rediscovery of the same dead end. A deleted rejection is a phantom hypothesis that inflates the apparent pass rate of surviving strategies.
  • Live deployment requires a retirement rule too: The decision criteria are not only for the promote moment, they should include a rule specifying what live performance pattern would trigger suspension and review, set before deployment begins.
  • The mechanism must survive, not just the number: A strategy that passed a significance test but whose mechanism cannot be explained should be held at a higher evidentiary standard before promotion. Numbers can be flukes; a plausible mechanism increases the prior probability of a genuine effect.
  • First deployment should be at reduced scale: Regardless of how rigorous the research was, first live deployment at a fraction of intended scale limits the capital at risk while real live performance data is collected. The initial live period is additional evidence, not confirmation of the backtest.
  • The promote decision is documented, not asserted: The research log should contain a formal promote decision entry that references the pre-specified criteria, the evidence on which each criterion was evaluated, and the conclusion. A promote decision that exists only in the researcher's memory is not documented.

Core Concepts

Pre-specified decision criteria for promotion

The promote criteria should be written as part of the research protocol, in the pre-test log entry, before any backtest is run. They specify the conditions that must all be satisfied before the strategy can proceed to live trading. A complete set of promote criteria typically includes: (1) the primary metric threshold (Sharpe above X, t-statistic above Y, information ratio above Z); (2) the required robustness checks and their pass conditions; (3) a minimum number of trades in the sample (to ensure statistical power); (4) a maximum drawdown condition (the in-sample maximum drawdown must be survivable with the intended position sizing); (5) a cost assumption that must be realistic relative to the live execution environment; and (6) a qualitative requirement that the strategy has a credible, theoretically motivated mechanism.

The qualitative mechanism requirement is important and often overlooked. A purely data-driven strategy that passes statistical tests but has no explanation for why it should work deserves more skepticism than one grounded in a persistent structural or behavioral phenomenon. Momentum strategies have been explained by investor underreaction and gradual price discovery. Mean-reversion strategies have been explained by overreaction, liquidity provision premiums, and microstructure effects. A strategy that produces a positive Sharpe with no mechanism explanation might be fitting to a historical accident, a one-time liquidity episode, a regulatory regime that has since changed, or a data artifact.

The criteria must be all-or-nothing: a strategy that passes five of six criteria and the researcher wants to deploy anyway has not met the promote standard. If the criteria are genuinely too strict, if they would reject every reasonable strategy, they were written incorrectly and should be revised in a future protocol, not relaxed on the current strategy because the researcher wants to proceed.

At the promote moment, the researcher writes a formal promote decision entry in the research log. This entry lists each criterion, the evidence against it, and an explicit pass/fail notation. The entry is then signed off with a timestamp. This creates a document that can be reviewed later, when the strategy is performing live, to verify that the decision was based on pre-specified standards and not on post-hoc rationalization.

Revise: when and how to redesign

A revise decision is appropriate when a strategy fails the primary hypothesis but the failure reveals a specific, actionable improvement that has a theoretical motivation independent of the results. The key test is counterfactual: would the revision have been proposed if the original test had produced different results? If the answer is no, if the revision was designed to address the specific failure observed. It is a data-driven revision, not a theoretically motivated one, and it must be tested on fresh data.

The revision procedure: log the original hypothesis as failed with a complete post-test entry. Document the insight the failure revealed and the theoretical motivation for the proposed revision. Write a new pre-test entry for the revised hypothesis, with a new timestamp, before running any analysis. If possible, test the revision on data not used in the original test (a different time period, a different market). If the same data must be used, apply a multiple-testing correction that accounts for the number of hypotheses tested including the original.

Common legitimate reasons to revise: the original hypothesis tested a broad version of a signal and the failure revealed a specific subset where the mechanism should be stronger; a data quality audit revealed a look-ahead bias that, once corrected, changes the signal's characteristics; new theoretical work suggests a different signal computation is more theoretically motivated than the original. Common illegitimate reasons: the signal worked in one decade but not another and a regime filter is being added to exclude the bad decade; the Sharpe was 0.55 when 0.6 was required and the threshold is being lowered; the transaction cost assumption was too high and it is being reduced post hoc.

The total number of hypothesis tests in a research line, the original plus all revisions, must be tracked in the research log. If the cumulative count grows large relative to the number of passing results, the Bonferroni-adjusted significance threshold for the line must be applied. A research line that has produced 30 failed hypotheses and 1 passed one is in a different evidential position than a research line that produced 3 failed hypotheses and 1 passed one, even if the passing result's t-statistic is identical in both cases.

Reject: archiving as discipline

A reject decision is the conclusion that a strategy idea has been sufficiently tested and the evidence does not support deployment. "Sufficiently tested" requires that the tests were conducted with scientific discipline, pre-specified hypotheses, clean data, appropriate cost assumptions, and multiple-testing corrections, not just that many tests were run and all failed.

Archiving a rejected strategy is as important as logging the original hypothesis. The archive entry should contain: the original hypothesis and all subsequent revisions, the results of all tests, the reasons for rejection, and a date. It should also contain any insights from the failure that might be useful for related research, for example, "the signal appears to work during regime X but not universally" is a valid insight from a failed universal test, even though it cannot be acted on without a new hypothesis and test.

The file drawer problem, the tendency for negative results to disappear without being documented, directly inflates the apparent quality of the surviving strategies. If 20 strategy ideas are tested in a year, 1 is deployed, and 19 are not documented, the 20-hypothesis count that should be used for multiple-testing correction is invisible. Future reviewers of the portfolio see one successful strategy and have no way to know how many failures preceded it. Full archive discipline is the only way to maintain an honest accounting of the research process's overall record.

Live deployment with a retirement rule

The decision criteria do not end at the promote moment. Every live strategy should be deployed with a pre-specified retirement rule, the conditions under which the strategy will be suspended and reviewed before continuing. The retirement rule prevents the common failure mode of allowing a live strategy to continue underperforming indefinitely because the researcher is reluctant to admit it has failed.

A retirement rule might look like: the strategy is suspended and reviewed if (a) the rolling 12-month Sharpe ratio falls below -0.2 for two consecutive months, (b) the drawdown from the live inception high exceeds 1.5× the maximum drawdown in the backtest period, or (c) 6 months of live performance produces a cumulative return below the risk-free rate with no evidence that the drawdown is within the expected range of backtest drawdown distribution. The specific thresholds should be derived from the backtest distribution, if the strategy historically experienced drawdowns up to 15%, a 25% drawdown rule is reasonable. If the backtest drawdowns were never above 5%, a 25% rule is too generous.

Initial live deployment at reduced scale (10-25% of intended allocation) is also good practice regardless of the promote criteria's rigor. The first 6-12 months of live trading provides real evidence of live performance that the backtest cannot provide, actual execution costs, actual market impact, and the strategy's behavior in the genuinely unknown future. This live track record, at reduced scale, informs the decision to deploy at full scale or retire the strategy before significant capital is at risk.

Worked Scenario

A researcher's pre-specified promote criteria for a momentum strategy are: (1) Sharpe ≥ 0.7, (2) t-statistic of alpha vs benchmark ≥ 2.0, (3) maximum drawdown ≤ 25%, (4) profitable in at least 2 of 3 sub-periods, (5) Sharpe ≥ 0.4 at 2× baseline costs, (6) a theoretically motivated mechanism documented before the test.

Bitcoin and investment strategy visualized with smartphone, tablet, and coins.
Photo by Leeloo The First via Pexels
  1. Result: Sharpe = 0.82 (pass), t-stat = 2.4 (pass), max drawdown = 22% (pass), sub-period profitability = 3/3 (pass), Sharpe at 2× costs = 0.55 (pass), mechanism: well-documented momentum premium with behavioral explanation (pass).
  2. Decision: All 6 criteria met. Promote decision entry written in research log with each criterion's evidence and an explicit pass notation. Timestamp: 2026-08-07.
  3. Live deployment: Deployed at 20% of intended allocation. Retirement rule: suspend if rolling 6-month Sharpe drops below -0.2, or drawdown from live high exceeds 33% (1.5× backtest maximum).
  4. Live review (6 months later): Live Sharpe = 0.45 (below backtest but positive). Drawdown = 11% (within backtest range). Decision: no retirement trigger met; scale up to 50% of intended allocation per the scale-up rule documented at promotion.

The pre-specified criteria meant no post-hoc justification was possible at the promote stage, and the pre-specified retirement rule meant the live period's modest underperformance did not trigger a premature termination or an indefinite continuation without a decision point.

Measurement Framework

Decision criterionQuestion it answers
Primary metric vs pre-specified thresholdDoes the core evidence meet the standard set before the test?
Number of robustness checks passed / total specifiedDid the strategy pass the diagnostic tests, or only the headline metric?
Total hypothesis count in this research line (for multiple-testing adjustment)How many tests preceded this result, and what is the adjusted significance threshold?
Mechanism documented in pre-test entry (yes/no)Was there a theoretical reason for the strategy before the results confirmed it?
Live performance vs retirement rule thresholdsHas the live strategy triggered the conditions for suspension and review?
Archive entries for rejected strategies in same research lineIs the search breadth being honestly documented?

Common Failure Modes

Lowering the bar for promotion because the strategy "looks promising"

A strategy produces a Sharpe of 0.55 when the pre-specified threshold was 0.7. The researcher believes the strategy has potential and promotes it anyway, with a note that "the Sharpe was close to target and the other metrics look strong." This is a post-hoc revision of the promote standard. The 0.7 threshold was chosen because a Sharpe of 0.55 was not considered sufficient evidence for promotion. If the threshold was wrong, it should be updated in the next protocol, not lowered ad hoc to let the current strategy pass.

The cost of adhering to the standard is not deploying a strategy that might have worked. The cost of not adhering is deploying strategies with insufficient evidence, which is the root cause of most live trading disappointments. A strategy deployed with a 0.55 in-sample Sharpe under the original 0.7 standard will live-trade at some fraction of the in-sample Sharpe. If the degradation is 30-50%, the live Sharpe is 0.28-0.39, not a viable strategy.

Revising the strategy every time it underperforms a criterion

A strategy fails the sub-period stability criterion (profitable in 2/3 sub-periods; the strategy was profitable in only 1). The researcher adds a regime filter, tests again, the strategy now passes the criterion. A new criterion is introduced for cost sensitivity; the strategy barely passes. Another revision is made. After five iterations, every criterion is met, but the multiple-testing count is now five times the original plus all the revisions. The final result appears rigorous but is the output of an extensive search.

The fix is to count every revision as a new hypothesis test and apply the appropriate multiple-testing correction. If the research line has produced five revised hypotheses and one pass, the Bonferroni threshold for that pass (N=6) is approximately t > 2.4 rather than t > 1.96. If the final result barely clears 2.0, it fails the adjusted standard even though it looks significant on the unadjusted one.

Deleting or losing track of rejected strategies

A researcher spends six months testing 25 momentum variants. Twenty-four fail. The twenty-fifth passes, and the researcher clears out the failing research directories to declutter their workspace. A year later, an allocator asks how the strategy was developed. The researcher describes the one passing test but cannot reconstruct how many were tried before it. The allocator cannot apply the appropriate skepticism because the search breadth is no longer documented.

Research archives should be treated as permanent records. A simple convention: all research directories are named with the strategy ID and a status tag (e.g., "momentum-v12-rejected-2025-11" or "momentum-v18-archived-2026-02"), and none are deleted. The disk space cost of retaining research directories is trivial. The cost of losing the record of negative results is a permanently overstated hit rate.

No retirement rule leading to indefinite underperformance

A strategy promoted six months ago is now in its fourth consecutive month of negative returns, with a year-to-date loss of 8%. The researcher continues running it because "momentum sometimes has rough patches" and "the backtest showed this is normal." Without a pre-specified retirement rule, there is no trigger for a structured review. The strategy continues to lose until the researcher's subjective discomfort reaches a threshold, which may be much higher than the evidence-based threshold for suspension would have been.

A pre-specified retirement rule removes the subjective judgment from the suspension decision. If the rule is triggered, the strategy is suspended and a structured review determines whether to resume, revise, or retire permanently. If the rule is not triggered, the researcher has evidence that the current drawdown is within the expected range and can continue with appropriate confidence. The rule converts an emotional decision into a documented, systematic one.

Write the Rejection Criteria First

The criteria that need writing down first are the ones for rejection, because those are the ones that come under pressure later. Specifying what a good result looks like is easy. Accepting, after months of work, a rule saying the work should be shelved is not, and committing to it in advance is the only version of this that survives contact with a disappointing outcome.

Close-up of a business professional analyzing data trends using a tablet and laptop.
Photo by George Morina via Pexels

The rule also has to name a decision rather than a sentiment. Which measure, over what period, at what level, produces which of the three outcomes? A criterion that cannot be evaluated mechanically is a criterion that will end up being negotiated.

The middle option quietly absorbs everything if it is left open. Revision is a legitimate outcome and it is also the natural home for results that should have been rejected, so it needs a limit: how many revisions are permitted, and what makes the next one a rejection instead.

A decision rule governs process, not correctness. Following one carefully still permits deploying a strategy that fails, and its value is that the failure ends up informative rather than confusing.

Frequently Asked Questions

What is a promote decision in strategy research?

A promote decision is the formal conclusion that a strategy has met the pre-specified criteria for live deployment: it passed the primary metric threshold, passed the specified robustness checks, has a mechanism explanation that makes economic sense, and has no disqualifying red flags in data quality or the research process. The promote decision should be written in the research log alongside the evidence on which it was based, not as a summary judgment but as a traceable conclusion from pre-specified criteria.

What is a revise decision and when is it appropriate?

A revise decision is appropriate when the primary hypothesis failed, but the failure reveals a specific, theoretically motivated improvement that was not simply the result of looking at the data. The revision must be logged as a new hypothesis entry with a new timestamp, and the new test should use data that was not used to design the revision if possible. A revise decision that simply changes parameters to get a better result, without a substantive theoretical reason for the new parameters, is more honestly classified as a continuation of the search, not a principled revision.

What is a reject decision and should rejected strategies be archived?

A reject decision is the conclusion that a strategy idea has been sufficiently tested and does not demonstrate a real edge under the conditions specified. Rejected strategies should be archived, not deleted, for three reasons: they contribute to the denominator for multiple-testing corrections on future related work; they prevent the same idea from being re-investigated without memory of the prior failure; and the failure sometimes reveals genuine information about what does not work, which is valuable research even if not actionable directly.

How do I prevent post-hoc justification of a promote decision?

Write the promote criteria in the research protocol before the backtest runs. The criteria should include specific quantitative thresholds for the primary metric, the required robustness checks, any disqualifying red flags, and the minimum out-of-sample evidence required. A promote decision made by consulting these pre-written criteria is structurally protected from post-hoc justification; a promote decision made by looking at the results and deciding they seem good enough is not.

What evidence is required before promoting a strategy to live trading?

At minimum: (1) the primary hypothesis passed on pre-specified metrics with the pre-specified threshold; (2) all pre-specified robustness checks passed or were documented with their outcome; (3) data quality was verified (point-in-time data, survivorship-free universe, appropriate lag); (4) the result survived a multiple-testing adjustment for the number of variants tested; (5) there is a credible, theoretically motivated mechanism explaining why the strategy should work, not just a statistical pattern; and (6) costs in the live implementation will not exceed the cost assumption used in the backtest.

What should a strategy retirement rule look like?

A retirement rule specifies the conditions under which a live strategy will be stopped and archived, as distinct from the conditions under which its parameters will be revised. A reasonable retirement rule might be: if the live strategy's rolling 12-month Sharpe ratio falls below -0.3, or if the maximum drawdown since live inception exceeds twice the maximum drawdown in the backtest period, the strategy is suspended and reviewed. Writing this before going live prevents the researcher from endlessly lowering the bar for acceptable performance to avoid admitting the strategy failed.

How do you distinguish a genuine strategy failure from bad luck?

Poor live performance can reflect either a genuine edge that is temporarily out of regime, or a strategy that was never robust to begin with. The distinction requires examining: whether the live period represents a regime that was rare in the backtest period; whether the strategy's signal continues to work directionally even if the returns are poor; and whether the research process met scientific standards. This is a judgment call, but it should be explicitly stated before underperformance begins to avoid motivated reasoning during the loss period.

Can a rejected strategy ever be revisited?

Yes, but the revisit must be logged as a new hypothesis entry with an explanation of why the conditions for revisiting have been met. Legitimate reasons to revisit include: new data covering a regime not present in the original test period; a new theoretical mechanism that predicts the strategy should work differently than the original specification; significant changes in market microstructure that affect the strategy's execution assumptions. Revisiting because "the parameters look different now" without a substantive reason is re-running a search on the same data, which adds to the multiple-testing count without generating new information.

Who should apply the decision rule when the researcher and the decision-maker are the same person?

Separating the roles in time is the available substitute for separating them between people. Writing the criteria on one date, running the test on another, and evaluating against the written version without editing it imposes some of the distance a reviewer would provide. Sharing the pre-written criteria with someone outside the project before results arrive adds more, because it makes a later revision visible to a second party rather than being a private adjustment.

References

  • Lopez de Prado, M. (2018). Advances in Financial Machine Learning. Wiley. Chapter 14 covers strategy lifecycle management and the backtesting-to-production transition.
  • Harvey, C.R. & Liu, Y. (2014). "Evaluating Trading Strategies." Journal of Portfolio Management, 40(5), 108-118. Discusses promotion criteria and the minimum evidence standard for strategy deployment. Available at jpm.pm-research.com.
  • Aronson, D. (2007). Evidence-Based Technical Analysis. Wiley. Part III covers hypothesis testing standards for technical trading rules and the criteria for rule acceptance vs rejection.
  • Pardo, R. (2008). The Evaluation and Optimization of Trading Strategies (2nd ed.). Wiley. Chapter 11 covers the decision to deploy a strategy and the parameters of the initial live deployment period.
  • Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux. Chapters 19-24 discuss the systematic biases that affect research-to-deployment decisions, including the planning fallacy and optimism bias, the behavioral reasons pre-committed decision rules are necessary.

Educational Disclaimer

This guide is for educational purposes only and does not constitute investment, financial, or trading advice. All examples are illustrative. Live trading involves significant risk of loss including the possible loss of all principal. Consult a qualified financial professional before making investment decisions.