Keep some evidence out of the workshop

A researcher uses one period to explore, build and improve an idea. Before opening a separate test period, the rule is fixed. The test period is then used once to see how the finished rule behaves. Because those examples did not shape the method, they provide a cleaner challenge.

The boundary applies to human choices as well as computer training. If somebody repeatedly checks the holdout and adjusts the rule, the holdout has joined the workshop. It is no longer a fresh test.

Later is not automatically unseen

A chronological split is often sensible for markets because it respects time. But a later date is not enough. If the researcher had already inspected the entire archive before choosing the rule or split, the later observations may already have influenced the idea.

That later slice can still reveal instability and provide useful descriptive evidence. It should be labelled as an exposed replay rather than promoted to out-of-sample validation.

One clean test is not the finish line

A rule can pass unseen data by chance, especially when many rules were developed. The test period may also represent only one market environment. Robustness checks, independent replication and prospective records add different kinds of pressure.

Fresh-data performance should be judged with the same result measure, costs and exclusions fixed in advance. Changing the target after seeing the test creates a new version that needs its own check.