Historical validation
Backtesting
How TrydingDay tests market signals against a growing historical record — and what those tests can and cannot establish.
What backtesting does
From signals to evidence
Backtesting asks whether a signal retains useful information when it is applied to many historical situations and tested under rules fixed before the outcome is known. It does not turn an isolated indicator into an automatic instruction. It helps distinguish repeatable patterns from coincidences, quantify uncertainty and decide what information can usefully support the research.
From early tests to a broad historical base
Evidence that grew with the research
The first tests were deliberately narrow. One of them compared 75 observations from the daily scan corpus available at the time, including candidates, alternatives and excluded companies. It explored different volume bands and generated the next questions for the research.
That first sample was a starting point, not the complete archive and not a manual selection of preferred cases. The research later expanded through the capture and consolidation of a much larger historical record. The consolidated episode catalogue reached 49,794 episodes, while the factual dataset used for machine learning approached 80,000 observations. These are different data layers, but together they show the move from an initial exploration to a broader historical evaluation.
Indicators and signals tested
Different inputs, specific questions
The work tested trend and relative-position signals, including moving averages, RSI and distance from recent references. It also examined volume, liquidity, company size, volatility, ATR and the context of a reference market.
Further analyses considered events, sentiment, insider activity, short interest and price-compression patterns. Each block was studied with a concrete question: whether it helped explain what happened afterwards, improved the description of adverse risk or added information beyond a simple historical reference.
Validation stages
A progressively stricter review
The process began with exploratory comparisons to identify plausible relationships between indicators and later outcomes. Promising hypotheses were then tested against a wider historical record, with observations at 1, 5, 10 and up to 15 sessions.
Temporal evaluations followed: each period was assessed using only the information that would have been available at that time. Signals were compared with simple references, while the research also examined costs, next-open execution, favourable and adverse paths, incomplete data, price adjustments and robustness across different samples.
An external audit reviewed twelve families of tests on the historical engine. No variant showed a sufficiently robust advantage to become a new selection or execution criterion. That conclusion is part of the value of backtesting: it keeps apparent improvements out of the process when they do not withstand independent review.
How the approach changed
Why the score and ranking were removed
In the early workflow, the engine calculated a score and produced a ranking to order the signals. A bounded selection was then made and the Desk deliberated over the companies included in that review.
The first backtesting cycle showed that the score was not a sufficiently reliable basis for discriminating between alternatives. The response was not to refine it or retain the ranking as an auxiliary ordering: the score and ranking were removed from the workflow.
Since then, research prepares the candidate set for the Desk without an automatic classification determining which company should prevail. The Desk reviews the sources and available evidence, compares alternatives and deliberates before making the editorial selection.
Machine learning
Quantitative context from machine learning
The machine-learning model is a separate quantitative layer trained on factual historical observations. It does not use editorial decisions or narrative outcomes as shortcuts. Its role is to provide context across different horizons, approximately 3, 8 and 15 sessions, rather than to replace Desk selection or decide an operation automatically.
The temporal evaluation covered 495 days and more than 1.1 million publication-to-session examples. The results indicate that the model is more useful for describing a range of favourable and adverse scenarios than for predicting one central outcome. Its lower and upper estimates, P10 and P90, improved against simple historical references and increased central coverage by between 1.09 and 2.45 percentage points. The central P50 estimate and trajectory persistence did not show a consistent improvement.
What this means for TrydingDay
Useful evidence, not an automatic promise
Backtesting does not turn a signal into a permanent truth. It helps identify which indicators deserve attention, which hypotheses need more evidence and which ideas should remain outside the decision process.
Historical results are therefore used as context. They help calibrate expectations, describe possible scenarios and recognise risk, but they do not guarantee a future outcome. The factual learning base is not frozen: as new trajectories mature, it expands every day and the model is updated and retrained against that growing record.
Continue through the record