Two things to state up front
Technical analysis without quantified measurement cannot work durably. A chart contains no information until you have counted. A support level does not exist because it “looks solid”: it exists if, across a large number of comparable cases, price reacted there more often than chance would predict. Without that counting, you are not reading a market, you are reading a shape — and the human eye always finds a shape.
Fundamental analysis, for its part, gives no entry signal. It serves to estimate what a company is worth, not to decide whether to buy on Tuesday rather than Thursday. By the time a retail investor discovers a piece of public information — earnings, a change of management, a macroeconomic figure — the price has already moved. Professionals had it earlier, processed it faster, priced it in.
This is what Eugene Fama formalised in 1970 as the informational efficiency of markets: in its semi-strong form, all public information is already reflected in prices. What is public is no longer exploitable. The theory remains debated — the 2013 Nobel was shared between Fama, Lars Peter Hansen and Robert Shiller, the latter defending an opposite reading — but the practical conclusion does not change for a retail investor: reading a press release is not an edge.
That leaves a single serious path: measuring. This page presents the five tests every strategy must pass, and the five mistakes that consist of not passing them.
Count before deciding
Before opening a position, only one question matters: out of the many times this situation has already occurred, how many times did I win?
And you can only count what is defined in advance. “Price touched the support” cannot be counted: two people will not see the same support in the same place. Three things must therefore be written down before any observation:
- How I identify the situation — a rule a computer could apply, with no human judgement
- Where I exit if I am wrong — the stop
- Where I exit if I am right — the target
Without these three elements, no percentage means anything.
The textbook example
A classic situation: in an uptrend, price comes back to touch its 200-day moving average, then moves on.
200-day moving average: the average of the last 200 closes, recalculated each day. A line that follows the price while smoothing it.
The rule under test:
| Market | a broad equity index (fictional) |
|---|---|
| Period | twenty years |
| Situation | price has been above its 200-day average for at least 10 sessions, then comes back to touch it |
| Buy | at the close of the session of contact |
| Stop | 80 points below |
| Target | 160 points above — risking 1 to win 2 |
The result: the situation occurs 74 times in twenty years. The target is reached 26 times, i.e. 35%.
Risking 1 to win 2, you need to succeed on more than 33% of trades to break even. On paper, then, it is a winner:
- 26 wins × 160 points = 4,160 points
- 48 losses × 80 points = 3,840 points
- Gross result: +320 points
- Costs (74 round trips × 3 points): −222 points
- Net result: +98 points in twenty years
It is this figure — the net — that will be used throughout. Many would stop at the gross. That is where the mistakes begin.
The five tests that follow each ask a different question. A single failure is enough to discard a strategy. We run them all anyway: knowing why it fails is what makes it possible to build a better one.
Mistake #1 — Not comparing against buy-and-hold
Am I doing better than doing nothing?
Over the same period, our index goes from 3,800 to 7,400 points: +3,600 points, with no screen and no decision. Thirty-six times the strategy's net result.
But comparing two gross gains means nothing: buy-and-hold ties up the capital for twenty years, the strategy is only in a position a few days a year. They have to be brought to a common unit.
The most telling one to start with: what you earn, relative to the worst loss you had to endure along the way.
| Net gain | Worst drawdown endured | Gain per point of drawdown | |
|---|---|---|---|
| Buy-and-hold | +3,600 | −3,600 (a −60% crash) | 1.00 |
| The strategy | +98 | −600 | 0.16 |
The strategy's drawdown can be computed: its longest losing streak is seven consecutive trades, i.e. 7 × 80 = 560 points, plus costs — about 600 points.
Buy-and-hold recovers 1 point for every point of drawdown endured; the strategy, 16 cents. It is therefore beaten six times over once the comparison is made fair — and beaten by someone who did nothing.
This ratio is a simplification. Professionals use annualised measures such as the Calmar ratio, which account for duration. With twenty identical years on both sides, the raw comparison is enough here.
Takeaway: a backtest must always display, side by side, the strategy and buy-and-hold over the same period, scaled to the risk actually endured.
Mistake #2 — Not comparing against chance
Does my signal add anything?
The principle: keep everything except the signal. Same market, same period, same number of positions, same stop, same target. Only the entry dates are drawn at random. The operation is repeated thousands of times by computer, and you obtain what chance produces under the same conditions.
Beware of how you draw. If you draw any day across twenty years, you compare purchases made in an uptrend (the signal only appears there) with purchases made sometimes in the middle of a crash. These are not the same conditions, so it is not a valid test.
The right method: draw only among the days when the market was in the same state — above a rising moving average. You keep the context, and you replace only what you want to evaluate: the fact of touching the average on that precise day. One step further, you can also require the drawn days to be at a comparable distance from the average.
In our example, these draws produce on average +68 net points, and 4 out of 10 do better than the strategy's +98.
Conclusion: the 200-day moving average adds nothing. The 98 points do not come from the signal, but from the underlying trend the signal passively follows.
This is the test that eliminates the most strategies, because it isolates the only thing that matters: the signal's own contribution, once trend and luck have been removed.
Mistake #3 — Confusing the past with the future
Does it hold on data I have never looked at?
A backtest proves nothing if it runs on the data that was used to choose the settings. By trying enough combinations — a 180 average instead of 200, a stop at 70 instead of 80, discarding the crash year “which was unusual” — you always end up finding a fine result on the past.
Overfitting: tuning a strategy so finely to the past that it describes the past instead of describing the market. It is learning the answers to an exam instead of learning the subject.
The remedy is simple and non-negotiable: cut the data in two. A first part is used to design and tune. The second is set aside, never looked at, and serves only for a single test, at the end.
| Period | Occurrences | Success |
|---|---|---|
| First 12 years (design) | 45 | 44% |
| Last 8 (held out) | 29 | 21% |
At 21% with a 1-to-2 ratio, the strategy loses money. The 35% was only an average between a well-tuned period and a period where nothing works any more.
It is not necessarily a failure of method: markets change. Volatility, participants, the weight of algorithms — nothing is stable over twenty years. A rule validated on one decade can stop working in the next, all the faster because it is simple, hence known and exploited. That is the meaning of Andrew Lo's “adaptive markets” hypothesis: inefficiencies appear, get exploited, then disappear.
The practical consequence: a strategy put into service must come from day one with an abandonment criterion written in advance. Otherwise you keep it out of attachment, blaming every bad streak on bad luck.
Mistake #4 — Forgetting the costs
Is anything left once everything is paid?
The costs were already deducted above: 222 points, two thirds of the gross gain. This section explains why they weigh so much, and how they combine with a counter-intuitive mechanism.
They are paid on every trade, winning or losing. Between the spread, the commission and execution slippage, let us count 3 points per round trip. That is little.
Spread: the gap between the buying price and the selling price at the same instant. You pay it without seeing it go by.
Execution slippage: the difference between the intended price and the obtained price, frequent when the market moves fast.
Now the mechanism: the farther the target, the less often it is reached.
| Target | Risking 1 to win… | Winners | Gross | Net of costs |
|---|---|---|---|---|
| 80 points | 1 | 41 of 74 (55%) | +640 | +418 |
| 160 points | 2 | 26 of 74 (35%) | +320 | +98 |
| 240 points | 3 | 19 of 74 (26%) | +160 | −62 |
(Computed over the twenty years to show the mechanism. Reminder: test #3 has already disqualified the strategy.)
Two lessons. The “1-to-3” version, often presented as the most prudent, loses money. And the risk/reward ratio is not a dial you turn to become profitable: the success rate degrades in exactly the proportion that cancels the expected gain.
This is not a quirk of our example. It is the point on which several academic studies turned: apparently profitable technical rules stopped being so once costs were deducted (Fama & Blume, 1966).
Mistake #5 — Ignoring statistical biases
Is my result distinguishable from noise?
The sample that is too small. 74 occurrences in twenty years sounds like a lot. It is very little. Same problem as a poll: surveying 74 people gives a margin of error of about 11 points. Our 35% therefore only means: the true rate lies between 24% and 46%. At 46% the method is clearly winning; at 30% it loses 860 net points. We have not measured an edge, we have measured noise.
Multiple testing. Try 20 rules with no real edge: one of them will look remarkable by pure chance. Testing 40 levels on 20 assets and 3 timeframes makes 2,400 tests — dozens of false positives guaranteed in advance. The number of attempts made before keeping a rule is part of the result and must be disclosed.
Hindsight bias. On a chart already drawn, the supports that held leap to the eye; those that gave way are invisible. The right reflex: could the rule have been applied on that day, without knowing what came next? A trendline drawn after the fact by joining two well-chosen points fails the test.
Survivorship bias. Testing an equity strategy on the current composition of an index means testing it on the only companies that survived. Bankruptcies and delistings have vanished from the sample — and with them, the worst losses.
The verdict
| Test | Question | Result |
|---|---|---|
| 1. Buy-and-hold | Better than doing nothing? | No |
| 2. Chance | Does the signal add anything? | No |
| 3. Out of sample | Does it hold on fresh data? | No |
| 4. Costs | Is anything left after costs? | Almost nothing |
| 5. Sample | Distinguishable from noise? | No |
A strategy that looked like a winner passes none of the five tests. That is not a failure: it is a result, obtained before committing a single euro.
The temptation, here, is to tweak the settings until the tests pass. That is exactly the overfitting described above. The defensible behaviour is the opposite: discard the rule, and only come back to it with new data.
What this implies
It would be easy to read the above as “nothing works, give up”. It is the opposite.
Nothing here says that no strategy works. It says you only know after measuring, and that the overwhelming majority of what gets measured does not pass.
That is the job. Nobody finds the right idea on the first try. You test hundreds, discard nearly all of them, and sometimes something remains. The failures are not accidents along the way: they are the normal output of the work. Every discarded rule eliminates a hypothesis and narrows the search. A researcher with forty negative results has not wasted time — they have forty answers.
What separates this process from tinkering is the order of operations. You test before committing money, not after. Forty rules discarded on historical data cost nothing. A single rule tested live was paid for — dearly, for information that was available for free.
The market is not a training ground. It is where you apply work already done.
And you have to be lucid about proportions: it is not ten ideas, it is hundreds. Most people give up long before, which goes a fair way towards explaining the numbers in the next section.
Two factors that make everything worse
The five tests bear on the quality of a strategy. Two elements, independent of that quality, are enough to ruin an account even when the strategy is good.
Leverage. It lets you take a position ten, twenty or fifty times your capital, and multiplies gains and losses in the same proportion. A 2% move against you, with 30× leverage, wipes out 60% of the account. It is the central mechanism of the products covered by the AMF study cited below, and the reason those losses are so fast.
Add a structural point: on retail CFDs and Forex, the counterparty is often the broker itself. What the client loses, someone gains. Net of costs, the clients as a whole cannot win collectively.
Position size. A winning strategy applied with positions that are too large still leads to ruin, because losing streaks exist. Take a strategy that wins one trade in three — a common proportion: the probability of chaining five losses is about one in eight on each attempt. Over a hundred trades, it will happen several times. Whoever risks 20% of their capital per position is eliminated at that moment, whatever the quality of their signal.
Position sizing is therefore not an execution detail: it is part of the strategy, and is tested like everything else.
What the numbers show about retail investors
These demands may seem excessive. The data suggests the opposite. In France, the AMF published in 2014 a study of 14,799 retail clients active on Forex and CFDs over four years between 2009 and 2013, some 16 million transactions. The result: 89% losing clients, with an average loss of around €10,900 and a cumulative loss of about €161 million. More disturbing still: the most active clients are the ones who lose the most, which does not point towards learning by practice.
Among professionals, the picture is no better. The SPIVA study, published twice a year by S&P Dow Jones Indices, compares actively managed funds with their benchmark index. Over ten years, the overwhelming majority of European equity funds fail to beat theirs. In the United States, over fifteen years, not one of the 22 equity fund categories saw a majority of managers succeed.
One objection deserves to be addressed, because it is well founded: funds charge 1.5 to 2% in management fees per year, and a sizeable part of their lag comes from that rather than from incompetence. A retail investor does not pay those fees.
They pay others. Spread, commissions and execution slippage, multiplied by the number of round trips. A fund that trades little pays 1.5% a year; a retail trader taking two hundred positions a year can exceed that figure without noticing — exactly what test #4 showed. The disadvantage is not removed, it changes shape.
The nuance between trading and investing matters. These studies demonstrate the difficulty of beating the market. They do not say it is difficult to take part in it. Buying a broadly diversified index and holding it for decades has historically produced a positive return — that is the result of test #1. The difficulty is not in long-term, diversified, low-cost investing. It is in the claim to do better, faster, by trading often.
The tool is technical, the analysis is quantitative
The “gut feeling” approach is often opposed to the “mathematical” one, as if counting were opposed to observing. It is not: a study on historical data rests entirely on observation.
The real difference is between observing at random and observing methodically. Between remembering the times it worked, and counting all the times — including the ones you would rather forget. That is what modern medicine did: it did not replace clinical observation with equations, it organised observation so that it became countable and verifiable.
Hence the conclusion about vocabulary. The tools — moving averages, RSI, Fibonacci, Ichimoku — are technical: they belong to the mechanics of the chart. The analysis, for its part, is quantitative: counted occurrences, percentages, margins of error, costs.
Saying “I do technical analysis” suggests the tool is enough. Five questions must have a numbered answer before clicking:
- Am I doing better than doing nothing, at comparable risk?
- Am I doing better than random entries under the same conditions?
- Does it hold on a period I did not use to design the rule?
- Is anything left once costs are deducted?
- Is my sample large enough for this not to be noise?
And a sixth, which concerns not the strategy but the person applying it: how much can I lose in a row without being eliminated?
To go further
- Market efficiency (Fama, 1970): weak, semi-strong and strong forms, depending on the information already embedded in prices.
- Benchmark: the reference to beat. Buy-and-hold is one, random entries are another.
- Permutation test: replace the signal with chance, holding the context constant, and compare the distributions.
- Out-of-sample validation: set aside a period not used for design.
- Overfitting and multiple testing: the more you try, the more you find by chance.
- Confidence interval: for a proportion p over n trials, margin of error ≈
1.96 × √(p(1−p)/n). Wilson or Agresti-Coull intervals are more reliable on small samples. - Risk of ruin: the probability of exhausting your capital before the edge materialises.
Sources
- Fama, E. F. (1970). Efficient Capital Markets: A Review of Theory and Empirical Work. Journal of Finance, 25(2), 383-417.
- Fama, E. F. & Blume, M. E. (1966). Filter Rules and Stock-Market Trading. Journal of Business, 39(1), 226-241.
- Shiller, R. J. (1981). Do Stock Prices Move Too Much to be Justified by Subsequent Changes in Dividends?. American Economic Review, 71(3), 421-436.
- Brock, W., Lakonishok, J. & LeBaron, B. (1992). Simple Technical Trading Rules and the Stochastic Properties of Stock Returns. Journal of Finance, 47(5), 1731-1764.
- Sullivan, R., Timmermann, A. & White, H. (1999). Data-Snooping, Technical Trading Rule Performance, and the Bootstrap. Journal of Finance, 54(5), 1647-1691.
- Lo, A. W. (2004). The Adaptive Markets Hypothesis. Journal of Portfolio Management, 30(5), 15-29.
- Park, C.-H. & Irwin, S. H. (2007). What Do We Know About the Profitability of Technical Analysis?. Journal of Economic Surveys, 21(4), 786-826.
- Bailey, D. H., Borwein, J., López de Prado, M. & Zhu, Q. J. (2014). Pseudo-Mathematics and Financial Charlatanism. Notices of the AMS, 61(5), 458-471.
- Harvey, C. R., Liu, Y. & Zhu, H. (2016). … and the Cross-Section of Expected Returns. Review of Financial Studies, 29(1), 5-68.
- AMF (2014). Étude des résultats des investisseurs particuliers sur le trading de CFD et de Forex en France. Autorité des marchés financiers, 13 October 2014.
- S&P Dow Jones Indices. SPIVA Scorecards (Europe and United States), semi-annual publication.
- Aronson, D. R. (2006). Evidence-Based Technical Analysis. Wiley.
- López de Prado, M. (2018). Advances in Financial Machine Learning. Wiley.