How many stocks make a diversified portfolio?
A stock count alone does not tell you how diversified a portfolio is. Weights, shared exposures and behaviour during bad periods matter too.
Someone shows you an investment strategy. Its rules are clear, and the line on the chart rises steadily across five decades while convincingly beating the market. There are no forecasts and no story about a miracle company, only rules and a long history. Where is the catch?
Most readers look for it in the rules. Perhaps they are too simple to work, or too complicated to trust. Yet the problem with a persuasive backtest usually lies in what the chart leaves out: which companies disappeared from the data, how many similar rules were tested and quietly discarded, why the period begins and ends where it does, and how much trading costs would have consumed.
In What Works on Wall Street, James O'Shaughnessy asked a different question. Instead of asking who had the best story, he asked which stock-selection rules had actually worked in the past. He tested them systematically over a long history. The fourth edition went further, combining several measures of cheapness and adding checks on financial strength and earnings quality. To his credit, O'Shaughnessy says plainly that every strategy he tested had bad periods.
Later independent research supported his broad lessons. Boudoukh and co-authors showed in 2007 that a broad measure of payouts to shareholders, counting buybacks as well as dividends, says more about future returns than dividend yield alone. Novy-Marx showed in 2013 that profitability deserves attention alongside price. For the foundation, see our explanation of value, momentum and quality. None of those studies independently confirmed the book's precise formulas, exact weights and exact number of portfolio holdings. A broad principle found in many forms and markets is therefore worth more than a recipe that works in one implementation. The book also shaped how we think about evidence at JonatanMars Invest: respectful of data and suspicious of recipes.
Before looking at the result of any backtest, ask five questions about the data.
First, where are the losers? If a database contains only companies and funds that still exist today, failed ones have silently disappeared from history and the past looks safer than it was. This is survivorship bias. Brown, Goetzmann, Ibbotson and Ross showed in 1992 how excluding failed funds can create the appearance of persistently successful management.
Second, could an investor have known this at the time? A test is broken if a 2005 decision uses information published in 2006 or an accounting correction released later still. Ask when each data point became public and when the test used it.
Third, how many rules were tested before this winner appeared? Test thousands of combinations of measures, periods and thresholds, and some results will look exceptional through luck alone. Researchers such as Harvey, Liu and Zhu therefore propose much stricter statistical standards for new findings. More searching requires stronger evidence.
Fourth, why does the chart start and end where it does? A period beginning after a major fall and ending near a peak flatters almost any strategy. Demand the full available history and several different starting points.
Fifth, who could really have traded it? A result built on the smallest, hardest-to-access stocks, or on constant portfolio changes, behaves differently on paper than it does in a world with trading costs, taxes and limited liquidity.
A simple ladder helps to separate the questions. Every backtest result sits on one of four levels, each with a different meaning.
An original result is what a researcher measured in one data sample. By itself it is only a hypothesis.
A reproduced result means that an independent researcher using the same data and definitions obtains the same numbers. When Andrew Chen and Tom Zimmermann closely followed the original methods of published studies, they reproduced a large majority of the results reported as statistically significant in the original papers. Reproducibility is necessary, but it is a low bar.
A robust result survives changed conditions: different definitions, weighting methods, periods and markets. The picture becomes less friendly here. Hou, Xue and Zhang placed hundreds of published effects into a consistent, more conservative implementation that reduced the influence of the smallest stocks. Most failed even the conventional statistical threshold. An international study by Jensen, Kelly and Pedersen was more forgiving. It found that most effects also appear around the world and cluster into a smaller number of common themes. The studies disagree less than it first seems because they answer different questions. Whether an original result can be reproduced is a different question from whether it survives changed conditions.
McLean and Pontiff showed that published strategies tend to weaken outside their original periods. Linnainmaa and Roberts showed the same. Part of the weakness comes from testing many rules before publication. Part comes after other investors begin exploiting the published opportunity.
An investable result is the last and strictest level. It asks what remains after trading costs, taxes, liquidity constraints and the bad periods an investor must endure. Novy-Marx and Velikov showed that costs reduce the results of every strategy. Strategies that changed holdings infrequently generally held up better, while few high-turnover strategies survived.
Each level reduces uncertainty. None removes it.
This leads to a standard that does not require specialist training. The rule should be written before the result, together with an economic reason for why it might work. The data should reflect what was known at the time and include companies that failed. The result should survive nearby versions of the rule and different periods. Evidence should also come from periods and markets the researcher did not use. Finally, show the net result after costs and include the bad periods, not only the most attractive slice.
One caveat belongs here, because without it the rest would mislead. Even a strategy that passes every test is not a promise. Markets adapt, definitions change and a genuine effect can lag for years. Investment values fluctuate and can fall despite a disciplined process. Evidence filters out weak ideas; it does not insure against loss.
Our guide chapter on the investment process shows how to use these standards in your own decisions.
So you judge a backtest by asking these five questions before you trust the curve, including when the chart is ours. At JonatanMars Invest, our Alpha strategy rebalances monthly and includes smaller companies, which is exactly where trading costs and thinner liquidity eat into results on paper. It can also fall further, so we offer it only to investors with risk profiles 6 and 7. Our tests therefore deduct estimated trading costs and bid-ask spreads, and only stocks that trade in enough volume make the portfolio.
If you would like to talk it through, book a free introductory consultation.
What is a backtest?
A calculation of how an investment rule would have performed if used in the past. It shows how the rule fitted history. It cannot show how the rule will perform in the future.
What is survivorship bias?
An error in which a test includes only companies or funds that still exist and omits those that failed. Because the worst outcomes disappear, the test makes history look safer and more profitable than it really was.
Does a good historical result imply a good future return?
No. Research shows that published strategies tend to weaken outside their original data periods. A good historical result reduces uncertainty but does not remove it. Investments can fall, and an investor can receive less than they invested.
Why do studies of strategy replication reach such different conclusions?
Because they measure different things. One type asks whether the original calculation can be reproduced with the original method. Another asks whether the result survives stricter and more consistent conditions. The first gives encouraging answers; the second is much harsher. Both questions matter.
How does JonatanMars Invest use these findings?
As a filter. We write rules before seeing the result, require evidence from several periods and markets, account for costs and expect bad periods in advance. These rules are process discipline and do not promise a return. Risk remains part of every investment.
What should I ask first when I see a persuasive chart?
Ask where the losers are. Does the dataset include companies and funds that failed during the period? Without an answer, the rest of the analysis has little value.
This is a marketing communication and general educational material, not personal investment advice.
Investing involves risk. The value of investments may fall as well as rise, and you may receive less than you invested. Past and any simulated or tested returns are not a reliable indicator of future returns. Tax treatment depends on personal circumstances and applicable law, both of which may change.
Luka Gubo is director and lead investment manager at JonatanMars Invest. He has been trading and investing since 2006.
More insights
A stock count alone does not tell you how diversified a portfolio is. Weights, shared exposures and behaviour during bad periods matter too.
Volatility is useful, but it does not capture permanent loss, liquidity, inflation or rare events. A smooth chart is not proof of safety.
What value, momentum and quality measure, why they may persist, and why each can disappoint for years.
Stay in touch