Skip to content
contested evidence

Factor premia are real in the sample and fragile out of it

The claim

Size, value and momentum premia appear in long historical samples, but the majority of published anomalies fail to replicate under consistent methodology, and the surviving premia have underperformed for stretches longer than most investors' patience.

Why we rate it contested

The original evidence is strong and the replication evidence is also strong, and they point in different directions. The disagreement is about how much of the historical premium was discovery and how much was data mining, and it is not settled.

Fama and French established that size and book-to-market explain variation in returns that a single market factor does not. That result reshaped asset pricing and is not seriously disputed as a description of the historical sample. A large industry now sells exposure to it.

The problem is what happened next. Hundreds of further anomalies were published, and Hou, Xue and Zhang re-tested a large set of them under one consistent methodology — value-weighting rather than equal-weighting, and screening out microcaps. The majority did not survive. The pattern of failure is diagnostic: anomalies concentrated in the smallest, least liquid stocks look strong when every stock counts equally and disappear when weighted by the money actually investable in them.

Harvey, Liu and Zhu made the statistical version of the argument. When hundreds of researchers test thousands of factors against the same data, the conventional significance threshold guarantees a stream of false discoveries. They argue the bar for a new factor should be roughly a t-statistic of 3 rather than 2, which retroactively disqualifies much of the published literature.

Then there is post-publication decay: premia tend to shrink after they are documented, consistent with capital arriving to exploit them. And even the survivors have brutal droughts — US value underperformed growth for well over a decade after 2007, long enough that any real investor holding it would have faced years of being told they were wrong by the only evidence they could see, their own account balance.

Where this leaves a DIY investor: a tilt towards value or small-cap is defensible, cheap to implement, and should be sized on the assumption that the future premium is a fraction of the historical one — and held for decades or not at all. What is not defensible is treating a backtested factor product as a higher-return version of the market index. The premium, if it exists, is compensation for a risk or a discomfort you have to actually bear.

Where this breaks down

  • This note is deliberately two-sided because the evidence is. Anyone who tells you the factor debate is settled, in either direction, is selling something.
  • Implementation costs, turnover and tax drag consume part of any premium, and they are certain in a way the premium is not.
  • A factor tilt raises tracking error against the index everyone else talks about, which is a behavioural cost as well as a statistical one.
  • The market factor itself — simply owning equities — has far stronger evidence behind it than any of its subdivisions, and is available for a few basis points.

Sources

Follow these rather than taking our word for the summary.

  • Eugene F. Fama and Kenneth R. French (1993). Common risk factors in the returns on stocks and bonds

    Journal of Financial Economics, 33(1), 3-56

    Finding: Size and book-to-market factors capture cross-sectional variation in average stock returns that the market factor alone does not explain.

  • Kewei Hou, Chen Xue and Lu Zhang (2020). Replicating Anomalies

    The Review of Financial Studies, 33(5), 2019-2133

    Finding: Re-testing 452 published anomalies with value-weighted returns and microcap screens, the majority failed to replicate at conventional significance levels.

  • Campbell R. Harvey, Yan Liu and Heqing Zhu (2016). ... and the Cross-Section of Expected Returns

    The Review of Financial Studies, 29(1), 5-68

    Finding: Accounting for the volume of factors tested, a new factor should clear a t-statistic near 3.0 rather than 2.0, invalidating many published results.

Test this on your own numbers

Related notes