What Makes a Stock More Volatile?
Research Question
Which structural and market characteristics are the strongest predictors of realized volatility in large-cap U.S. equities?
Background
Volatility, the standard deviation of log returns over a measurement window, is the central quantity in modern portfolio theory, options pricing, and risk management. Despite its central role, the cross-sectional determinants of why some stocks are persistently more volatile than others are less well understood than the time-series properties of volatility itself. We know from the GARCH literature that volatility clusters over time and mean-reverts, but this tells us little about the structural characteristics that place a stock in the high- or low-volatility regime in the first place.
The CAPM framework decomposes total variance into systematic (market-related) and idiosyncratic components. The systematic portion is explained by a stock's beta relative to the market portfolio. But the idiosyncratic component, which can account for a substantial fraction of total variance particularly for smaller and less covered firms, remains unexplained by factor models. Ang et al. (2006) produced the puzzling finding that stocks with high idiosyncratic volatility subsequently earn lower returns, a result that remains actively debated and has spawned a large empirical literature.
Beyond the academic puzzle, understanding volatility drivers has practical implications for portfolio construction, hedging, and options pricing. A model of expected realized volatility that incorporates structural firm characteristics would allow investors to anticipate volatility regime changes rather than simply adapting to them. This investigation approaches the question empirically: across a large cross-section of S&P 500 constituents, which observable firm and market characteristics most reliably predict next-quarter realized volatility?
Methodology
We analyze a cross-section of S&P 500 constituents using FRED and Yahoo Finance data spanning 2014 to 2024. For each stock, we compute realized volatility as the annualized standard deviation of daily log returns over rolling 63-trading-day (approximately one quarter) windows.
Predictor variables are collected at the start of each quarter and include: log market capitalization, price-to-book ratio, analyst coverage count (from I/B/E/S consensus data), the number of earnings announcements in the prior 12 months, average daily dollar volume (as a liquidity proxy), sector membership as fixed effects, and rolling 12-month beta versus the S&P 500. We also include a binary indicator for whether the company issued guidance in the prior quarter's earnings call.
We estimate a cross-sectional OLS regression with Fama-MacBeth standard errors, running the regression separately in each of 40 quarterly cross-sections and averaging the coefficients. This approach controls for time-fixed effects and produces standard errors that are robust to cross-sectional correlation of residuals. The R² values reported are the time-series average of quarterly cross-sectional R² values.
Visualizations
Log Market Cap vs. Realized Volatility
Volatility Distribution by Sector
Chart data computed from public sources.
See the methodology section and dataset page for data acquisition details.
Key Findings
Market capitalization has the strongest negative relationship with volatility (β of roughly -0.41, t = -12.3)
Analyst coverage provides additional explanatory power beyond size alone
Earnings announcement months show about 23% higher realized volatility on average
Technology and Biotech sectors contribute disproportionately to high-volatility outcomes
Limitations
Analysis is limited to S&P 500 constituents, introducing both survivorship bias (companies that failed or were delisted are excluded) and size bias (the smallest U.S. companies are not represented). Volatility is estimated from return data rather than observed directly, so measurement error in the dependent variable is present. The Fama-MacBeth approach controls for time-series correlation but assumes the coefficient estimates are stable across time, which may not hold across different volatility regimes such as 2020 versus 2017. Causality cannot be established from cross-sectional regression; the observed associations are predictive but not necessarily structural.