Industry: Banking and Financial Services Product: ModelRisk Application: Resampled mean-variance optimization with parameter uncertainty
An institutional asset-management team running mean-variance optimization on a 6-asset universe — US Equity, International Equity, EM Equity, US IG Bond, US Govt, Alternatives — had a long-standing problem: every quarterly rebalance produced a materially different "optimal" portfolio. Across 12 quarterly data refreshes the classical optimizer's weights swung by an average of 38 percentage points per asset (the worst single asset moved 71pp, low to high) on what felt like minor changes in expected-return inputs. The classical Markowitz optimizer was behaving as expected — it is famously sensitive to its inputs — but the resulting portfolio was not behaving as the team needed it to. The simulation rebuilt the optimizer with explicit parameter uncertainty, generating a 400-draw resampled (Michaud) allocation and a Black-Litterman-style shrunk allocation, then measured each candidate against the classical solution on two questions: how steadily the weights hold across data refreshes, and how the out-of-sample 5-year Sharpe distributions compare head-to-head.
The starting point is the efficient-frontier cloud — the long-only feasible set, coloured by Sharpe ratio — with the three candidate allocations marked on it.
Expected real returns and volatilities (annualized):
Correlations are realistic post-2010: equities tightly coupled (0.65-0.80 across regions), US Govt mildly negative against equities, IG Bonds neutral, Alternatives positively correlated with equities at 0.35-0.45. The correlation matrix is not the data the team disputes — it is the expected returns that move quarter to quarter and produce the weight instability.
The classical Markowitz solution maximizes the Sharpe ratio assuming the expected return vector and the covariance matrix are known. They are not. Empirically, the standard error on a sample estimate of an asset's mean return runs from about 0.5pp/year on US Govt to 2.4pp/year on EM Equity — small in absolute terms, yet large enough that the optimizer's "preferred" weight can swing by tens of percentage points between two adjacent rebalance dates that differ only in the last quarter of data. The stability experiment confirms it: averaged across the 12 refreshes, the classical lead-asset weight has a standard deviation of 12.1pp.
In the frontier cloud above (built from 8,000 random-Dirichlet long-only portfolios), the three candidate allocations cluster near the high-Sharpe ridge but differ in where they sit. The classical max-Sharpe portfolio (red dot) lands at 5.2% return / 8.8% volatility. The Black-Litterman-shrunk allocation (blue square) sits just to its right at 5.2% / 9.1% — accepting marginally more volatility in exchange for staying closer to an equal-weight prior. The resampled-frontier allocation (green triangle) is further up and to the right at 5.4% / 9.8%: a slightly higher expected return for slightly more volatility, the diversification it buys showing up as a fuller spread of weights rather than a leftward shift on the frontier.
The team's implementation runs the optimizer 400 times, each time on a draw of the expected-return vector perturbed by its estimation noise (Normal noise with sd matching the standard error of each return estimate — 1.6pp/yr on US Eq, 2.4pp/yr on EM, 0.7pp/yr on US IG, down to 0.5pp/yr on US Govt). For each draw, the max-Sharpe allocation is recorded. The final resampled allocation is the mean across the 400 draws.
The effect on weight stability is decisive.
Each bar is a method's average allocation across the 12 quarterly refreshes; each whisker is the standard deviation of that weight across the refreshes — a direct read on how much the allocation jumps quarter to quarter. The classical (red) solution concentrates and lurches: on its single max-Sharpe point estimate it pours 42% into Alternatives and 30% into US IG Bonds while pinning US Equity, International Equity and US Govt each near or below 2% (the corner-solution behaviour the Markowitz literature warns about), and its weights swing by an average of 38pp across the refreshes (sd 12.1pp). The resampled (green) bars show the Michaud average: a far more even, fully funded allocation — US Equity 18%, International 9%, EM 17%, IG Bonds 21%, US Govt 7%, Alternatives 29% — with no asset dominating.
The error bars carry the verdict. The resampled allocation swings by an average of only 26.5pp across the same 12 refreshes (sd 8.1pp), with its worst-moving asset travelling 45pp versus the classical optimizer's 71pp. Resampling does not eliminate movement — the inputs really are noisy — but it holds the allocation about a third tighter on every stability measure: average swing (26.5 vs 38pp), worst-asset swing (45 vs 71pp), and cross-refresh standard deviation (8.1 vs 12.1pp). That is the difference between a portfolio the investment committee can actually trade quarter to quarter and one it cannot.
The correlation matrix is not in itself disputed, but it is worth showing because every diversification argument lives there.
US-International-EM Equity cluster sits at 0.65-0.80 — the diversification benefit across geographies is small. US Govt is the only asset with consistently negative correlations to the equity bucket. Alternatives are not the diversifier the marketing brochures suggest — at 0.35-0.45 correlation with equities, they soften risk but do not insulate against an equity drawdown. This matrix is what makes the resampled answer realistic; the classical optimizer's 42%-Alternatives bet leans hard on that 0.35-0.45 correlation staying benign in a drawdown, when in fact Alternatives move with the equity bucket exactly when it hurts.
The question is not "which allocation has the highest ex-ante Sharpe on paper" but "which allocation produces the most reliable Sharpe over the next 5 years of actually-realized returns." A 1,500-trial out-of-sample simulation answers exactly that.
The result is a useful corrective to the hope that resampling buys a free lunch in realized return. All three allocations produce nearly indistinguishable out-of-sample Sharpe distributions: median 0.29 classical, 0.28 Black-Litterman, 0.30 resampled, with P10s clustered tightly around −0.27 to −0.32 and P90s near 0.84-0.89. The gaps between the medians are well inside the noise of the experiment — a statistical dead heat. On the single metric of realized 5-year Sharpe, the diversified resampled allocation neither beats nor trails the concentrated classical one in any way the team could trade on. The honest reading is that all three are drawing from the same realized-return process, so once you condition on the true mu/cov, the choice of allocation barely moves the realized-Sharpe distribution. The case for the resampled allocation is therefore not "higher realized Sharpe"; it is stability, fundability of every asset, and the avoidance of corner solutions — virtues that show up in turnover and governance, not in this Sharpe histogram.
The resampled portfolio has a second institutional virtue: it is stable across estimation refreshes. The classical solution moves substantially each quarter, generating turnover that the team had been quietly absorbing in transaction costs of roughly 20-30bp annualized. The resampled solution moves much less because the noise in any one quarter's mu estimate is averaged across the 400 draws. The team's measured rebalancing turnover fell by roughly two-thirds in the year after adoption, recovering 15-20bp/year in execution costs — a return improvement larger than the entire alpha budget of some of the underlying mandates.
Three operational consequences:
The lesson the institutional team had to internalize: a Markowitz optimum on point estimates is a single answer to a single question that nobody asked. The right question is "what is the distribution of optimal allocations across the uncertainty in my inputs", and the right answer is the resampled allocation — the one whose weights are diversified and fundable across every asset, and whose turnover is lowest. The simulation also keeps the team honest about what resampling does not buy: on realized 5-year Sharpe the three approaches are a statistical dead heat, so the case for resampling rests on stability and governance, not on a promise of higher return.