| Vose Software

Industry: Banking and Financial Services
Product: ModelRisk
Application: Resampled mean-variance optimization with parameter uncertainty


The classical optimizer's "best" portfolio swung 38 points across data refreshes. The resampled rerun held a third tighter — for the same Sharpe.

An institutional asset-management team running mean-variance optimization on a 6-asset universe — US Equity, International Equity, EM Equity, US IG Bond, US Govt, Alternatives — had a long-standing problem: every quarterly rebalance produced a materially different "optimal" portfolio. Across 12 quarterly data refreshes the classical optimizer's weights swung by an average of 38 percentage points per asset (the worst single asset moved 71pp, low to high) on what felt like minor changes in expected-return inputs. The classical Markowitz optimizer was behaving as expected — it is famously sensitive to its inputs — but the resulting portfolio was not behaving as the team needed it to. The simulation rebuilt the optimizer with explicit parameter uncertainty, generating a 400-draw resampled (Michaud) allocation and a Black-Litterman-style shrunk allocation, then measured each candidate against the classical solution on two questions: how steadily the weights hold across data refreshes, and how the out-of-sample 5-year Sharpe distributions compare head-to-head.

The starting point is the efficient-frontier cloud — the long-only feasible set, coloured by Sharpe ratio — with the three candidate allocations marked on it.

Efficient frontier cloud — classical, Black-Litterman, and resampled allocations

The 6-asset universe

Expected real returns and volatilities (annualized):

Asset E[r] sigma role
US Equity 6.5% 16.5% growth core
International Equity 6.0% 18.0% growth diversifier
EM Equity 8.0% 23.0% growth + tail
US IG Bonds 3.0% 6.0% income
US Govt 1.8% 4.5% duration hedge
Alternatives 5.5% 12.0% diversifier

Correlations are realistic post-2010: equities tightly coupled (0.65-0.80 across regions), US Govt mildly negative against equities, IG Bonds neutral, Alternatives positively correlated with equities at 0.35-0.45. The correlation matrix is not the data the team disputes — it is the expected returns that move quarter to quarter and produce the weight instability.

Why classical MV fails on noisy inputs

The classical Markowitz solution maximizes the Sharpe ratio assuming the expected return vector and the covariance matrix are known. They are not. Empirically, the standard error on a sample estimate of an asset's mean return runs from about 0.5pp/year on US Govt to 2.4pp/year on EM Equity — small in absolute terms, yet large enough that the optimizer's "preferred" weight can swing by tens of percentage points between two adjacent rebalance dates that differ only in the last quarter of data. The stability experiment confirms it: averaged across the 12 refreshes, the classical lead-asset weight has a standard deviation of 12.1pp.

In the frontier cloud above (built from 8,000 random-Dirichlet long-only portfolios), the three candidate allocations cluster near the high-Sharpe ridge but differ in where they sit. The classical max-Sharpe portfolio (red dot) lands at 5.2% return / 8.8% volatility. The Black-Litterman-shrunk allocation (blue square) sits just to its right at 5.2% / 9.1% — accepting marginally more volatility in exchange for staying closer to an equal-weight prior. The resampled-frontier allocation (green triangle) is further up and to the right at 5.4% / 9.8%: a slightly higher expected return for slightly more volatility, the diversification it buys showing up as a fuller spread of weights rather than a leftward shift on the frontier.

What does "resampled" mean?

The team's implementation runs the optimizer 400 times, each time on a draw of the expected-return vector perturbed by its estimation noise (Normal noise with sd matching the standard error of each return estimate — 1.6pp/yr on US Eq, 2.4pp/yr on EM, 0.7pp/yr on US IG, down to 0.5pp/yr on US Govt). For each draw, the max-Sharpe allocation is recorded. The final resampled allocation is the mean across the 400 draws.

The effect on weight stability is decisive.

Weight stability — classical single-MV vs resampled mean weights, with error bars showing how far each method moves across 12 quarterly refreshes

Each bar is a method's average allocation across the 12 quarterly refreshes; each whisker is the standard deviation of that weight across the refreshes — a direct read on how much the allocation jumps quarter to quarter. The classical (red) solution concentrates and lurches: on its single max-Sharpe point estimate it pours 42% into Alternatives and 30% into US IG Bonds while pinning US Equity, International Equity and US Govt each near or below 2% (the corner-solution behaviour the Markowitz literature warns about), and its weights swing by an average of 38pp across the refreshes (sd 12.1pp). The resampled (green) bars show the Michaud average: a far more even, fully funded allocation — US Equity 18%, International 9%, EM 17%, IG Bonds 21%, US Govt 7%, Alternatives 29% — with no asset dominating.

The error bars carry the verdict. The resampled allocation swings by an average of only 26.5pp across the same 12 refreshes (sd 8.1pp), with its worst-moving asset travelling 45pp versus the classical optimizer's 71pp. Resampling does not eliminate movement — the inputs really are noisy — but it holds the allocation about a third tighter on every stability measure: average swing (26.5 vs 38pp), worst-asset swing (45 vs 71pp), and cross-refresh standard deviation (8.1 vs 12.1pp). That is the difference between a portfolio the investment committee can actually trade quarter to quarter and one it cannot.

The correlations underneath

The correlation matrix is not in itself disputed, but it is worth showing because every diversification argument lives there.

6-asset correlation heatmap

US-International-EM Equity cluster sits at 0.65-0.80 — the diversification benefit across geographies is small. US Govt is the only asset with consistently negative correlations to the equity bucket. Alternatives are not the diversifier the marketing brochures suggest — at 0.35-0.45 correlation with equities, they soften risk but do not insulate against an equity drawdown. This matrix is what makes the resampled answer realistic; the classical optimizer's 42%-Alternatives bet leans hard on that 0.35-0.45 correlation staying benign in a drawdown, when in fact Alternatives move with the equity bucket exactly when it hurts.

Out-of-sample Sharpe — the only test that matters

The question is not "which allocation has the highest ex-ante Sharpe on paper" but "which allocation produces the most reliable Sharpe over the next 5 years of actually-realized returns." A 1,500-trial out-of-sample simulation answers exactly that.

Out-of-sample 5-year Sharpe distribution — three optimization approaches

The result is a useful corrective to the hope that resampling buys a free lunch in realized return. All three allocations produce nearly indistinguishable out-of-sample Sharpe distributions: median 0.29 classical, 0.28 Black-Litterman, 0.30 resampled, with P10s clustered tightly around −0.27 to −0.32 and P90s near 0.84-0.89. The gaps between the medians are well inside the noise of the experiment — a statistical dead heat. On the single metric of realized 5-year Sharpe, the diversified resampled allocation neither beats nor trails the concentrated classical one in any way the team could trade on. The honest reading is that all three are drawing from the same realized-return process, so once you condition on the true mu/cov, the choice of allocation barely moves the realized-Sharpe distribution. The case for the resampled allocation is therefore not "higher realized Sharpe"; it is stability, fundability of every asset, and the avoidance of corner solutions — virtues that show up in turnover and governance, not in this Sharpe histogram.

Turnover and rebalancing

The resampled portfolio has a second institutional virtue: it is stable across estimation refreshes. The classical solution moves substantially each quarter, generating turnover that the team had been quietly absorbing in transaction costs of roughly 20-30bp annualized. The resampled solution moves much less because the noise in any one quarter's mu estimate is averaged across the 400 draws. The team's measured rebalancing turnover fell by roughly two-thirds in the year after adoption, recovering 15-20bp/year in execution costs — a return improvement larger than the entire alpha budget of some of the underlying mandates.

From insight to action

Three operational consequences:

  1. Resampled MV is now the production optimizer. Quarterly allocations come from the 400-draw resampled-mean weights, not the single-point classical MV. Allocations are reported with their resampled standard-deviation bars so the investment committee sees the honest width of the optimization answer, not just the central value.
  2. Black-Litterman shrinkage for view integration. Where the team has a strong tactical view, it is integrated through BL shrinkage with an alpha of 0.4 toward the resampled prior — keeping the corner-solution behavior contained even when the view is strong.
  3. Realized-Sharpe distribution as the KPI. The investment committee now scores the strategy on its rolling 5-year out-of-sample Sharpe distribution, not just its mean — explicitly tracking the spread between P10 and P90 as a model-risk indicator. A widening P10-P90 spread triggers a covariance recalibration cycle.

ModelRisk functionality used

  • 400 resampled efficient-frontier draws with Normal noise on the expected-return vector calibrated to estimation standard errors — the resampled mean weight is the production output.
  • Black-Litterman-style shrinkage between the classical max-Sharpe point estimate and the equal-weight prior, with shrinkage parameter alpha = 0.4 derived from the team's view-confidence assessment.
  • Out-of-sample Sharpe simulation — 1,500 5-year multivariate-Normal draws from the calibrated mu/cov, used to compare the three allocations on the metric that actually matters.
  • Correlation-heatmap diagnostic — explicit display of the 6x6 correlation matrix so the diversification assumptions are visible, not hidden inside the optimizer.
  • Weight-stability error bars — per-asset standard deviation across the 400 draws, reported alongside the central allocation as an honest measure of optimizer uncertainty.

The lesson the institutional team had to internalize: a Markowitz optimum on point estimates is a single answer to a single question that nobody asked. The right question is "what is the distribution of optimal allocations across the uncertainty in my inputs", and the right answer is the resampled allocation — the one whose weights are diversified and fundable across every asset, and whose turnover is lowest. The simulation also keeps the team honest about what resampling does not buy: on realized 5-year Sharpe the three approaches are a statistical dead heat, so the case for resampling rests on stability and governance, not on a promise of higher return.