Financial Risk

Market risk measurement with VaR, expected shortfall and stress testing

A technical overview of market risk measurement, covering VaR, expected shortfall, stress testing, model validation, governance and reporting. The article links statistical design choices to risk appetite, capital assessment and supervisory expectations.

By Jonas Osman AbdelghafourPublished August 1, 2026Last reviewed August 1, 2026

Market risk is the potential for losses arising from movements in prices, rates, spreads, volatilities and correlations across traded and balance sheet exposures. Value at risk, expected shortfall and stress testing remain central tools, but they serve different purposes. VaR is a quantile estimate, expected shortfall measures average tail loss beyond a confidence threshold, and stress testing examines specified severe but plausible conditions. A sound framework does not treat any one metric as sufficient. It links model design, data quality, valuation control, limit structures, capital assessment, escalation and independent validation into a coherent governance process.

The measurement problem in market risk

Market risk measurement begins with a definition of the portfolio, the risk factors and the decision use of the metric. A trading desk limit metric, an economic capital metric and a board risk appetite metric may use related inputs but should not be assumed to be interchangeable.

The principal modelling challenge is that market returns are not stable, normally distributed or independent over all horizons. Tail dependence, volatility clustering, illiquidity, basis risk and non-linear payoffs create exposures that are not fully captured by a single daily loss distribution. This is why market risk frameworks commonly combine statistical metrics with scenario analysis and expert judgement.

Key design questions include:

  • which positions are in scope, including hedges, basis exposures and valuation adjustments;
  • which risk factors are mapped explicitly and which are proxied;
  • whether the model is full revaluation, sensitivity based or a hybrid;
  • the holding period and confidence level;
  • the treatment of liquidity horizons, optionality, stochastic volatility and correlation breakdown;
  • the frequency of calculation, monitoring and escalation.

These choices should be tied to governance. For example, a board-approved appetite statement should define how market risk capacity is allocated and which limit breaches require management action. Related governance considerations are discussed in market risk governance, risk appetite that actually guides decisions and board risk reporting.

VaR as a quantile measure

Value at risk estimates the loss level that should not be exceeded with a specified probability over a specified horizon, assuming the model and data are representative. A one-day 99% VaR of 10 million means that, under the model, losses greater than 10 million are expected on about 1% of trading days. It does not say how large losses may be when that threshold is exceeded.

Common VaR methods

The main implementation approaches are:

  • Historical simulation: applies historical risk factor changes to the current portfolio. It is transparent and captures empirical non-normality, but depends heavily on the lookback window and may understate new regimes.
  • Parametric or variance-covariance VaR: estimates loss using means, volatilities and correlations, often with distributional assumptions. It is computationally efficient but may be weak for options, credit spread jumps and non-linear exposures.
  • Monte Carlo VaR: simulates risk factor paths from specified stochastic processes. It can address non-linearity and path dependence, but introduces model risk through calibration and distributional assumptions.

In practice, a bank or insurer may run several variants for control purposes. A production VaR model may be used for limits, while sensitivity measures, stress measures and profit and loss attribution are used to understand drivers.

Main weaknesses of VaR

VaR has several well-known limitations. It is not subadditive in all cases, meaning diversification effects may be inconsistently represented. It can create cliff effects around the quantile. It can also be insensitive to the severity of losses beyond the threshold. Two portfolios may have the same VaR but materially different tail loss profiles.

VaR is therefore better viewed as one control statistic rather than a complete description of market risk. It is useful for day-to-day monitoring because it is compact and comparable, but it should be interpreted alongside expected shortfall, stress losses, sensitivities and scenario narratives.

Expected shortfall and tail loss estimation

Expected shortfall estimates the average loss conditional on losses exceeding a selected quantile. If VaR identifies the edge of the tail, expected shortfall estimates the mean of the tail beyond that edge. This makes it more informative for capital and solvency discussions, particularly where tail severity is a concern.

The Basel market risk framework has incorporated expected shortfall in its revised internal models approach for market risk capital. Supervisory frameworks also place emphasis on backtesting, profit and loss attribution and risk factor modellability. The technical point is that a tail-sensitive metric must still be supported by observable data, robust valuation and credible model governance.

Estimation issues

Expected shortfall is more tail-sensitive than VaR, but that is also its difficulty. Tail observations are sparse. Estimates can be unstable where market history does not include enough severe observations or where the current portfolio differs from past exposures. For illiquid instruments, risk factor histories may be short, stale or proxied.

Important design controls include:

  • clear rules for selecting and refreshing historical data windows;
  • documented proxy selection and proxy performance testing;
  • separate treatment of modellable and non-modellable risk factors where relevant;
  • review of tail scenarios to identify whether the statistical tail is economically plausible;
  • comparison of expected shortfall with stress losses and realised loss experience.

Expected shortfall also raises communication issues. A single tail average may be interpreted as a worst-case loss, which it is not. Board and senior management reporting should explain the confidence level, horizon, model limitations and actions linked to limit utilisation. This is part of broader risk ownership and accountability.

Stress testing and scenario analysis

Stress testing asks a different question from VaR or expected shortfall. Instead of estimating a loss distribution from historical or simulated data, it evaluates the portfolio under defined adverse conditions. These may be historical, hypothetical, reverse or sensitivity based.

Types of market risk stress tests

A mature market risk programme usually includes several categories:

  • Historical scenarios: replay past market disruptions using risk factor moves from specified periods.
  • Hypothetical macro-financial scenarios: apply internally designed or supervisory scenarios involving rates, inflation, credit spreads, equities, foreign exchange, commodities and volatility.
  • Sensitivity stresses: shock one or more risk factors by defined magnitudes to identify concentrations and convexity.
  • Reverse stress tests: identify scenarios that would breach capital, liquidity, collateral or risk appetite thresholds.
  • Idiosyncratic desk scenarios: assess exposures that are specific to a product, basis trade, hedge strategy or market segment.

Stress testing should not be treated as a periodic reporting exercise only. It should inform hedging, limit calibration, new product approval, contingency planning and capital adequacy assessments. Alignment with stress testing programmes, liquidity risk management and asset and liability management is important because severe market moves often interact with funding, collateral and behavioural assumptions.

Severity and plausibility

The phrase severe but plausible requires discipline. A scenario may be severe in one dimension but internally inconsistent across asset classes. Conversely, a scenario may be historically observed but insufficiently severe for the current risk profile. Governance should require a rationale for the scenario narrative, the calibration of shocks, the treatment of second-order effects and the assumed management actions.

For non-linear portfolios, full revaluation under stress is preferable where feasible. Sensitivity-based approximations can be acceptable for screening, but they may understate losses when gamma, vega, credit migration or correlation effects dominate.

Framework

A robust market risk framework connects measurement, governance and assurance. The following structure can be used by risk functions and model validators when assessing VaR, expected shortfall and stress testing arrangements.

1. Scope and risk taxonomy

Define the portfolio perimeter and risk taxonomy. The taxonomy should cover interest rate risk, credit spread risk, equity risk, foreign exchange risk, commodity risk, volatility risk, inflation risk and basis risk where relevant. It should also distinguish trading book market risk from structural balance sheet risks such as interest rate risk in the banking book, even where data and risk factors overlap.

2. Data and valuation control

Market risk metrics are only as reliable as the prices, curves, volatilities, correlations and position records used to calculate them. Controls should include independent price verification, stale price checks, corporate action controls, curve construction standards and reconciliations between front-office and risk systems.

Data lineage should be documented for material risk factors. Where proxies are used, the framework should require justification, performance monitoring and periodic challenge. Proxy risk can be material for illiquid credit, private assets, structured products and emerging market exposures.

3. Model design and implementation

The model methodology should specify horizon, confidence level, weighting schemes, distributional assumptions, revaluation approach, risk factor mapping and aggregation. Implementation should be controlled through version management, change approval, testing environments, independent code review where proportionate and documented production controls.

Model changes should be assessed not only for statistical performance but also for business impact. A change that materially reduces measured risk requires evidence that the reduction reflects improved measurement rather than an unintended weakening of controls.

4. Limits, escalation and use test

A model that is not used in decisions is unlikely to be an effective control. Limit structures should link VaR, expected shortfall, stress losses and sensitivities to desk mandates and risk appetite. Escalation triggers should be clear, including breach classification, permitted remediation, approval authority and reporting timelines.

The use test should cover daily monitoring, new trades, hedge assessment, stress loss review, capital planning and senior management reporting. Governance weaknesses often arise where limits exist formally but are not aligned with actual trading behaviour or risk transfer.

5. Independent validation and assurance

Independent validation should assess conceptual soundness, data integrity, implementation accuracy, outcomes analysis and ongoing monitoring. It should also review model limitations and compensating controls. Useful reference points include independent model validation standards, model risk management frameworks and model limitations and compensating controls.

Validation should be risk based. A simple sensitivity limit tool may require less extensive review than an internal capital model, but all material tools require ownership, documentation and performance monitoring.

Worked numerical illustration

Consider a simplified trading portfolio with daily profit and loss generated from a historical simulation model. Losses are expressed as positive values. Assume 250 daily observations are used and the losses are sorted from smallest to largest.

For a 99% one-day VaR, the relevant point is near the 99th percentile. With 250 observations, 1% corresponds to 2.5 observations in the tail. A simple conservative convention may select the third-largest loss. Suppose the three largest historical simulated losses are:

Rank by loss severitySimulated daily loss
118.0 million
215.5 million
314.0 million
411.0 million
59.5 million

Under this illustrative convention, the 99% VaR is 14.0 million. If expected shortfall is calculated as the average of losses beyond the VaR threshold using the three largest losses, expected shortfall is:

(18.0 + 15.5 + 14.0) / 3 = 15.83 million.

The difference is material. VaR reports the selected quantile at 14.0 million, while expected shortfall reports the average tail loss at 15.83 million. If a stress test applies a combined spread widening, equity fall and volatility increase that produces a full revaluation loss of 28.0 million, the stress loss is substantially larger than both statistical measures.

This example is intentionally simplified. A production model would need defined interpolation rules, valuation controls, treatment of overlapping horizons, risk factor mapping, liquidity horizons and validation of the historical window. The example illustrates why risk committees should not infer maximum loss from VaR and why stress testing remains necessary.

Backtesting, attribution and validation evidence

Backtesting compares realised or hypothetical profit and loss with model estimates. For VaR, exceptions occur when losses exceed the VaR estimate. Exception counts, clustering and size should be reviewed. A model with an acceptable number of exceptions may still be weak if exceptions cluster during volatility regimes or if losses beyond VaR are systematically large.

Expected shortfall backtesting is more complex than VaR backtesting because the metric concerns conditional tail averages. Validators therefore often use a combination of tests, including exceedance analysis, tail loss comparison, sensitivity review and benchmarking. Profit and loss attribution is also important: risk-theoretical P&L should explain actual or hypothetical P&L movements to a reasonable degree, subject to documented valuation and data differences.

A validation checklist for market risk models should include:

  • scope: complete inventory of positions, risk factors, products and legal entities;
  • data: price source hierarchy, missing data controls, proxy rationale and outlier treatment;
  • methodology: confidence level, horizon, weighting, distribution and aggregation assumptions;
  • valuation: evidence that non-linear and path-dependent instruments are adequately revalued;
  • implementation: reconciliations, code review, user acceptance testing and change logs;
  • outcomes: VaR exceptions, expected shortfall stability, stress loss sensitivity and benchmarking;
  • governance: limits, breach escalation, committee review and management action tracking;
  • limitations: documented weaknesses, compensating controls and remediation plans.

Validation should be independent from model development and business ownership. However, independence does not mean isolation. Validators need access to traders, quants, finance, technology and risk managers to understand how the model behaves in production.

Governance, reporting and capital use

Market risk reporting should support decisions at different levels of the organisation. Desk heads need timely exposure and limit information. Risk committees need utilisation, breaches, emerging concentrations, stress results and model performance. Boards need a concise view of appetite, material changes, tail vulnerabilities and management actions.

Good reporting distinguishes between measurement movement and risk movement. A VaR increase may reflect new positions, higher volatility, correlation changes, model recalibration or data corrections. Without attribution, management may misread the risk signal.

Capital processes should connect market risk measurement to the internal capital adequacy assessment process or equivalent enterprise capital framework. The ECB guide to ICAAP emphasises internal capital adequacy, risk identification and governance expectations for significant institutions under its supervision. Insurers will also consider market risk within solvency, ORSA and stress testing processes, with sector-specific expectations.

Risk appetite should define both quantitative and qualitative constraints. Quantitative metrics may include VaR, expected shortfall, stress loss, sensitivities and loss triggers. Qualitative constraints may address product complexity, permitted hedging strategies, basis risk, model reliance and liquidity of risk factors.

Limitations

VaR, expected shortfall and stress testing all have limitations that should be explicit in model documentation and risk reporting.

First, historical data may not represent future regimes. A calm lookback period can understate risk, while a crisis-heavy period can overstate current day-to-day volatility. Weighted approaches reduce but do not eliminate this issue.

Second, tail dependence and correlation breakdown are difficult to estimate. Diversification benefits may disappear under stress, particularly where liquidity withdrawal, margin calls and crowded hedging amplify market moves.

Third, valuation uncertainty can dominate statistical uncertainty. This is relevant for illiquid credit, structured products, complex derivatives and instruments that require model-based valuation inputs.

Fourth, management actions are uncertain. Stress tests may assume hedging, asset sales or balance sheet actions that are not feasible at the assumed prices or time horizons. Assumptions should therefore be challenged and, where material, supported by operational playbooks.

Fifth, model outputs can create false precision. Reporting a VaR figure to multiple decimal places does not imply that the underlying loss distribution is known precisely. Effective governance requires ranges, sensitivities and clear limitation statements.

Frequently asked questions

How is market risk different from credit risk?

Market risk arises from changes in market prices and risk factors, such as rates, spreads, equities, foreign exchange, commodities and volatility. Credit risk arises from counterparty default or deterioration in credit quality. The two can interact: credit spread widening is a market risk factor for traded credit instruments, while borrower default is primarily credit risk.

Why use expected shortfall if VaR is already calculated?

VaR identifies a loss quantile but does not measure the severity of losses beyond that point. Expected shortfall estimates the average loss in the tail beyond the selected quantile. This makes it more informative for tail risk and capital discussions, although it is also more sensitive to sparse data and modelling choices.

Can stress testing replace VaR and expected shortfall?

No. Stress testing and statistical measures answer different questions. VaR and expected shortfall estimate risk from a modelled loss distribution. Stress testing evaluates specified scenarios, including conditions that may not be well represented in historical data. A complete framework uses both.

What should a board focus on in market risk reporting?

Boards should focus on appetite utilisation, material concentrations, limit breaches, stress vulnerabilities, model limitations and management actions. They do not need every desk-level sensitivity, but they do need enough information to challenge whether the risk profile remains consistent with strategy and capital capacity.

Professional disclaimer: This article is for general technical information and does not constitute legal, regulatory, actuarial or investment advice.

Frequently asked questions

What should risk leaders know about the measurement problem in market risk?

Market risk measurement begins with a definition of the portfolio, the risk factors and the decision use of the metric. A trading desk limit metric, an economic capital metric and a board risk appetite metric may use related inputs but should not be assumed to be interchangeable.

What should risk leaders know about vaR as a quantile measure?

Value at risk estimates the loss level that should not be exceeded with a specified probability over a specified horizon, assuming the model and data are representative. A one-day 99% VaR of 10 million means that, under the model, losses greater than 10 million are expected on about 1% of trading days. It does not say how large losses may be when that threshold is exceeded.

What should risk leaders know about expected shortfall and tail loss estimation?

Expected shortfall estimates the average loss conditional on losses exceeding a selected quantile. If VaR identifies the edge of the tail, expected shortfall estimates the mean of the tail beyond that edge. This makes it more informative for capital and solvency discussions, particularly where tail severity is a concern.

What should risk leaders know about stress testing and scenario analysis?

Stress testing asks a different question from VaR or expected shortfall. Instead of estimating a loss distribution from historical or simulated data, it evaluates the portfolio under defined adverse conditions. These may be historical, hypothetical, reverse or sensitivity based.

What should risk leaders know about framework?

A robust market risk framework connects measurement, governance and assurance. The following structure can be used by risk functions and model validators when assessing VaR, expected shortfall and stress testing arrangements.