ESG Measurement: Why Rating Agencies Can't Agree on Who's Sustainable

This post is based on my seminar thesis at the Karlsruhe Institute of Technology (June 2022).

ESG ratings are supposed to tell investors how well a company handles environmental, social and governance risks. Large amounts of capital are screened and allocated with them. The catch: the major rating agencies disagree so much that the rating you pick can decide whether a company counts as a sustainability leader or a laggard.

This post explains how two of the largest agencies build their ratings, how large the disagreement is, where it comes from, and what that means if you use ESG data.

What ESG measures

The term goes back to a 2004 report by a group of financial institutions and has since become the standard frame for sustainable investing. The idea behind a rating is an independent assessment that does not rely only on what a company says about itself.

How MSCI builds its ratings

MSCI measures a company’s resilience to financially relevant, long-term ESG risks, relative to its industry peers, on a scale from AAA to CCC.

  1. For each industry, MSCI selects key issues, such as carbon emissions for utilities or product safety for pharmaceuticals.
  2. For each key issue, it scores exposure (0 to 10), based on business lines, locations and supply chain, and management (0 to 10), based on policies, programs and track record. High exposure requires strong management to score well.
  3. Environmental and social key issues each weigh between 5% and 30%. Governance always weighs at least 33%.
  4. The weighted average is normalized against industry peers and mapped to a letter grade.

How Sustainalytics builds its ratings

Sustainalytics measures how much of a company’s economic value is at risk from unmanaged ESG issues. Lower is better, and the score is absolute, so it can be compared across industries. It maps to five categories from negligible to severe.

The score combines corporate governance (about 20% of the risk on average), industry-specific material ESG issues and company-specific incidents. Exposure is expressed as a “beta” relative to the sub-industry average: above 1 means more exposed than a typical peer. The final number is the sum of risk the company cannot manage and the gap between manageable risk and what it actually manages.

The same company, two different answers

Aspect MSCI Sustainalytics
Question asked How resilient is the company to ESG risks? How much value is at risk from unmanaged ESG issues?
Scale AAA to CCC, higher is better 0 upward, lower is better
Comparison Relative to industry peers Absolute across industries
Governance weight At least 33% About 20% on average

Both methods assess exposure and management and both account for industry context, so they look similar at first glance. But they answer different questions on different scales, which already makes it likely that they rank companies differently.

How large the disagreement is

Berg, Kölbel and Rigobon (2022) compared six raters (KLD, Sustainalytics, Moody’s ESG, S&P Global, Refinitiv and MSCI) on a common sample of 924 companies. The pairwise correlations of their ratings range from 0.38 to 0.71. For comparison, credit ratings from Moody’s and S&P correlate at 0.99.

In practice, a portfolio screened with one provider can contain companies that another provider rates as high risk, and exclude companies the other considers leaders.

Where the disagreement comes from

The authors mapped all indicators of the six raters to a common taxonomy of 64 categories and split the divergence into three sources:

Source Question Share of divergence
Measurement Do raters measure the same attribute differently? 56%
Scope Do raters consider different attributes? 38%
Weight Do raters weight the same attributes differently? 6%

The main driver is measurement: raters look at the same attribute and arrive at different values. Weights matter surprisingly little. The raters’ most heavily weighted categories barely overlap, yet because scope and measurement differ so much, aligning the weights alone would not bring the ratings together.

Measurement disagreement reaches even simple facts. Berg et al. report a correlation of only 0.92 between raters on membership in the UN Global Compact, a yes-or-no fact from a public list. For categories that require judgment, correlations are much lower. The categories where disagreement is both large and heavily weighted are central ones: climate risk management, product safety, corporate governance, corruption, and environmental management systems.

The rater effect

Part of the measurement gap looks like a halo effect: once a rater sees a company as good or bad overall, that impression spills into its scores for unrelated categories. Berg et al. find that this rater effect explains roughly 15% of the variation in category scores. One plausible cause is organizational: analysts who cover whole companies, instead of specific ESG topics, form an overall impression that colors every category they score.

What this means in practice

Can standardization fix it?

Common definitions and mandatory disclosures would reduce the measurement problem, and regulation such as the EU’s Corporate Sustainability Reporting Directive moves in that direction. Standardization has limits, though. ESG has several dimensions that do not reduce to one number the way credit risk reduces to probability of default, assessing management quality always involves judgment, and rating agencies compete partly on their methodologies.

Takeaway

If you use ESG ratings, do not treat any single score as ground truth. Know which question your provider is answering, relative resilience or absolute unmanaged risk, compare several sources for decisions that matter, and look at the underlying category scores instead of only the headline rating.


References