Direct Answer
A market sentiment composite is a single index built by putting several indicators on a common scale and combining them. A defensible one groups correlated indicators into independent families, normalizes each against its own history, accounts for how fresh each input is, and keeps every component visible. It describes the state of the market and should never be read as a hidden buy or sell signal.
Market Sentiment Composite: Avoid Double Counting
Article
What is a sentiment composite, and why do people build one?
A sentiment composite is a derived index. It takes measures of market behavior that have different units (volatility points, basis points of credit spread, ratios, dollars of borrowing) and converts them into a comparable reading, then combines them into one headline number or label. The Fear and Greed composite is the best-known public example, and the idea behind it is simple: one number is easier to track than ten charts.
That convenience is also the danger. A composite looks more sophisticated than its inputs, and the reduction to one number can hide several design problems:
- Five indicators drawn from the options market may be counting the same information five times.
- A quarterly filing can be blended with an intraday statistic as if both were equally current.
- A noisy social-media score can dominate the result because it moves a lot, not because it is reliable.
- A neutral average can conceal a sharp disagreement between markets.
The rest of this guide walks through how a careful composite avoids those traps. Think of it as a reading guide too: if you use someone else's composite, the same questions tell you how much weight it deserves. The composite should be the last layer you look at, after the underlying evidence has been examined on its own. The source ladder explains how to judge each input first.
Why do sentiment composites double count information?
Double counting happens when several inputs are really measuring the same thing and each is treated as an independent vote. Suppose a composite includes the VIX level, the slope of the VIX futures curve, an equity put/call ratio and an options skew measure. When a selloff hits, all four tend to move together because all four come from the same place: the pricing and trading of index options. The composite then shows a dramatic reading, but it is one observation repeated four times, not four observations that agree.
The fix is structural rather than statistical. Instead of weighting indicators, weight information families. A family is a group of indicators that draw on the same underlying market or behavior. Each family gets one top-level weight no matter how many indicators it contains. Adding a fifth options indicator then does not make options five times more influential.
A useful beginner test: ask, "If the options market moved and nothing else did, how many of my inputs would change?" If the answer is most of them, they belong in one family.
Step 1: What information families should a composite use?
Start with families, not indicators. One reasonable taxonomy has nine, and the number can grow or shrink as long as the grouping principle holds:
- Options and volatility: the VIX level, the shape of the VIX futures curve, put/call ratios and skew-related measures. See the VIX term structure guide for how that family behaves.
- Futures positioning: normalized readings from the CFTC Commitments of Traders report. The Commitments of Traders guide covers how to read it.
- Credit risk appetite: investment-grade and high-yield spreads, covered in credit spreads and risk appetite.
- Leverage and financing: margin balances and related normalized measures. See margin debt and leverage.
- Breadth and participation: advance/decline measures, the share of stocks above moving averages, and new highs versus new lows. The market breadth section explains these.
- Fund and flow behavior: ETF and mutual fund flows, where the methodology is reliable enough to use.
- Survey sentiment: transparent recurring surveys such as the AAII bull-bear spread.
- Insider and institutional filings: Form 4 and Form 13F data, adjusted for their reporting lag and their limited economic meaning.
- Alternative and social: search interest, news tone and forum activity. These are explicitly lower trust and deserve a capped weight (more on that below).
The reason for separating families is not tidiness. It is that different families can fail differently. A volatility spike and a widening credit spread are two separate pieces of evidence; two put/call variants are not.
Step 2: How do you put different indicators on one scale?
Raw units cannot be averaged. The VIX is in volatility points, a high-yield spread is in percentage points, a put/call ratio is a plain ratio and margin debt is in dollars. Before any combination, each series is converted into a comparable historical state. Common choices:
- Rolling percentile: where today's value ranks within its own recent history. This is the easiest to explain to a general audience.
- Standard z-score: how many standard deviations from the historical average. It works when the distribution is reasonably stable.
- Robust z-score: the same idea using the median and the median absolute deviation, which is less distorted by crisis outliers.
- Rank transform: a close cousin of the percentile.
Percentiles are simple but throw away information about magnitude. A reading at the 99th percentile could be barely above the 98th or far above it; the percentile cannot tell you which. A robust z-score keeps more of that magnitude. For public education, a percentile as the default and a robust z-score as an advanced view is a sensible pairing, provided the composite records which one it used.
Step 3: How do you make every component point the same way?
Some high values mean stress and others mean risk appetite. Before combining, the direction has to be aligned so that every normalized component reads in the same semantic direction. For example:
- A high VIX percentile means more stress.
- A wide high-yield spread percentile means more stress.
- A strong breadth percentile means less stress.
- A high bullish-survey percentile means more optimism, which may or may not mean lower forward risk.
The last item shows why direction is not always obvious. A survey extreme should not be forced onto a good-or-bad scale. The honest approach is to label each component by what it actually measures, and to treat the composite as a market-state index rather than an expected-return forecast. "This reading describes elevated hedging demand" is a statement about the present. "This reading means stocks will fall" is a forecast the data cannot support.
Step 4: How do family weights control for correlation?
The standard structure has two layers:
- Family score = the weighted average of the normalized component scores within one family.
- Composite = the sum of each family weight multiplied by its family score.
Because each family gets one fixed top-level weight, component count does not bias the result. The inputs inside a family can change over time without shifting how much the family matters overall.
A worked example with illustrative numbers
The numbers below are made up to show the mechanics. They are not real market readings.
Imagine five families at equal 20% weights: volatility, credit, breadth, positioning and leverage. The volatility family contains four correlated components, with percentile scores of 80, 85, 78 and 82, so its family score is about 81. Credit has three rating buckets scoring 40, 45 and 35, so about 40. Breadth scores 30, positioning 55 and leverage 60.
| Family | Components | Family score (illustrative) | Weight | Contribution |
|---|---|---|---|---|
| Volatility | 4 | 81 | 20% | 16.2 |
| Credit | 3 | 40 | 20% | 8.0 |
| Breadth | 1 | 30 | 20% | 6.0 |
| Positioning | 1 | 55 | 20% | 11.0 |
| Leverage | 1 | 60 | 20% | 12.0 |
| Composite | 53.2 |
Now compare a naive average that gives each of the ten components equal weight. Volatility's four components would carry 40% of the total simply because there happen to be four of them, and the result would be 59.0 under the same inputs (the sum of all ten scores, 590, divided by ten). That is higher than the family-weighted 53.2, and the entire gap comes from the volatility family having more indicators, not from the market saying anything different. Component count has quietly decided which market matters most.
Step 5: How should data freshness affect a composite?
Inputs arrive at very different speeds. Cboe publishes daily market statistics, and the FRED series for the ICE BofA US High Yield option-adjusted spread is a daily close. The CFTC publishes its Commitments of Traders data weekly, describing positions as of an earlier day. FINRA's margin statistics are compiled from monthly reporting by member firms. Form 13F is due within 45 days after each calendar quarter ends. A 13F observation that is 40 days old is not "stale" in the same sense as a day-old options statistic, because its natural cadence is quarterly.
A transparent way to handle this is a freshness factor:
Effective weight = base weight x freshness factor
The factor should decline according to each source's own cadence, not one universal decay rate. A weekly positioning report does not become worthless after two days. A social feed might lose its relevance within hours. Whatever rule is used, a good composite displays the factor so the reader can see why a component contributes less than its base weight. The data latency and vintages guide goes deeper on release timing and revisions.
Step 6: Should source quality change the weights?
Source quality should influence what is included and how much confidence to place in it, but multiplying everything into a mysterious confidence score is a mistake. A clearer approach is to show, for each component:
- its evidence family
- its source class (official regulator or exchange data, a licensed index, a survey, a social feed)
- its freshness
- its current normalized state
- its contribution to the composite
If alternative or social data are included, they deserve a capped family weight until their usefulness has been validated over a long period. A volatile series should not earn influence just because it moves a lot.
Step 7: Why should a composite show its raw inputs?
A one-number display with no component detail cannot be audited. For each input, a trustworthy composite lets you see:
- the current raw value
- the normalized percentile or score
- the observation timestamp
- the source
- the transformation applied
- the direction convention
- the contribution to the headline
- the known limitations
If you cannot expand a composite and trace the headline back to its parts, treat it as an opinion rather than a measurement.
Step 8: What do good state labels look like?
Emotionally loaded labels such as "extreme greed means sell" imply a rule the evidence does not contain. Better labels describe conditions:
- broad risk appetite
- mixed or neutral evidence
- elevated hedging demand
- cross-market stress
- positioning crowding
- data conflict
The "data conflict" label matters most. A composite should not erase disagreement. If equity volatility is calm while credit spreads are widening, that divergence is information, and a good display surfaces it rather than averaging it away. The equity and credit divergence page explores why that pattern gets attention.
How does a conflict score work?
A conflict score measures dispersion among the normalized family states. High dispersion means independent evidence families disagree strongly. This matters because uncertainty may be highest precisely when the average looks neutral.
For example, volatility may be calm while credit is stressed and breadth is weak. A simple average could land near the middle. A conflict flag tells the reader that the middle is not consensus; it is disagreement. A dashboard can then show several separate facts side by side:
- State: mixed
- Evidence completeness: 86%
- Cross-family conflict: high
- Stale families: leverage
That tells you far more than a bare "sentiment score: 52."
Keep confidence separate from direction. Completeness describes how many required families are current and available. Source confidence describes provenance and methodology. Neither is confidence in future market direction, and none of them should be read that way.
How should missing or stale data be handled?
A robust composite keeps working when one source is unavailable, and it does so without hiding the gap. A sound approach:
- Mark the unavailable component explicitly.
- Never substitute zero, which would silently read as a bearish or bullish value.
- Re-normalize weights across the available components within the same family only if a minimum data threshold is met.
- Lower the displayed completeness status.
- Keep the last valid observation separately for reference, clearly labelled as old.
What a composite should never do is silently carry a stale value forward at full weight. A reader who sees a smooth line has no way to know that one input stopped updating weeks ago.
How does the lookback window change the answer?
Normalization depends on the historical window chosen. A one-year percentile adapts quickly but can label ordinary values "extreme" after an unusually calm year. A ten-year percentile gives broader context but can mix structurally different regimes, such as a zero-interest-rate era and a high-rate era. No single lookback is correct for every series.
The best practice is to choose a default window suited to each series, publish it, and where enough data exist, show medium-term and long-term context side by side. Data availability can also constrain the choice. For instance, the FRED page for the high-yield spread notes that, starting in April 2026, the series will include only three years of observations, which is a reminder that free public history can be shorter than the history a long-lookback method needs. Always check how much history a source actually provides before computing a percentile from it.
For new series with short histories, a careful composite avoids manufacturing a precise percentile from too few observations. It marks the component as having limited history and caps its contribution until enough data accumulate. That is more honest than assigning a confident rank to a tiny sample.
How should macro regime affect interpretation?
The same reading can mean different things in different environments. A given VIX percentile may carry different information when interest rates are stable than when they are repricing rapidly. A given credit spread may mean something different when defaults are low but Treasury market volatility is high. The financial conditions, credit spreads and liquidity page covers the macro backdrop.
The safe way to use regime information is as explanatory context, not as a prediction. A statement such as "this stress reading is occurring in a tightening financial-conditions regime" describes the setting. A statement that a regime implies a fixed future return claims more than historical data can deliver. Regime context belongs next to the composite as metadata, not inside it as a hidden adjustment.
How should revisions and vintages be treated?
Some sources revise their history. A composite built on the latest corrected series is not the same thing as what a person would have seen at the time. Two kinds of history should be kept distinct:
- As-published history: what was known on each date.
- Latest-vintage history: the newest corrected series.
Backtests and historical comparisons should prefer as-published history where it is available. Research charts can use latest-vintage values but must say so. Mixing the two quietly makes past readings look cleaner and more informative than they were in real time.
What are the alternatives for choosing weights?
Equal family weighting is a strong default because it is easy to explain and resistant to overfitting. Other defensible approaches include inverse-volatility weighting of normalized family scores, reliability weighting based on source quality, and fixed policy weights set before any market outcomes are examined. Each has tradeoffs: inverse-volatility weighting leans on recent history, reliability weighting requires subjective quality scores, and fixed weights can drift out of date.
What is not defensible is choosing weights by searching for the combination that best predicted past returns. That turns a descriptive state index into an optimized trading model and invites data mining: with enough combinations, something will fit the past by chance and then fail afterward. If research-weighted versions are explored at all, they belong in a clearly separate methodology exercise with genuine out-of-sample testing.
How do you validate a composite without overfitting it?
Optimizing a composite to maximize future returns is the wrong test. Better questions are about whether the construction is sound:
- Is it internally consistent?
- Does it stay stable when the lookback window changes?
- What are the correlations among the families, and are any nearly redundant?
- How sensitive is the result to a missing component?
- How does it behave during known stress episodes?
- Is it stable across data revisions?
- Do the labels match observable market conditions?
Checking whether families are truly independent
Family taxonomy should be tested periodically with rolling correlations. If two families become nearly redundant for long stretches, the grouping may no longer add information. Correlation alone does not prove duplication, since two distinct markets can move together during stress, but it is a prompt to look again. A more advanced diagnostic is the effective number of independent families, which can be estimated from the correlation matrix or a principal-component analysis. This is usually a research metric rather than something to show a general audience.
Testing whether a new component earns a place
When someone proposes adding an indicator, the question is whether it changes the information after controlling for what is already included. If a new options indicator is almost entirely explained by the existing volatility family, it may still deserve its own educational page, but it does not deserve an additional top-level weight. This rule keeps feature accumulation from quietly distorting the index. The combining breadth, volatility and sentiment without double counting guide applies the same logic to practical chart reading.
Sensitivity checks any reader can run
You can stress-test any composite mentally or in a spreadsheet:
- Remove one family at a time and see how much the state changes. A large change means the result depends heavily on that family.
- Change the lookback window and see whether a component is "extreme" only because of an unusually short window.
- Mark one family stale and recompute under a documented missing-data rule.
If the headline flips because of one removal, the composite is fragile, and that fragility is worth knowing about before relying on the reading.
Can you use a composite to look at what happened historically?
People naturally ask what happened after similar composite states. The responsible answer is a distribution, not a promise. Define the state in advance, show all historical matches rather than selected famous crashes or rallies, report the range of outcomes, and state the sample size. Because market structure and macro regimes change, these conditional histories are descriptive and say nothing certain about the next episode.
A related discipline concerns changes to the method itself. When a family or weight changes, old charts should not be silently recomputed under the new definition. Keep versioned outputs or at least a changelog with the effective date, and if history is recomputed, label it as backfilled under the new methodology. Changing weights because a recent episode made the composite look wrong is hindsight dressed up as research; the better test is whether a change represents the intended information families more faithfully and transparently, judged before looking at future outcomes.
What should a well-designed composite let you inspect?
Putting the steps together, a composite deserves trust to the extent that a reader can check the following:
- Every input is visible, with its source and timestamp.
- Formulas are versioned and changes are logged.
- Family-level anti-double-counting is explicit.
- Missing-data behavior is published.
- No hidden discretionary overrides exist.
- No fixed "buy" or "sell" threshold is attached to any reading.
- A component table accompanies every headline number.
- Color is never the only signal: text labels, numeric values and tables carry the same meaning for people who cannot distinguish a red-to-green gradient or who use a screen reader.
If a redesign makes the headline prettier while making the evidence harder to inspect, it is a step backward. The measure of a composite is explanatory value: which families are driving the state, which are stale, which disagree, and which methodology version produced the result.
A short exercise you can try
Build a five-family mock composite with equal family weights using made-up percentile scores. Then add three more options indicators to the volatility family without changing its weight, and compare the result with a naive average where every indicator counts equally. The gap shows how component count can bias a composite toward whichever market has the most indicators available.
Next, mark one family stale, recalculate under the missing-data rules above, and explain in plain language why the output changed. If you cannot explain the change simply, the method is too opaque to lean on. For a live look at how inputs are presented side by side, the sentiment dashboard shows individual indicators, and the market sentiment hub maps the full set of guides. You can also look up unfamiliar terms in the glossary.
What a sentiment composite does not tell you
- It does not forecast returns. It describes conditions that already exist.
- It does not reveal causes. Two identical readings can come from very different situations.
- It does not replace the underlying data. Percentiles hide magnitude and hide which input moved.
- It does not stay comparable across methodology versions unless the versions are tracked.
- It is not a trading rule. This material is educational and is not a recommendation to buy, sell, hedge or hold any security, option, futures contract, fund or digital asset, and past patterns do not guarantee future results.
The practical value of a composite is that it makes assumptions explicit and helps you pressure-test a view against evidence from more than one independent source.
Frequently Asked Questions
Why not just average every sentiment indicator?
Because several indicators often measure the same underlying activity, such as options volume and volatility. Averaging them as independent votes overweights that one source of information and makes the result depend on how many indicators happen to exist for each market. Grouping into families and giving each family one weight avoids that bias.
Should a sentiment composite predict returns?
A descriptive composite should report the market's current state, not claim to forecast a specific return. If historical outcomes after similar states are shown, they should come with the full set of matches, the sample size and a clear note that market structure changes. Tuning a composite to predict past returns is a common route to overfitting.
What happens when one component is stale?
A stale component should be labelled as stale and handled under a published rule, for instance by lowering its effective weight or excluding it and reducing the completeness reading. It should never keep full weight silently. Different sources have different natural cadences: a quarterly filing is not stale at 40 days, while a daily statistic may be after a week.
Why show a conflict score instead of one average?
A neutral average can hide strong disagreement. If volatility is calm, credit is stressed and breadth is weak, the average may sit in the middle even though no two families agree. A conflict measure reports how much the families disagree, which is a different and often more useful fact than the average itself.
Is a percentile or a z-score better for normalization?
Neither is universally better. Percentiles are easy to explain and handle odd distributions, but they discard magnitude. Robust z-scores based on the median and median absolute deviation keep more information about extremes and resist distortion from crisis outliers. A composite should state which it uses and over what lookback.
How much history does a composite need?
Enough that the percentile rests on a meaningful sample across more than one kind of market environment. Check the actual history each source provides, because some free series are limited in length. When history is short, mark the component as limited and cap its influence rather than reporting a falsely precise rank.