Direct Answer
A market sentiment source ladder is a ranking of sentiment data by how well you can trace, reproduce, and defend each number. Exchange statistics, regulator-collected datasets, and official filings usually sit highest, surveys and licensed index series sit in the middle, and derived composites and social-media scores sit lowest because more can go wrong between the raw observation and the number on your screen. The ladder does not tell you whether a signal predicts anything; it tells you how much confidence the underlying evidence deserves.
Market Sentiment Source Ladder: Which Data to Trust
Article
Why does sentiment data need its own source ladder?
Most weak sentiment charts are not weak because of their formula. They are weak because of their inputs. A line can look precise while the numbers behind it come from an unstable web page, a sample whose makeup changes over time, a transformation nobody documented, or a release that arrives days after the event it describes. A dramatic chart is not automatically a trustworthy one.
The general method for matching any financial claim to the right kind of evidence is covered on the Swoopr Source Ladder, the parent methodology for this site. That page asks which authority owns a given statement, whether a rule, a filing, a dataset, or a piece of commentary. This page applies the same thinking to one narrow job: judging the data that feed sentiment, positioning, and risk-appetite analysis. The question changes from "who has authority over this claim?" to "how far is this number from the raw behavior it claims to measure?"
That second question matters because sentiment data are unusual. A tax limit is either correct or it is not. A sentiment reading is always a stand-in for something nobody can observe directly: the mood, fear, or crowding of millions of participants. Every step between the real behavior and the published number (sampling, aggregation, timing, scoring) is a place where information is lost or distorted. The ladder makes those steps visible. It belongs alongside the broader market sentiment hub, where each indicator family is explained.
What are the five levels of the ladder?
Each level is defined by what kind of organization produced the observation and how much processing sits between the observation and the reader. Higher is not "true" and lower is not "useless." The levels describe the default amount of skepticism a reader should bring.
| Level | Source type | Typical examples | Main strength | Main weakness |
|---|---|---|---|---|
| 1 | Primary market, regulatory, or official data | CFTC Commitments of Traders, Cboe options statistics, SEC insider and 13F filings, FINRA margin statistics | Direct provenance and documented definitions | Delays, partial coverage, reporting errors |
| 2 | Licensed institutional series with published methodology | Index-provider credit spreads distributed through FRED, exchange-derived volatility indexes | Consistent professional calculation | Licensing limits, complex methodology |
| 3 | Transparent surveys and research datasets | Recurring investor surveys with documented questions | Measures stated opinion directly | Sample and non-response bias |
| 4 | Derived public indicators | Breadth composites, put/call percentiles, volatility spreads | Turns raw data into context | Transformation choices, double counting |
| 5 | Social, media, and alternative data | Social-post polarity, search interest, news tone, forum activity | Fast, captures emerging narratives | Bots, selection bias, unstable access, manipulation |
Level 1: primary market, regulatory, and official data
Level 1 is data produced by the body that runs the market, collects the filing, or regulates the activity. Four examples show the pattern.
The Commodity Futures Trading Commission publishes the Commitments of Traders report, which shows how large traders are positioned in futures markets. The CFTC states that it releases the report weekly using data from the preceding Tuesday. Its explanatory notes also say that reportable positions typically represent 70 to 90 percent of total open interest, which means the report is a large sample of the market, not the whole market.
Cboe, which operates the exchange where many US options trade, publishes daily options statistics including put/call ratios and volume by product category. Its own disclaimer says the data were obtained from sources believed to be reliable but that accuracy is not guaranteed. That is worth reading carefully: even an exchange page does not promise perfection.
FINRA collects margin data from its member firms under a monthly reporting requirement and publishes aggregated statistics. The firms report as of the last business day of each month, so the series is a monthly snapshot, not a live reading of leverage. The margin debt and leverage page explains how to read it, and the FINRA margin debt indicator page documents the series.
The SEC publishes insider-transaction filings (Forms 3, 4, and 5) and quarterly Form 13F holdings reports from large institutional managers. The 13F is due within 45 days after the end of a calendar quarter, and the filer threshold is investment discretion over $100 million or more in covered securities. A 13F therefore describes a portfolio as it stood on a past quarter-end date, which can be six weeks stale by the time it becomes public.
What Level 1 gives you is accountability: you know who produced the number, you can usually read the definition, and the record is generally the closest inspectable version of what happened. What it does not give you is perfection. Late filings, amendments, missing values, and definition changes all occur.
Level 2: licensed institutional series with published methodology
Level 2 covers data that a professional index or analytics provider calculates from market prices under a documented method, then distributes to the public through another channel. The standard example is a corporate-bond option-adjusted spread: the extra yield that bonds of a given credit quality pay over comparable Treasuries after adjusting for embedded options. The ICE BofA US High Yield spread, distributed by the St. Louis Fed's FRED service, is widely used as a read on credit stress. The credit spreads and risk appetite page covers how analysts interpret it.
The tradeoffs at this level are practical. The method is consistent and well known, but you did not collect the data and you may not be allowed to copy them. FRED's page for that series says the underlying data come from ICE Data Indices, that reproduction is prohibited without ICE's permission, and that the series on FRED carries a limited window of history with longer history available from the original source. A chart built on it is only as complete as the license allows. This is a separate quality dimension from accuracy, and it is covered further below.
Level 3: transparent surveys and research datasets
Surveys measure what respondents say they feel, not what they have done with money. That is both their appeal and their limit. A long-running survey with a stable question, a published sampling approach, and a consistent schedule can be a useful read on opinion over time. The AAII bull-bear spread is one familiar example, and the investor survey signals article discusses how people use such readings.
The careful phrasing is "survey respondents reported," not "investors believe." Respondents are a self-selected group, not a random slice of all capital, and stated opinion can diverge from actual positions. A methodology change, such as a new question wording or a new sampling frame, can also break comparability with older data without any visible sign on the chart.
Level 4: derived public indicators
Level 4 is where raw inputs get processed into something easier to read. Put/call percentiles, volatility spreads, breadth composites, and any blended "fear and greed" style index belong here. A derived indicator is only as strong as its weakest input plus the quality of its recipe. Two issues recur.
The first is transformation sensitivity. If one version of an indicator ranks today's value against the past 252 trading days and another ranks it against five years, the two can disagree about whether conditions are extreme, even though both start with identical raw data. Neither is wrong; they answer different questions. Reputable indicators say which window they use.
The second is double counting, covered in its own section below. Both problems are easy to miss because the output looks like a clean, single number.
Level 5: social, media, and alternative data
Level 5 includes social-post polarity, news-tone scores, forum activity, and search-interest data. These sources are fast and can show a narrative forming before slower data confirm it. They also carry the heaviest list of cautions: bots, coordinated posting, paid promotion, unrepresentative platform audiences, classification errors from automated language models, and data access that can change without notice.
Regulators have said as much. A FINRA and SEC investor alert on social sentiment tools warns that information from such tools may be inaccurate, incomplete, or misleading, that stale posts can distort readings, and that posts can be used to spread false information to try to move a price. The same alert says investors should not rely solely on these tools. Google's own FAQ for Google Trends explains that the service uses a sample of search requests, scales results to a 0 to 100 index relative to each query's share of searches, and filters out very low-volume queries, so a reading of 100 means "peak interest for this query in this window" and not a count of searches. The search interest and narrative page goes deeper on those mechanics.
Level 5 is not worthless. It is the level that most needs its limitations stated next to the number, and it is the level where reading a value in isolation is most likely to mislead.
How can you rate a single source on five dimensions?
A ladder level is a starting point. Two sources at the same level can differ a lot, so it helps to rate each one on five questions instead of collapsing everything into one "trust" score.
Provenance. Can you name the organization that created the raw observation? "An exchange," "a regulator," and "a website that aggregates several vendors" are three very different answers.
Reproducibility. Could another careful person rebuild the number from the published source and a written formula? If the formula is proprietary or undocumented, the number is a claim, not a result.
Timeliness. How long after the event does the number become available, and does the source say so? A weekly report that describes Tuesday and appears days later is useful for context, not for what is happening this minute.
Coverage. Does the dataset represent the relevant market, or a narrow slice? An index-options statistic says little about single-stock behavior, and the reverse is also true.
Manipulation resistance. How easy is it for participants to distort the measurement, deliberately or accidentally? A regulated exchange print is hard to fake. A count of social posts is not.
Keeping these as five separate descriptors is more honest than a single badge, because a source can be strong on one and weak on another. The CFTC report is strong on provenance and manipulation resistance and weak on timeliness. A social-media score can be excellent on timeliness and poor on nearly everything else.
Why do observation date and publication date matter?
Every sentiment number has at least two dates, and mixing them up is one of the most common errors in the field. The observation date is the date the data describe. The publication date is the date the data became available to you. Often there is also a retrieval date: when you or a tool last pulled the number.
COT positions describe a Tuesday but arrive later in the week. A 13F describes a quarter-end portfolio but may become public up to 45 days later. Margin balances describe a month-end. A Form 4 insider filing is much timelier than a 13F, but it still reflects one reported transaction and not a live portfolio. The data latency and vintages page works through this in detail, including why a value can change after its first publication.
A careful reader checks the dates before reading the value. If the observation date is missing, the number cannot be placed on a timeline, and any comparison with price action is guesswork. A good habit is to write the three dates beside any figure you plan to reference.
Does "primary" mean "perfect"?
No. Official sources are the closest inspectable record, not a guarantee of correctness. The SEC says, on its insider-transactions data page, that because the data derive from information provided by individual filers it cannot guarantee their accuracy, and that the data sets are not a substitute for the filings themselves. Cboe's disclaimer says something similar. Filers can submit late, amend earlier entries, or make errors.
The right framing is that a primary source gives you the clearest accountability trail. If something looks odd, you know where to go and whom to ask. A secondary website that republishes the same figure may be faster or prettier, but it adds one more place for mistakes to enter and removes your ability to check the original easily.
What is the difference between source confidence and signal confidence?
These are separate questions, and mixing them up is a quiet source of overconfidence.
Source confidence asks whether the number is a faithful record of what it claims to measure. Signal confidence asks whether that number is informative about something you care about, such as future returns or risk. The CFTC positioning data can rate very high on source confidence (a regulator collected it with clear definitions) while the idea that extreme positioning predicts a reversal can rate low on signal confidence, because the historical relationship is inconsistent and depends on market and period.
The reverse also happens. A social-media measure may be a messy record of what people said, yet still point to a real narrative surge that moved a small asset. The data were poor but the signal was real that time.
A reader who keeps the two apart can say, "This is a reliable measurement of positioning, and I remain unsure what it implies." That sentence is more useful than either blind trust or blanket dismissal.
How do related indicators double count the same information?
Many dashboards give a false sense of breadth because several indicators come from the same underlying market. The VIX, put/call ratios, option skew, and the gap between implied and realized volatility all use inputs drawn from options markets. A composite that counts them as five separate votes is really counting one market several times.
A cleaner approach is to group indicators by what kind of information they carry and think about the groups first:
- Options and volatility
- Futures positioning
- Credit
- Leverage
- Breadth and price participation
- Surveys
- Fund flows
- Insiders and institutions
- Social and alternative data
Agreement across groups is stronger evidence than agreement within one group, because separate groups draw on separate sources of information. The sentiment composite framework shows how to weight groups, and the article on combining breadth, volatility, and sentiment without double counting walks through concrete cases. For the distinction between opinion and participation, see sentiment versus breadth and the overview of positioning, sentiment, and breadth.
How does manipulation risk change the ladder?
The more a measure depends on voluntary, anonymous, or unaudited input, the easier it is to distort. That is why social data sit at the bottom. Coordinated campaigns, automated accounts, paid promotion, and plain duplication can inflate message counts and polarity scores without any change in genuine investor opinion.
A reader assessing a social or alternative sentiment measure can ask whether it discloses:
- How concentrated activity is among a small number of authors
- How many posts are reposts or near-duplicates
- The age mix of the accounts involved
- Sudden spikes in message volume
- How much activity comes from a single platform
- Whether the access method has changed
If none of those are reported, the score may be measuring spam intensity as much as opinion. Exchange and regulator data are far harder to move this way, which is a large part of why they rank higher.
What does licensing have to do with data quality?
A high-quality source can still come with limits on what you can do with it. The ICE credit spread example above is typical: the numbers are professionally produced, but redistribution is restricted. This does not make the source unreliable. It does change how it can be used. Someone building a chart may be able to show a derived percentile and link to the licensed original, but may not be allowed to republish the raw series.
Licensing matters for a second reason. Restrictions can limit how much history you can retrieve, as with the short FRED window noted above. A reader comparing two cycles should check that the history actually reaches back far enough to include both.
What happens when two sources disagree?
Two credible sources can publish different numbers for what looks like the same thing. The first reaction should not be to average them or to pick the one that fits a view.
A careful approach checks, in order: the universe (which products or markets are included), the timestamp, the unit, any adjustment, and the revision policy. For example, an exchange may publish option volume by product category, while a vendor publishes a market-wide put/call series with different inclusion rules. Those are different measurements, not contradictory ones. Presenting both with labels is accurate. If one is clearly stale, say so. If an unresolved difference remains after the checks, saying so is better than presenting a tidy figure that hides it.
The habit to build is asking "same definition?" before asking "which is right?"
How should citations match claims?
A single citation rarely does every job. Three kinds of claim need three kinds of support:
- A claim about what a dataset is or how it is defined should cite the organization that owns the dataset. A statement about the CFTC's trader categories should cite the CFTC; a statement about the 13F deadline should cite the SEC; a statement about FINRA margin reporting should cite FINRA.
- A claim about what an academic study found should cite that study, with its sample and period, rather than a summary of it.
- A claim about how a particular transformation was computed should cite the document that describes the calculation.
When these are blurred, a secondary article ends up standing in as the authority for a regulatory definition, and errors travel from page to page. Separating definitional authority from interpretive evidence makes a claim easier to check and easier to correct.
Research papers deserve their own caution. A paper's conclusion depends on its sample, period, method, and publication status, and a finding from one market and era is not a timeless rule. When studies disagree, the honest summary says that and explains the methodological differences where they are known.
A worked example with illustrative numbers
Suppose a reader sees three claims in one morning and wants to rank them. The numbers below are invented for illustration and are not real market readings.
- A website shows a "retail fear score" of 71.38 out of 100, built from scraped social posts.
- A weekly regulator report shows that a category of large speculators moved from a modest net long position to a larger net short position in a futures contract.
- A monthly leverage statistic reports that investor margin balances rose for a third straight month.
Using the ladder, the reader notes the following.
The fear score is Level 5, or Level 4 if derived from a documented formula. Its decimal places imply more precision than the method can support, since scraped social data have unknown sampling and bot exposure. Its observation window is unclear. The reader treats it as a hint about narrative and looks for a stated methodology before leaning on it.
The positioning report is Level 1. Its provenance is strong and its definitions are published. Its timeliness is limited: the positions describe a past Tuesday, so a market move since then is not in the data. Coverage is partial, because only reportable positions are included. The reader describes it as "what reportable traders held on that date" and nothing more.
The margin statistic is also Level 1, with a month-end observation date. It is slow, but it measures an actual balance rather than an opinion. It says how much borrowed money was in margin accounts at that time, not what investors will do next.
Now the reader asks the cross-family question. These three items come from social data, futures positioning, and leverage, which are three different information families. If they point the same way, that is more meaningful than three options-based indicators pointing the same way. If they conflict, the conflict itself is information: perhaps narrative moved faster than positioning. Either way, the reader has not been told what to do. The ladder improved the quality of the question, which is its job.
What does the ladder not tell you?
The ladder is a tool for judging evidence quality, and it is easy to ask it for more than it can give.
- It does not predict returns. A high-level source can describe a crowded trade without telling you when or whether the crowd will be wrong.
- It does not make an indicator useful. A perfectly documented measure of something irrelevant remains irrelevant.
- It does not replace judgment about context. The same reading can mean different things in different rate, credit, and volatility regimes, which the macro-economics and market regimes section explores.
- It is not a recommendation about any security, fund, or position. It describes the quality of evidence, not what anyone should do with it.
Sentiment research is best used to make assumptions explicit, notice crowded or stressed conditions, and test a thesis against more than one independent source.
A short checklist for any new sentiment number
Before putting weight on a sentiment figure, a careful reader tries to answer these questions:
- Who produced the raw data?
- What population or market does it cover?
- What date does the observation represent, and when did it become public?
- Can it be revised after first publication?
- How are missing values handled?
- Could another person reproduce it from the stated formula?
- Does it overlap with a source already in use?
- Is there a licensing restriction on reuse?
- What question does it help answer?
If the last answer is weak, the indicator may not deserve attention at all. Collecting data is easy. Knowing why you are looking at them is harder.
Where does this fit with the rest of the sentiment pages?
The ladder is a reading aid, so it pairs with the pages that cover specific indicators. For hands-on explanations of how common signals work, start with how sentiment indicators work, then see put/call ratio and options sentiment and social sentiment and retail activity signals for the options and Level 5 cases. The per-indicator pages under sentiment indicators include the Cboe Volatility Index and the COT net position series, and the VIX term structure page covers the options-derived volatility curve. For definitions of unfamiliar terms, the glossary is the quickest reference.
Frequently Asked Questions
Is primary data always better than secondary analysis?
Primary data are usually best for definitions and raw observations, because they come from the body that owns the measurement. Secondary research can still be valuable for interpretation, historical context, and empirical testing. The practical rule is to use each for the claim it is qualified to support: the owner of a dataset for what the data are, and a careful study for what the data may imply.
Can an official filing or statistic be wrong?
Yes. Filings can be late, amended, or incorrect, and the SEC notes that data derived from filer-submitted information cannot be guaranteed accurate. Exchange statistics carry similar disclaimers. Official provenance improves traceability and accountability. It does not eliminate error, so a reader still checks dates, definitions, and whether a figure has been revised.
Why is social sentiment lower on the ladder?
It is more exposed to sampling bias, bots, coordinated activity, platform changes, and unclear population coverage. A FINRA and SEC alert warns that these tools can be inaccurate, incomplete, misleading, or manipulated. That does not make the data useless as evidence of a narrative, but the limitations should be stated beside the number, and it should not be the sole basis for any conclusion.
What makes a derived indicator trustworthy?
Public inputs, a written formula, a stated lookback window, a clear rule for missing data, a version or date for the method, and a path that lets someone else reproduce the result. If any of those are missing, the output is better described as an estimate than as a fact. A derived number should also say which information family each input belongs to, so overlap is visible.
Why do the observation date and publication date matter so much?
Because they decide what a number can legitimately be compared with. A reading that describes a Tuesday cannot be lined up against a Thursday price move without noting the gap. A quarterly holdings report describes one past date, not the present. Without both dates, a reader cannot tell whether a signal was knowable at the time it is being compared to.
What should I do when a source goes offline or changes format?
Treat the series as degraded or unavailable until the issue is resolved, rather than quietly swapping in a lower-quality copy. A substitute can change definitions, timing, revision behavior, and licensing even when the numbers look similar. If a replacement is adopted, it should be documented as a change of source so that later comparisons account for it.
References
- CFTC: Commitments of Traders
- CFTC: Commitments of Traders Explanatory Notes
- Cboe: U.S. Options Daily Market Statistics
- FINRA: Margin Regulation
- SEC: Frequently Asked Questions About Form 13F
- SEC: Insider Transactions Data Sets
- FRED: ICE BofA US High Yield Index Option-Adjusted Spread
- FINRA: Social Sentiment Investing Tools, Think Twice Before Trading Based on Social Media
- Google: FAQ about Google Trends data