Oct 9, 202617 min read

Sources of Alt Data

Equity research used to run almost entirely on documents:

Filings, earnings releases, management guidance, the quarterly call. A growing share of it now runs on evidence that companies never publish about themselves, such as satellite photographs of their car parks, aggregated card spending at their tills, and the jobs they are quietly advertising for. This article is about that shift. Where it has reached in 2026, what it is genuinely being used for, which good ideas are still lying around unused, and where the next generation of signals is likely to come from. One theme runs through all of it. The advantage no longer belongs to whoever buys the data. It belongs to whoever can turn it into a decision before everyone else arrives at the same answer.

Picture an analyst covering a large retail chain. Earnings are three weeks away and management has said nothing. Meanwhile a satellite is photographing car parks across four hundred stores, a payments provider is watching card transactions settle in close to real time, and a web-analytics platform is counting how many shoppers fill a basket in the retailer's app and then walk away from it. Read those three things together and you may already have a rough idea of what the company will report. The rest of the market finds out in three weeks.

This is alternative data, and it is no longer the preserve of a handful of elite quant funds. Where you sit decides what it is worth to you. For a hedge fund or an asset manager it is an independent early read on revenue, which means positioning before the crowd. For a bank research-desk, it is something to hand clients that the desk next door cannot. For risk and credit teams it is an early warning on supply-chain stress. For ESG teams it is a way to check what a company claims against what is observably happening. The shift is the same across the industry; only the angle of advantage changes.

What counts as alternative data

At its simplest, alternative data (or Alt Data - as is sometimes referred) is any information about a company that does not come from the traditional sources:

The annual report, the earnings release, the regulatory filing, the terminal. It is the trail that economic activity leaves behind.

Source

What it actually measures

The question it helps answer

Card and payment transactions

Aggregated spending by merchant or category

Is revenue running ahead of or behind consensus?

Satellite and aerial imagery

Car park occupancy, factory activity, port traffic, crop condition

Is physical activity rising or slowing before the company says so?

Web and app analytics

Site visits, app-store rank, reviews, price changes

Is demand for a product building or fading?

Shipping and logistics telemetry

Vessel positions, port congestion, cargo records

Are inputs arriving on time, and at what cost?

Job postings and employee reviews

Hiring volume, the skills being hired, staff sentiment

Where is the company actually investing next?

Social and search data

Brand attention, consumer mood, emerging controversy

Is sentiment turning before it reaches the numbers?

Sensor and IoT readings

Energy use, emissions, equipment activity

Does reported behaviour match real behaviour?

From niche to normal

Adoption of Alt data has moved quickly, although it pays to be careful about which numbers measure what.

In a survey of 107 investment firms run by the law firm Lowenstein Sandler in late 2025, 90% of respondents said they use alternative data or plan to, against 67% a year earlier and 62% the year before that. It is a self-selected sample weighted towards hedge funds, private equity and venture capital, so it is not a measurement of the whole industry. Even read cautiously, close to thirty points of movement in two years looks like a trend rather than survey noise. 89% of the same group expected to raise their alternative-data budgets during 2026.

The money is following. Investment managers paid data vendors about US$2.8 billion in 2025, up 17% in a year, and Neudata expects that to reach about US$23 billion by 2030. Count corporate and government buyers as well and the market looks far bigger again Grand View Research puts it at about US$136 billion by 2030.

Neudata counted more than 2,800 datasets on offer in 2025, yet the average one is now used by about twenty investment clients, down from twenty-five.

More supply, choosier buyers:

The difference between firms is shifting from which feeds they buy to what they do with them.

The market's plumbing is moving the same way. Bloomberg has put alternative-data entitlements on the Terminal, with metrics from providers like Placer.ai and Similarweb, sitting inside the workflow analysts already use. Once a category reaches the Terminal, it is no longer a specialist add-on. North America still leads on adoption. Asia-Pacific is growing fastest, helped by mobile-first economies and the sheer volume of digital activity they throw off.

Three forces sit behind this, and none of them looks likely to reverse:

Everything leaves a trace. Almost every economic act now creates a digital record. The businesses that collect those records and sell them on have grown into a real industry.

AI made the messy part affordable. Access to data was never the hard bit. Pulling a usable signal out of noisy, incomplete input was, and machine learning has changed what that costs.

The edge decays. Alpha, meaning return above the broad market, fades as information spreads. A two-week head start on consumer traffic is worth paying to defend.

What it is being used for today

Predicting earnings surprises. The most established use. Map alternative signals onto a company's key numbers, compare the result with the Street consensus, meaning the average of what analysts have published, and look for the gap. Card spending combined with car-park imagery gives a workable revenue estimate for a listed retailer weeks before the quarter closes. Practitioners who do this well, report success rates on the direction of the surprise that are comfortably better than a coin flip. Current estimates are of 65% to 70% accuracy in forecasting with 2% to 4% of return captured around the actual announcement. This blend of judgment-led fundamental work and quantitative technique has a name in the industry, the quantamental approach, and the teams that practice it usually put analysts and data scientists in one pod rather than in separate departments.

How you know the model is any good. This is where most of the honest work sits. You rebuild history. Stand the model at each past earnings date and feed it only the data you would actually have had that day, with no sight of what came later.

Then ask three things of every call it made:

Did it point the right way against consensus, how close did it land to the reported number, and, the one that counts, was it closer than the consensus figure itself? A model that cannot beat the consensus it is meant to improve on has not earned its subscription. Test it next on stretches of history it never saw while it was being built, then run it quietly alongside a live earnings season before anyone trades on it. Clear those hurdles and a claimed hit rate stops being a claim.

Watching supply chains and margins. When shipping lanes clog or ships sit longer in port, the cost lands on corporate margins, and companies rarely report it while it is happening. Vessel tracking through AIS, the transponder system ships broadcast, together with cargo and bill-of-lading records, the shipping documents that log what is moving where, lets an analyst watch inventory build and lead times stretch as it happens. For anyone covering electronics, industrials or any business with a long global supply chain, that changes the quality of a margin forecast.

Reading a launch in real time. The first few days of a product's life tell a story that management guidance will not capture for months. Web traffic, app-store ranking, review volume and the shape of social chatter assemble into a live picture of traction, and an analyst who sees it can adjust position before consensus catches up.

Checking ESG claims against behaviour. Probably the most underrated application. Satellite imagery picks up deforestation, pollution plumes and expanding mining footprints that nobody has disclosed, employee-review data surfaces labour problems, and social listening catches a controversy while it is still small. As ESG work moves from box-ticking towards genuine portfolio risk, the ability to observe corporate behaviour independently, rather than accept the reported figure, becomes a real difference in quality.

Ideas still lying around

The quiet signal in hiring. Job-posting volume is already tracked. The subtler read is which skills a company is hiring for. A traditional manufacturer quietly advertising for machine-learning engineers and supply-chain software architects is announcing a strategic shift months before any investor presentation does. Mapping skill-level hiring to expected capital spending or product roadmaps is a genuinely differentiated signal.

Shipping data at the level of the individual input. Most supply-chain work looks at the whole picture, whether ports are jammed and lanes backed up. The unused opportunity is to zoom in and follow the specific raw materials a company depends on, ship by ship and port by port. Say for a specialty chemicals firm, a car-parts maker or a food processor manufacturer, the price of certain specific raw materials is one of the largest drivers of its margin. Tracking those materials in transit shows how much the company is buying, whether its supply is steady, and whether it has quietly switched to a costlier supplier or a longer route. Now, layer these volumes on top of known market prices and you get an independent read on input costs, often before the company reports them, which means spotting margin pressure while it is still forming.

Property footprint as a leading indicator. Commercial property data, meaning lease signings, vacancies, store openings and closures, is public but rarely built into equity models systematically. A retailer signing ten leases in high-footfall suburban malls is saying something about its ambitions that management commentary has not caught up with. A company quietly shrinking its office space across several cities is signalling a cost restructure before it appears in operating expenses.

The same signals, pointed at credit quality. Almost all alternative data in equity research is aimed at revenue and earnings. The same inputs, meaning payment behaviour, supply-chain stress and staff attrition, can be layered into credit-quality work. For an analyst covering companies with heavy debt loads, folding alternative signals into credit-spread models bridges equity and fixed-income research.

Further out: the body as a data source

The ideas above are available now. It is worth looking five to ten years ahead as well, because the most powerful sources of the next decade may come from somewhere most analysts are not yet looking, point in case - the human body. Two frontiers stand out. Both are promising, and both are ethical minefields, which is precisely why they deserve thought now rather than a scramble later.

Reading demand at its source. Neuromarketing, which uses brain imaging, eye tracking, skin-response sensors and emotion-recognition software to measure how people actually react to a product or an advert, is still a modest industry. Market researchers size it somewhere between US$2 billion and US$4 billion, a range wide enough to tell you the category is not yet properly defined, and it mostly happens inside laboratories. What is changing is the hardware. Non-invasive wearable brain-computer interfaces, the EEG headbands sold by firms such as Emotiv and Neurable, along with consumer AR and VR headsets, are moving the capability out of the lab and into ordinary life. Precedence Research puts the brain-computer interface market at about US$3.3 billion in 2026, growing towards US$14 billion by 2035.

Extrapolate from there. As biometric sensing becomes ambient, built into watches, rings, earbuds and glasses, anonymised and aggregated readings of genuine subconscious engagement, meaning attention, emotional arousal and the brain's reward response, could become a leading indicator of which products, games, shows or brands are about to catch on. The appeal is straightforward. People misreport their preferences in surveys constantly, but physiology is much harder to fake. An analyst covering consumer goods, media or gaming could one day watch an aggregated engagement index and get the earliest possible read on a launch. A hint of the next viral product weeks before it shows up in sales data or even in social chatter.

Data this personal raises serious privacy questions, and the rules governing it are only now being written. Consent and responsible use have to be designed in from the start rather than bolted on later.

The body as a forward indicator. The second frontier is physiological data at population scale. Wearables and connected health devices, from continuous glucose monitors to sleep trackers and fitness rings, already produce a continuous picture of how millions of people move, rest and metabolise. Aggregated and anonymised, a measurable shift in activity, sleep or metabolic patterns across a population could become a leading indicator for whole categories: pharmaceutical demand, medical devices, wellness, food and drink, even insurance underwriting.

Biotechnology could push it further. At-home health tests, DNA kits and gut-health analysis are becoming ordinary purchases, and pooled and stripped of names, the data they generate could act as an early signal for healthcare: how a new drug is being received, or which health trend is about to take hold, before official results or sales figures arrive. Picture a sudden coordinated jump in purchases of a new supplement, backed by evidence that people are actually still using it. That could flag a breakout brand a full quarter before it reaches anyone's earnings model. The conditions are the same as before: consent, anonymity and responsible use, built in from the beginning.

Neither of these is a next-quarter trade. They are directional bets on where the information advantage migrates next, and that is the point. By the time a data source is obvious, the edge has gone. The question is not whether these signals will matter, but whether a firm will have built the responsible infrastructure to use them when they do.

Why most Alt data implementation programmes stall

The obstacles are well known, and they sit in the same place:

Not in the data, but in everything that has to happen to it. Raw alternative data is messy. The work is cleaning it, for example - resolving entities so that “Apple Inc.”, “AAPL” and “Apple” are recognised as one company, mapping it to a particular firm's numbers, and making it comparable with consensus. That needs data engineering and capital-markets judgment in the same room, a rarer pairing than it sounds.

Compliance is the second constraint and a serious one. Web-scraping law, personal data handling and, above all, the risk of accidentally acquiring material non-public information, or MNPI, mean governance cannot be bolted on at the end. MNPI risk ranked among the top concerns across every asset class in the Lowenstein survey. Cost is the third: premium datasets can run into seven figures a year, and investment committees increasingly want evidence rather than enthusiasm, meaning a better hit rate, an earlier or timely exit as a result of using the data sets.

Underneath all the three constraints mentioned above, sits something firms rarely name. Most treat alternative data as a purchasing decision. They find a clever dataset, negotiate the licence and wait for the edge to appear. What decays is not the data but the edge, and it decays fastest when a signal sits in a data lake for six weeks while engineering, compliance and the research desk work out who owns it. By the time it reaches an analyst's screen, the earnings event it was meant to anticipate has passed. The firms that get value out of alternative data are not the ones with the better satellite. They are the ones that have shortened the distance between a signal arriving and a decision being taken.

Put positively, the firms that sustain an advantage have usually built five things that reinforce each other, a signal-to-alpha chain if it needs a name: dependable ingestion, signal engineering that turns raw input into clean company-mapped measures, models validated against history, integration into the tools analysts already have open, and honest measurement of what the programme produced. Most stop after the first two, and end up owning the data without owning the capability.

Where Zensar fits

As we saw, most of the firms do not have an awareness problem. The senior people at any hedge fund or bank research desk can already explain why alternative data matters. The difficulty turns up on a Monday morning, when the satellite feed and the card-spend feed will not talk to each other, compliance has not signed off, and nobody owns the job of turning either of them into something an analyst can put into a model by Thursday.

That gap between knowing and doing is where Zensar works, and not by selling another dataset. The value is in the unglamorous machinery around it: ingestion pipelines and entity resolution that turn changing vendor feeds into a reliable stream; expectation-gap models that compare an alternative-data estimate against published consensus and flag the differences worth acting on; the governance layer of sourcing policies, MNPI controls and audit trails that a regulated firm cannot operate without; and independent vendor evaluation, so a subscription is bought on tested signal quality rather than on reputation. The common thread is capital-markets context sitting next to data engineering, because either one on its own tends to produce something that works in a demonstration and not in a research process.

If you are starting

Start with one real question, not a strategy. “How do we use alternative data” goes nowhere. “Can we build a better revenue estimate for our five highest-priority names using transaction and web data” is answerable, with a testable hypothesis and a clear measure of success. Narrow beats grand.

Fix the plumbing before falling in love with the model. The temptation is to jump straight to the clever algorithm. The value of alternative data is decided almost entirely by how well it has been cleaned, mapped and tied to company-specific context, and skipping that is the most common reason programmes quietly fail.

Treat compliance as a design input, and measure from day one. Governance added at the end becomes the thing that stops the launch. It is recommended to build the governance framework along with the pipeline, starting from day one.

The bigger picture

Alternative data is not a passing trend. It is a change in how information reaches markets and gets priced into them, and the institutions that build a systematic capability around it will compound their advantage. Those that do not will find themselves working from a slower and narrower information set, a disadvantage that has nothing to do with how clever their analysts are.

The window is still open. Adoption is rising fast, but the gap between the early movers and everyone else remains wide. For hedge funds, bank research divisions, sell-side teams and brokers alike, the question is no longer whether to engage with alternative data but how to do it well. And that deserves more than a vendor subscription. It deserves a strategy.

Authored by

Chetan Kothavade

Lead Business Consultant, Capital Markets, Zensar Technologies

Let's talk

Please fill out the form to get in touch with us.