Research Automation4 min read3 Oct 2026

Building a Research Data Pipeline from NSE and BSE Filings

A blueprint for a research data pipeline on Indian listed companies: what filings exist, when they land, how to store and verify them, and what it unlocks.

Short answer

A research data pipeline for Indian equities collects exchange filings (results, annual reports, shareholding patterns, transcripts and announcements) on a schedule. It stores the raw documents unchanged, extracts structured data into a database with source references, verifies material figures, and feeds models, screens and alerts. Use licensed vendors or official exchange channels for data, and follow each source's terms of use.

Key takeaways

  • Indian disclosure runs on a predictable calendar. Design the pipeline around it.
  • Store raw documents permanently and unchanged. Every extracted number should point back to a file and page.
  • Standalone vs consolidated, restatements and changing segments are the three things that break naive pipelines.
  • Licensed data is often cheaper than it looks once you count the cost of maintaining your own collection scripts.

Every institutional research process starts with data. For Indian equities the raw material is unusually rich: regulator-mandated, standardised and published on a predictable calendar. Most of the work in a data pipeline is not getting the data. It is keeping it correct, traceable and comparable over time.

This guide sets out a blueprint you can build in-house or hand to a partner.

What gets filed, and when

Indian listed companies file on a predictable rhythm set mainly by SEBI’s Listing Obligations and Disclosure Requirements (LODR) regulations.

The disclosure calendar around each quarter
The disclosure calendar around each quarterDay 0Quarter endsThe clock starts for shareholding patterns and results.≤ Day 21Shareholding pattern filedPromoter, institutional and public holdings, including encumberedpromoter shares.≈ Day 15–45Quarterly results seasonResults within 45 days of quarter end (60 days for the auditedfull-year results after Q4). Most mid and large caps report in themiddle weeks.Within daysEarnings call recordings and transcriptsInvestor presentations are filed around the results. Call recordingsand transcripts follow within a short window set by the disclosurerules.ContinuousCorporate announcementsBoard meetings, credit ratings, acquisitions, management changes,pledges and insider trades are filed as events occur.AnnuallyAnnual reportFull statements, notes, auditor's report and governance report, sentto shareholders ahead of the AGM.
Deadlines summarised from SEBI LODR at the time of writing. Rules are amended periodically, so verify current timelines before relying on them operationally.

Pipeline architecture

Blueprint: from filings to research outputs
Blueprint: from filings to research outputsExchange filings(results, XBRL,announcements)Annual reports andtranscripts (PDF)Licensed vendor historySourcesUnchanged originals,hashed and datedNever overwritten;restatements are newversionsRaw storeFinancials: standalone+ consolidated, perperiodOwnership: promoter,pledge, FPI, DIIText facts: guidance,RPTs, KPIs, with sourcepageStructured layerModel history updatesScreens and peer tablesAlerts and change logsOutputsResearch datapipeline
Four layers. The raw store is the most important design decision: if you keep originals, every error downstream can be traced and fixed.

The data you need, by type

Data type Typical source Frequency Main gotcha
Quarterly P&L Results filing (PDF + XBRL) Quarterly Standalone vs consolidated; restated prior periods
Balance sheet & cash flow Half-yearly / annual filings Half-yearly / annual Limited quarterly balance-sheet detail
Segment data Results and annual report notes Quarterly / annual Segment definitions change
Shareholding & pledges Shareholding pattern filings Quarterly + events Category definitions; encumbrance types
Guidance & KPIs Transcripts, presentations Quarterly Unstructured; needs extraction and verification
Corporate actions Announcements Event-driven Splits/bonuses must adjust per-share history
Prices & volumes Exchange or vendor feed Daily Corporate-action adjustment

The three things that break naive pipelines

1. Standalone vs consolidated

Indian companies report both. Many PDF tables put them on adjacent pages with identical layouts. A pipeline that does not record which one it extracted will eventually mix them, and margins and growth rates will be quietly wrong. Make “basis” a required field on every financial fact.

2. Restatements

When a company restates prior periods after a merger, demerger or accounting change, the “same” quarter now has two values. Never overwrite. Store both, with the filing each came from, and choose explicitly which one a model uses.

3. Changing segments

Segments are reorganised more often than you would expect. Map old segment names to new ones in a lookup table, and flag any period where the mapping isn’t clean.

Verification: the step most pipelines skip

Illustrative: where extraction errors come from
Illustrative: where extraction errors come fromStandalone / consolidatedmix-up31%Wrong period or column24%Units (lakh vs crore vsmillion)17%Restated vs original figures12%Sign errors (expenses, otherincome)9%Genuine OCR / parsing errors7%
Illustrative distribution based on common failure modes in PDF and table extraction. Most errors are structural, not OCR, so they can be caught with simple rules and spot-checks.

Simple automated checks catch most of these before a human ever looks:

  • Arithmetic identities. Segment revenues sum to total. The balance sheet balances. Quarterly figures sum to annual.
  • Unit sanity. Flag any value that differs from the prior period by more than 10×.
  • Basis consistency. Alert when consolidated revenue is lower than standalone, which is possible but rare.
  • Cross-source match. Compare the XBRL value with the PDF-extracted value where both exist.

Whatever is still flagged goes to an analyst, who checks it against the source page in seconds because the pipeline stored the reference.

Build, buy or blend?

Approach Good for Watch out for
Licensed vendor only Fast start, deep history Cost per seat; less control over text data
Build on exchange filings Timeliness, text data, custom fields Maintenance; terms-of-use compliance
Blend Most serious desks Reconciling the two sources

Most teams end up blending: vendor data for deep, clean history, and their own extraction for timely and text-based data such as guidance, related-party transactions and pledge events.

What it unlocks

Once the data layer is reliable, the rest of the research stack becomes cheap. Screens run in seconds. Pledge monitors update themselves. Transcript trackers fill as calls are filed. Analysts can spend the week on the judgement work described in our guide to AI and research automation.

Frequently asked questions

Where can I get financial data for Indian listed companies?

Primary sources are the NSE and BSE corporate filing pages, which carry results, annual reports, shareholding patterns and announcements, including XBRL-tagged financial filings. Licensed vendors such as institutional terminals and Indian database providers offer cleaned historical data. Many teams combine a vendor for history with exchange filings for timeliness.

When do Indian companies publish quarterly results?

Under SEBI's listing regulations, listed companies generally must publish quarterly results within 45 days of quarter end, and annual audited results within 60 days of the financial year end. That puts results seasons roughly in the six to eight weeks after each quarter closes.

Is it legal to scrape NSE or BSE websites?

Each website has its own terms of use, and some restrict automated access. Before automating collection, read the terms, prefer official data products or licensed vendors, rate-limit any permitted access, and take legal advice if in doubt. This guide does not recommend bypassing any access controls.

What is XBRL in Indian financial filings?

XBRL is a standard format for tagging financial statement data so machines can read it. Indian listed companies file certain financial results and disclosures in XBRL with the exchanges, which makes extraction more reliable than parsing PDFs.