Key takeaways
- Indian disclosure runs on a predictable calendar. Design the pipeline around it.
- Store raw documents permanently and unchanged. Every extracted number should point back to a file and page.
- Standalone vs consolidated, restatements and changing segments are the three things that break naive pipelines.
- Licensed data is often cheaper than it looks once you count the cost of maintaining your own collection scripts.
Every institutional research process starts with data. For Indian equities the raw material is unusually rich: regulator-mandated, standardised and published on a predictable calendar. Most of the work in a data pipeline is not getting the data. It is keeping it correct, traceable and comparable over time.
This guide sets out a blueprint you can build in-house or hand to a partner.
What gets filed, and when
Indian listed companies file on a predictable rhythm set mainly by SEBI’s Listing Obligations and Disclosure Requirements (LODR) regulations.
Pipeline architecture
The data you need, by type
| Data type | Typical source | Frequency | Main gotcha |
|---|---|---|---|
| Quarterly P&L | Results filing (PDF + XBRL) | Quarterly | Standalone vs consolidated; restated prior periods |
| Balance sheet & cash flow | Half-yearly / annual filings | Half-yearly / annual | Limited quarterly balance-sheet detail |
| Segment data | Results and annual report notes | Quarterly / annual | Segment definitions change |
| Shareholding & pledges | Shareholding pattern filings | Quarterly + events | Category definitions; encumbrance types |
| Guidance & KPIs | Transcripts, presentations | Quarterly | Unstructured; needs extraction and verification |
| Corporate actions | Announcements | Event-driven | Splits/bonuses must adjust per-share history |
| Prices & volumes | Exchange or vendor feed | Daily | Corporate-action adjustment |
The three things that break naive pipelines
1. Standalone vs consolidated
Indian companies report both. Many PDF tables put them on adjacent pages with identical layouts. A pipeline that does not record which one it extracted will eventually mix them, and margins and growth rates will be quietly wrong. Make “basis” a required field on every financial fact.
2. Restatements
When a company restates prior periods after a merger, demerger or accounting change, the “same” quarter now has two values. Never overwrite. Store both, with the filing each came from, and choose explicitly which one a model uses.
3. Changing segments
Segments are reorganised more often than you would expect. Map old segment names to new ones in a lookup table, and flag any period where the mapping isn’t clean.
Verification: the step most pipelines skip
Simple automated checks catch most of these before a human ever looks:
- Arithmetic identities. Segment revenues sum to total. The balance sheet balances. Quarterly figures sum to annual.
- Unit sanity. Flag any value that differs from the prior period by more than 10×.
- Basis consistency. Alert when consolidated revenue is lower than standalone, which is possible but rare.
- Cross-source match. Compare the XBRL value with the PDF-extracted value where both exist.
Whatever is still flagged goes to an analyst, who checks it against the source page in seconds because the pipeline stored the reference.
Build, buy or blend?
| Approach | Good for | Watch out for |
|---|---|---|
| Licensed vendor only | Fast start, deep history | Cost per seat; less control over text data |
| Build on exchange filings | Timeliness, text data, custom fields | Maintenance; terms-of-use compliance |
| Blend | Most serious desks | Reconciling the two sources |
Most teams end up blending: vendor data for deep, clean history, and their own extraction for timely and text-based data such as guidance, related-party transactions and pledge events.
What it unlocks
Once the data layer is reliable, the rest of the research stack becomes cheap. Screens run in seconds. Pledge monitors update themselves. Transcript trackers fill as calls are filed. Analysts can spend the week on the judgement work described in our guide to AI and research automation.
Frequently asked questions
Where can I get financial data for Indian listed companies?
Primary sources are the NSE and BSE corporate filing pages, which carry results, annual reports, shareholding patterns and announcements, including XBRL-tagged financial filings. Licensed vendors such as institutional terminals and Indian database providers offer cleaned historical data. Many teams combine a vendor for history with exchange filings for timeliness.
When do Indian companies publish quarterly results?
Under SEBI's listing regulations, listed companies generally must publish quarterly results within 45 days of quarter end, and annual audited results within 60 days of the financial year end. That puts results seasons roughly in the six to eight weeks after each quarter closes.
Is it legal to scrape NSE or BSE websites?
Each website has its own terms of use, and some restrict automated access. Before automating collection, read the terms, prefer official data products or licensed vendors, rate-limit any permitted access, and take legal advice if in doubt. This guide does not recommend bypassing any access controls.
What is XBRL in Indian financial filings?
XBRL is a standard format for tagging financial statement data so machines can read it. Indian listed companies file certain financial results and disclosures in XBRL with the exchanges, which makes extraction more reliable than parsing PDFs.