How Geopits Built a Real-Time Data Platform for a Leading NBFC

Thirunavukkarasu Ramasamy

vector

For an Indian NBFC, getting a clear picture of loan disbursals, collections, and risk exposure was not always possible on the same day. Important information was scattered across core lending systems, customer relationship management platforms, payment gateways, credit-bureau APIs, collections tools, and file-based reports. Because of this, MIS reporting was delayed; business teams depended a lot on IT for data requests, and decision-makers often used old information.

To fix this problem, Geopits created an NBFC data platform that brought together 17 data sources into a governed, near-real-time analytics environment. With a GCP BigQuery BFSI setup, automated data pipelines, and self-service dashboards, the platform created a single source of truth for lending, collections, customer, and operational data. This allowed real-time analytics for NBFCs and built a strong base for future data engineering and reporting needs.

The Business Challenge: Fragmented Data Slowed NBFC Decision-Making.

The NBFC needed faster access to trusted data across lending, collections, and risk operations. However, information was distributed across legacy databases, SaaS applications, third-party APIs, and file-based reporting systems. Each platform captured a different part of the customer and loan journey. The loan origination system stored application and disbursal data, while CRM platforms tracked customer interactions and sales activity.

Payment systems recorded EMI collections and repayment status. Credit bureaus and KYC platforms provided important information for risk assessment and customer onboarding. These systems did not always update at the same speed. Some data arrived through APIs, while other information came through scheduled files or manual uploads.

As a result, MIS reports often required manual compilation. Business teams depended on IT and analysts for recurring reports and ad hoc data requests. This delayed visibility into repayment performance, overdue accounts, collections activity, and portfolio risk. It also made it harder for teams to respond quickly to changing business conditions.

The NBFC needed a governed data environment to improve auditability and compliance. It required clear data lineage, controlled access, and consistent definitions across every report.

Real-Time Credit Underwriting: Eliminating Disbursal Delays

Manual credit checks and legacy batch processing created a severe bottleneck in the loan origination flow. When evaluation took hours or days, high-intent applicants often abandoned their applications, resulting in lost business to faster competitors.

To solve this, real-time data pipelines were designed to continuously pull and normalize data from credit-bureau APIs, instant UPI transactions, and GST/ITR feeds into high-throughput databases like Google BigQuery and PostgreSQL. Replacing delayed batch updates with real-time data streams allowed automated underwriting engines to run instant risk checks, bringing loan approval and disbursal turnaround times down from days to under 60 seconds.

The NBFC also needed a governed data environment to improve auditability and compliance. It required clear data lineage, controlled access, and consistent definitions across every report.

The NBFC Data Landscape: Unifying 17 Disparate Sources

The NBFC’s data environment included 17 sources across databases, SaaS platforms, APIs, legacy reporting systems, and raw files. Each source captured valuable information, but the systems differed in structure, ownership, update frequency, and integration method.

Unifying these sources required more than connecting databases. The platform also had to process API responses, scheduled SFTP feeds, Excel reports, scanned documents, call recordings, communication logs, and field-agent data. The following categories show the breadth of data involved.

Core databases

  • Oracle: Loan origination system data, including applications and disbursals.
  • SQL Server: CRM and branch-operations data.
  • MongoDB: Mobile collections data and customer-interaction records.

SaaS and third-party applications

  • CRM platform: Customer lifecycle and sales-pipeline information.
  • Credit-bureau APIs: Credit-risk and borrower-assessment data.
  • eKYC and KYC platform: Identity verification and onboarding-compliance data.
  • Payment gateway: EMI collections and repayment transactions.
  • Collections dialer or telephony platform: Agent-call metadata and contact history.
  • HRMS or field-agent application: Agent productivity and collection-performance information.

Legacy reporting systems

  • Existing on-premises enterprise data warehouse: Historical MIS and portfolio data.
  • Legacy batch-reporting layer: Scheduled reporting workflows being phased out.

Raw and unstructured data

  • Scanned loan applications and KYC documents in PDF or image formats.
  • SFTP flat-file feeds from bureaus and lending partners.
  • Collections and customer-service call recordings.
  • Branch-level Excel and CSV MIS reports.
  • SMS and WhatsApp gateway logs.
  • Field-agent GPS and geo-tagging data for collection-visit verification.

What data sources do NBFCs need to unify for real-time analytics?

For real-time analytics, NBFCs must connect data from every stage of the lending and collections lifecycle. Geopits unified 17 sources, including structured databases, SaaS platforms, APIs, legacy reports, and unstructured files. This created a common data foundation for lending, collections, risk monitoring, and operational reporting. citefile:1

  • Loan origination systems: Geopits connected core loan data, including applications, approvals, disbursals, and borrower details.
  • CRM and branch systems: Customer profiles, sales activity, branch operations, and service interactions were brought together.
  • Collections applications: Mobile collections data and customer-interaction records were included.
  • Credit-bureau APIs: Credit-risk and borrower-assessment data was integrated into the broader analytics environment.
  • eKYC and KYC platforms: Identity verification and onboarding-compliance data were connected.
  • Payment gateways: EMI collections, repayment transactions, and payment statuses were incorporated.
  • Collections dialers: Call metadata and customer-contact history supported collections analysis.
  • HRMS and field-agent applications: Agent productivity, field activity, and collection performance were included.
  • Legacy data warehouses: Historical MIS and portfolio data were brought into the modern reporting environment.
  • SFTP and partner feeds: Scheduled bureau and partner files were processed through automated pipelines.
  • Branch Excel and CSV reports: Local operational reports were standardized for centralized analysis.
  • Loan and KYC documents: Scanned documents and images expanded the platform beyond structured data.
  • Call recordings: Audio data from collections and customer service became part of the broader data landscape.
  • SMS and WhatsApp logs: Customer communication trails helped provide additional engagement context.
  • GPS and geo-tagging data: Field-visit information supported collection verification and operational monitoring.
  • Credit-bureau, UPI, and GST feeds: Real-time data streams from bureau APIs, UPI payment logs, and GST/ITR filings are integrated straight into analytics platforms like BigQuery and PostgreSQL to support dynamic risk scoring.

Geopits used automated data engineering workflows to process these different formats and update the central platform. Apache Airflow coordinated ingestion, while GCP and BigQuery provided the analytical foundation. Apache Superset then presented the unified information through self-service dashboards.

How the GCP and BigQuery Architecture Supported Faster Analytics

  • The platform supports batch and near-real-time data processing across 17 sources using Apache Airflow to coordinate automated pipelines for databases, APIs, SaaS applications, SFTP feeds, and spreadsheets.
  • Google Cloud Platform (GCP) and BigQuery serve as the unified analytical layer to standardize lending, customer, payment, collections, and operational data.
  • Apache Superset enables self-service dashboards, allowing business teams to access recurring reports and ad hoc data independently.

The GCP BigQuery BFSI architecture establishes a scalable foundation for governed reporting, real-time analytics for NBFCs, and future data engineering for NBFC initiatives.

Why This Matters for Other NBFCs

Many NBFCs face similar challenges when building a modern NBFC data platform. Loan origination systems, CRM platforms, payment gateways, credit-bureau APIs, collections applications, and file-based reports often operate in separate environments. This makes it difficult to create a complete and timely view of lending, customer, repayment, and risk data.

Effective data engineering for NBFCs helps connect these fragmented systems without requiring every legacy application to be replaced. By combining structured, semi-structured, and unstructured data in a governed environment, NBFCs can create a stronger foundation for real-time analytics for NBFCs, self-service reporting, and future analytics initiatives.

The approach used in this project can be adapted for other BFSI organizations looking to modernize their data environment. A scalable GCP and BigQuery-based data platform can help organizations move from siloed reporting toward faster, more connected business insights.

How Geopits Built the NBFC Data Platform

Geopits brought the NBFC’s fragmented data into one governed platform for faster reporting and business decisions. The solution supported different source formats and processing requirements without replacing every existing system.

The result was a modern NBFC data platform that supported batch and near-real-time processing. It created a single analytical foundation for real-time analytics, collections visibility, risk monitoring, and future data-engineering initiatives. The implementation approach and technology flow are based on the client-approved outline. 

Business Outcomes: Faster Insights for NBFC Teams

Geopits unified fragmented NBFC data into a governed platform for faster reporting, collections decisions, and operational visibility. Here are the key outcomes:

Faster reporting: The unified data platform helped move reporting away from manual compilation. Final reporting-frequency improvements should be confirmed before publishing.

Less dependency on IT: Business teams could access self-service dashboards for recurring operational and portfolio reports. This reduced the need for IT teams to prepare every data request manually.

Better collections visibility: Collections teams could view customer, payment, call, and field-activity information together. This supported faster and more informed collection decisions.

Stronger audit readiness: Governed data pipelines created clearer visibility into data sources, transformations, and access. This improved traceability for reporting and audit processes.

Scalable analytics foundation: The platform created a reusable foundation for future use cases. These may include risk monitoring, portfolio analytics, collections optimization, and machine-learning initiatives.

Credit bureau, UPI, and GST feeds: Real-time data streams from bureau APIs, UPI payment logs, and GST/ITR filings are integrated straight into analytics platforms like BigQuery and PostgreSQL to support dynamic risk scoring.

Why Geopits for Data Engineering for NBFCs 

Geopits combines domain-focused data engineering with scalable cloud architecture for complex BFSI environments. In this project, Geopits connected 17 fragmented sources and created a governed analytics foundation using Airflow, GCP, BigQuery, and Superset. 

Built for complex data environments

Geopits worked across databases, APIs, SaaS applications, SFTP feeds, spreadsheets, documents, communication logs, and field-agent data. This demonstrates the ability to handle different formats, update frequencies, and integration requirements.

Focused on business outcomes

The platform was designed to reduce manual reporting and improve access to lending, collections, repayment, and risk information. The technology supported clear business goals rather than operating as a standalone technical deployment.

Designed for governed analytics

The solution included data access controls, masking, audit trails, validation checks, and traceable data movement. These controls are important for NBFCs managing sensitive customer and financial information.

Built for future growth

The platform provides a foundation for self-service reporting, portfolio analytics, collections optimization, and future machine-learning initiatives. Geopits’ approach can be adapted as NBFC data volumes, products, and reporting requirements expand.

Conclusion

If fragmented data, delayed reporting, and disconnected systems are limiting visibility across your lending and collections operations, Geopits can help. Our team works with organizations to build scalable NBFC data platforms using modern cloud and data engineering technologies.

Talk to our data engineering team to explore how your organization can build a governed foundation for real-time analytics, faster reporting, and better business decision.

FAQ

1. What is an NBFC data platform?

An NBFC data platform connects lending, customer, payment, collections, risk, and operational data in one governed environment. It helps teams access consistent information for reporting and decision-making.

2. How does a real-time data platform speed up NBFC loan approvals? 

Traditional batch processing delays disbursals by relying on periodic file syncs. A real-time data platform continuously streams data from credit bureau APIs, UPI transactions, and GST feeds into databases like BigQuery and PostgreSQL. This allows automated underwriting engines to run instant risk checks and deliver sub-minute loan approvals.

3. How did Geopits unify the NBFC’s data sources?

Geopits connected 17 sources, including databases, APIs, SaaS applications, SFTP feeds, spreadsheets, documents, communication logs, and field-agent data. Apache Airflow coordinated ingestion and pipeline workflows across these sources. 

4. Can NBFCs modernize data without replacing all legacy systems?

Yes. NBFCs can connect existing systems to a modern data platform while continuing to use core operational applications. This approach supports gradual modernization and reduces disruption to lending operations.

5. How does data engineering support real-time analytics for NBFCs?

Data engineering connects and standardizes data from databases, APIs, SaaS platforms, file feeds, and unstructured sources. In this project, Geopits used Apache Airflow, GCP, BigQuery, and Superset to create governed data pipelines and faster access to lending, collections, and risk insights.

6. What is the best architecture for an NBFC data platform?

A modern NBFC data platform typically connects data from loan origination systems, CRM platforms, payment gateways, credit bureaus, collections applications, and file-based sources. Automated data engineering for NBFCs can orchestrate these pipelines and bring the information into a unified analytical environment. In this project, Apache Airflow, GCP, BigQuery, and Apache Superset supported the flow from data ingestion to business reporting.

7. How does real-time data engineering accelerate NBFC loan approvals?

By replacing manual updates and scheduled batch processing with continuous data pipelines, data platforms stream credit bureau scores, UPI transactions, and GST data straight into analytical engines. This allows automated risk evaluation and cuts loan approval turnaround times from days to under 60 seconds.

We run all kinds of database services that vow your success!!