TL;DR
This is Part 6 of the Building the AI-First Lender series. Each part covers one function of the AI-first lender as a data architecture — not a feature list. Read Part 1,Part 2,Part 3, Part 4, Part 5 for the full context.
This series has traced the AI-first lender as a data architecture: origination, underwriting, credit decisioning, servicing, and disbursement. Collections is where every upstream decision either pays off or compounds into loss.
An AI-first collections stack runs on propensity-to-pay models, not call queues. It segments borrowers into self-cure, needs nudge, settlement candidate, legal track before a single agent is deployed.
The data feedback loop from collections back to the origination model is the single biggest efficiency gain available. Most lenders have not built it.
RBI's Fair Practices Code governs all collections contact — automated or human. The 7 AM to 7 PM contact window, the prohibition on third-party disclosure, and the Grievance Redressal Officer obligation apply whether the trigger is a voice bot or a field agent.
The best collections system is always the one used least. An AI-first lender that closes the data loop from collections back to origination will approve fewer loans that end in collection and that changes the economics of the entire business.
01/WHERE THE LOOP CLOSES
The argument across this series has been consistent: an AI-first lender is not a lender that deploys AI features. It is a lender built around a data architecture where every function from origination to disbursement generates signals that feed back into a continuously improving credit engine.
If you are reading this series for the first time, the short version is this. Most lenders in India still treat credit decisioning, loan servicing, and collections as separate functions with separate teams, separate systems, and separate data. An AI-first lender treats them as one loop. The January 2026 edition of this newsletter, AI Revolution in Collections, covered the tools available propensity scoring, voice bots, settlement logic, omni channel communication. This edition covers the harder architectural question: what does an AI-first lender build differently so that collections is an integrated function rather than a fire-fighting department?
Collections is the proof point for the AI-first architecture. Every credit decision made at origination, every servicing interaction, every Account Aggregator cash flow that was or was not pulled, all of it surfaces in a single number: recovery rate.
02/90 DAY IRACP CLOCK
RBI's IRACP Directions 2025 formalised the NPA classification timeline with no ambiguity. An account becomes overdue the day repayment is missed. At 30 days, it is SMA-1. At 60 days, SMA-2. At 90 days, NPA — provisioning requirements kick in, interest income reversal is mandatory, and the cost to the lender compounds.
That 90-day window is the collections function's operating theatre. Traditional collections teams work it by calling every delinquent account in DPD (days past due) order. The problem is structural: a borrower at 3 DPD who had a payment failure will self-cure by day 5. A borrower at 3 DPD who has lost employment will be at 90 DPD in six weeks. Treating them identically wastes agent time on accounts that need nothing and loses the window on accounts that needed immediate intervention.

Figure 1: RBI’s NPA Classification Timeline
AI-driven propensity-to-pay modelling flips this. The model scores each overdue account at day one of delinquency on the likelihood of self-cure, the probability of recovery given intervention, and the optimal channel and timing for outreach. The inputs are everything the lender already holds: payment history, bureau tradelines, Account Aggregator cash flow data, app engagement signals, and historical collections outcomes on comparable borrowers.
Credgenics, which processes over 11 million retail loan accounts in India, reports a 25% improvement in resolution rates and a 40% reduction in collections costs relative to manual approaches. CreditNirvana and Ezee AI are building similar capabilities — omni channel outreach, predictive scoring, promise-to-pay tracking — for lenders that want this as a managed service rather than a proprietary build. These are vendor claims and should be treated as directional rather than audited benchmarks. The direction, across all of them, is consistent. What varies is how well the underlying model is built and how cleanly it connects to origination data.
A propensity-to-pay model built on thin data is not better than a call queue. It is just a more expensive one.
03/BUILDING THE AI COLLECTIONS STACK
Propensity-to-pay modelling in digital loan collections is the practice of scoring each delinquent borrower at the point of first missed payment on three outputs: probability of self-cure, likelihood of recovery given intervention, and optimal channel and timing for outreach. In India, where Account Aggregator cash flow data and UPI transaction history are now available at origination, these models can draw on real-time behavioral signals rather than static bureau data alone — a material advantage over scoring models in markets where AA-equivalent infrastructure does not exist.
An AI-first lender does not bolt collections onto its credit stack. It designs collections into the architecture from day one. The difference is visible across four areas.

Figure 2: Collections Stack
Borrower segmentation. The first function of the collections stack is segmentation. Self-cure accounts — borrowers who have missed a payment due to technical or timing issues and carry strong repayment signals — need a reminder, not an agent. Accounts that need a nudge require digital outreach: WhatsApp, in-app notification, IVR. Settlement candidates — borrowers with genuine repayment difficulty and a recoverable position — need a structured offer and a human conversation. Legal-track accounts have exhausted the recovery window and require a separate process.
Most lenders do this segmentation manually, if at all. An AI-first lender automates it on day one of delinquency, using the credit and behavioral data it already holds.
Digital collections. RBI's Fair Practices Code and the Digital Lending Directions 2025 set the boundary conditions. Contact is permitted between 7 AM and 7 PM. Disclosure of a borrower's default to third parties — family members, employers, contacts — is prohibited. Every automated communication must carry the regulated entity's identity and the GRO (Grievance Redressal Officer) contact. An AI voice bot is not exempt because it is software. The conduct standards apply to the outcome, not the trigger.
This has a direct implication for architecture. The 7 AM to 7 PM window must be enforced at the system level. The WhatsApp message, IVR call, in-app push — each must be logged, timestamped, and auditable. Any settlement offer communicated digitally must go to the borrower directly, not to a guarantor or reference contact.
Legal collections. Currently, SARFAESI and DRT processes remain outside the scope of AI automation. The legal steps require human judgment and court timelines. What AI can do is flag accounts heading toward legal track earlier, ensure documentation is complete before referral, and track compliance timelines. The gap most lenders have is that their legal collections function operates without the behavioral and cash flow data the credit and collections models hold. Closing that gap is an architectural decision.
The outsourcing boundary. When a lender engages a recovery agency for field collections or tele-calling. RBI's outsourcing guidelines apply in full. The lender remains responsible for agent conduct. If the AI model flags an account for field visit and the agent violates the Fair Practices Code, the lender bears the compliance liability. Similarly if tele-calling is being done by an external voice agent, lender will again will bear the responsibility.
The regulatory boundary in collections applies to the entire chain from model output to last-mile agent. Building AI into collections without building compliance into that chain creates regulatory exposure, not efficiency.
04/COLLECTIONS FEEDBACK LOOP
This is the section that matters most for a lender's long-term economics and the one most lenders treat as a future initiative.
Every collections outcome is a data point about the credit model. A borrower who passed underwriting and landed in SMA-2 within four months is telling you something specific: which variable in the origination model was wrong. Was it the bureau score? The AA cash flow assessment? The loan-to-income ratio? The employer stability signal? Collections outcome data, mapped back to origination variables, is the most accurate model validation data a lender has. It is also the most underused.

Figure 3: Building the Collections Feedback into Origination
The model refresh cycle matters. Credit models are typically retrained quarterly or semi-annually. Vintage analysis — cohort-level portfolio performance tracked over time — gives a sharper signal earlier. If the Q3 FY25 cohort is running at 2x the NPA rate of Q3 FY24 at the same DPD point, something changed in underwriting or in the borrower pool. Catching that shift requires collections outcome data flowing to the credit team in near-real-time, not in a quarterly MIS report.
Most lenders have not built this loop. Collections teams and credit model owners sit in different functions, use different systems, operate on different reporting cadences. The collections team knows which borrower profiles are hardest to recover from. The credit team is retaining origination variables that look predictive at approval but correlate poorly with actual repayment. The disconnect is structural — and largely invisible until a vintage deteriorates.
An AI-first lender solves this at the architecture level. Collections outcome data — segmentation bucket, intervention type, resolution rate, recovery timeline — feeds back into the feature store the credit model draws from. The model learns that a borrower with a particular AA cash flow pattern and bureau profile ends up in the settlement bucket 40% of the time. That signal reduces origination to that profile, or reprices it, or changes the product structure for that segment entirely.
The closed-loop lender does not just collect better. It originates smarter.
05/COMPLIANCE BOUNDARIES
13.34 lakh complaints to RBI Ombudsman in FY25 — up 13.5% year-on-year. Loans and advances: the largest category at 29.25% of total complaints. Private sector banks: 37.5% of all grievances. Source: RBI Annual Report on the Integrated Ombudsman Scheme, 2024–25
The lenders driving these numbers are not legacy institutions with outdated processes. Private sector banks — the ones that expanded fastest into unsecured retail credit — account for the highest share of grievances. The expansion into personal loans, BNPL-style products, and short-tenure credit happened. The collections conduct did not keep up.
The specific prohibitions in the Fair Practices Code are worth stating plainly, because the same lenders deploying AI collections tools are the ones driving complaint volumes. Repeated contact designed to pressure is prohibited — not just abusive language, but frequency. Disclosing a borrower's default to family members, employers, or neighbours is a violation. Threats beyond legitimate recovery process cross the line. These standards apply to AI-driven communication exactly as they apply to human agents.
The GRO obligation adds an audit trail requirement that AI-first collections actually makes easier to meet, if the system is designed for it. Every digital communication is logged. Every call is recorded. Every settlement offer has a timestamp. An AI-first lender that builds compliance-by-default into its collections stack . Some examples: 7 AM to 7 PM enforcement hardcoded, GRO contact embedded in every outreach, disclosure prohibition enforced at the data layer.
DPDP compounds this. The DPDP Rules 2025, alongside RBI's Digital Lending Directions, create a dual compliance regime for collections data. Purpose limitation means data collected for credit assessment cannot be repurposed for collections communications beyond what was disclosed at consent. A digital lending platform that retains loan application data for seven years (per RBI Digital Lending Guidelines) but holds collections behavioral data e.g. call logs, WhatsApp threads, settlement correspondence without a defined retention policy is accumulating compliance risk that will surface as DPDP enforcement develops.
The interaction between DPDP and collections is still being worked through by legal teams across the industry. The principle is settled even if the operational implementation is not: data collected for one purpose at origination cannot be freely reused for collections communications without clear disclosure at the point of collection. Lenders that embedded broad consent language at onboarding to cover all future use cases will face challenge on those terms as enforcement develops.
RBI's ombudsman data describes what is actually happening. DPDP is arriving to constrain what is legally permissible. Lenders building compliant collections architectures now will have lower retooling costs later.
06/SIX FUNCTIONS, ONE ARCHITECTURE
This series has covered six functions: origination, credit underwriting, decisioning, disbursement, servicing, and collections. The argument across all six has been the same. The AI-first lender is defined by its data architecture, not by which vendor tools it deploys. The functions are interconnected. A weakness in origination data quality flows into a weaker credit model, which flows into a higher collections rate, which without the feedback loop never corrects the origination model.
The honest assessment of where AI can operate fully today versus where hybrid approaches remain necessary: origination, credit underwriting, and digital collections can be built AI-first within the current regulatory framework. Decisioning and servicing are close. Legal collections stays human-led. The constraint is not the technology it is that RBI has not yet provided clear guidance on AI-automated settlement authority, and SARFAESI and DRT processes require human actors.
Greenfield lenders building from scratch have one structural advantage over incumbents: no legacy of disconnected systems. A bank with 40 years of credit data across 15 core banking systems and three risk platforms cannot close the origination-to-collections feedback loop with a software update. A fintech NBFC with a modern stack, Account Aggregator integration, and a single data warehouse can close it in months.
The industry question follows from that: will incumbents rebuild, or will greenfield lenders take the market segment by segment? The lender that approves fewer loans that end in collection has a structural cost advantage. That advantage compounds over vintage cycles.
The AI-first lender is not better at collections. It is better at not needing collections.
