Skip to main content

Evaluating Machine Learning–Driven Fraud Detection: Evidence from AquaData in EPMAPS

The objective is to determine the causal impact of machine learning-driven inspection targeting, through the AquaData algorithm, on fraud detection effectiveness, inspector productivity, revenue recovery, and apparent losses at EPMAPS. Because the 2023 pilot did not generate an identifiable counterfactual, the design is prospective and experimental.

The main question is whether using the algorithm to prioritize inspections increases inspector productivity and fraud detection effectiveness relative to the current expert-based assignment process. Two complementary questions follow: the impact on inspector productivity, measured through inspections completed, hit rates, and recovered consumption per inspector; and the cost-effectiveness of the algorithm relative to the current process, accounting for implementation, operational, and inspection costs.

The knowledge gap is well defined. Utilities across the region increasingly use machine learning to detect anomalies in consumption, flow, and pressure data that may signal fraud (Lee and Kim, 2023; Romero-Ben et al., 2023), and the problem is not unique to water: electricity losses average roughly 17% of grid input in the region (Jimenez Mori et al., 2014). However, this literature remains largely descriptive and relies on retrospective comparisons that cannot separate the effect of the algorithm from selection bias, leaving utilities without credible evidence on the gains they can expect.

The theory of change is direct: using existing inspection staff, the resources to operate the algorithm, and historical billing data, the project designs a data collection protocol, institutionalizes the algorithm within the commercial losses unit, and trains inspectors. The output is an operational algorithm that prioritizes connections with the highest predicted probability of fraud, which should raise the share of inspections confirming fraud and the number of inspections per day and, over time, reduce apparent losses and improve revenue recovery, fine collection, and financial sustainability.

The methodology is a prospective RCT randomized at the operational sector level. Of the 110 sectors (some 720,000 active connections), eligible sectors will be randomly assigned to treatment, where inspection priority is set by AquaData, or control, where the existing process continues. The sample is restricted to active residential connections (89% of contracts and 80% of consumption) and excludes remote or unsafe sectors under an ex-ante geographic criterion. Sector-level randomization mitigates spillovers between neighboring connections. The intervention is the use of the algorithm to prioritize cases, not the inspection itself, which is the field validation mechanism. Impact will be estimated by intention to treat using OLS, with standard errors clustered at the sector level and baseline controls, verifying balance between groups.

Outcomes are measured with EPMAPS administrative data: (a) the share of inspections confirming an infraction; (b) inspections per inspector per working day; (c) revenue recovered through billing adjustments and fines; (d) effective collection of issued fraud notifications, against a baseline of 30-40%; and (e) the apparent loss index at the operational unit level.

Fund resources will finance the design of the data collection protocol, the digitization of inspection records, the implementation and monitoring of the randomized rollout, the analysis, and dissemination.

Project Detail

Country

Ecuador

Project Number

EC-T1647

Approval Date

-

Project Status

Preparation

Project Type

Technical Cooperation

Sector

WATER AND SANITATION

Subsector

WATER SUPPLY URBAN

Lending Instrument

-

Lending Instrument Code

-

Modality

-

Facility Type

-

Environmental and Social Impact Category (ESIC)

-

Total Cost

USD 150,000.00

Country Counterpart Financing

-

Original Amount Approved

USD 150,000.00

Jump back to top