Billing-Failure Anomaly Detection

Anomaly detection · Big data · Sprint (internship)

As a data-science intern at Sprint, I led research on a model-based anomaly-detection system for billing failures — and earned the only extension across an 80+ intern class.

5xDetectable failure types
MillionsRecords aggregated
Only 1Extension / 80+ interns

What I did

I designed a novel model-training framework for anomaly detection on customer billing data, clustering common billing issues from account data. To feed it, I built Hive and PySpark pipelines aggregating millions of customer billing records across remote servers — unlocking training variables the data-science team hadn't been able to reach before.

The result

The framework drove a 5x increase in the types of billing failure the system could detect. The internship turned into a three-month extension (remote from Dallas) — the only one offered across a class of 80+ interns.

What it demonstrates

Early proof of the pattern that runs through my career: build the data foundation yourself, then model on top of it — and turn a messy, large-scale data problem into something a team can actually use.