Anomaly detection · Big data · Sprint (internship)
As a data-science intern at Sprint, I led research on a model-based anomaly-detection system for billing failures — and earned the only extension across an 80+ intern class.
What I did
I designed a novel model-training framework for anomaly detection on customer billing data, clustering common billing issues from account data. To feed it, I built Hive and PySpark pipelines aggregating millions of customer billing records across remote servers — unlocking training variables the data-science team hadn't been able to reach before.
The result
The framework drove a 5x increase in the types of billing failure the system could detect. The internship turned into a three-month extension (remote from Dallas) — the only one offered across a class of 80+ interns.
What it demonstrates
Early proof of the pattern that runs through my career: build the data foundation yourself, then model on top of it — and turn a messy, large-scale data problem into something a team can actually use.