CASE STUDY · API DATA COLLECTION & PROCESSING
E-Commerce Product Data Collector
A Python data-collection workflow that retrieves product records from an API, validates and cleans the response, applies business-defined filters, and produces structured CSV and JSON outputs.
Turning API responses into controlled, analysis-ready product data.
Product APIs can return records that are incomplete, duplicated, or inconsistent with the rules of a specific analysis. Saving the raw response directly pushes those problems into every downstream report.
This collector creates a repeatable processing layer between the data source and the final dataset. Users define a product query, price range, minimum rating, and sorting rule; the application then retrieves all matching records and applies the same validation and transformation workflow to every run.
DummyJSON is used as a demonstration source, while the architecture can be adapted to permitted public APIs or other structured product feeds.
Validation at both the user-input and record level.
Numeric inputs are checked for valid ranges and rejected when they contain invalid values such as NaN or infinity. Returned product fields are validated independently so malformed API data cannot silently enter the exported dataset.
The workflow removes invalid and duplicate records, handles missing values, applies price and rating filters, and sorts the final collection by price, rating, or stock.
- Keyword-based API search
- Input and response-field validation
- Invalid-record and duplicate removal
- Price and rating filtering
- Price, rating, and stock sorting
- Network, timeout, data, and file-output error handling
Structured outputs with a transparent process summary.
Each run reports whether the API request succeeded and shows how many products were fetched, cleaned, rejected, deduplicated, removed by filters, and retained. This makes transformation results visible instead of treating data loss as a hidden step.
The final records are written to timestamped JSON and CSV files using safe filenames. Outputs preserve useful product attributes including identifiers, brand, category, price, discount, rating, stock, availability, shipping information, and image references.
Producing both machine-readable JSON and analysis-friendly CSV lets the same collection support application workflows, spreadsheet review, and downstream reporting.
One collection run, from validated request to reusable files.
The demonstration shows the processing summary alongside the two structured formats generated from the cleaned and filtered product collection.
Technology stack: Python · Requests · REST API · CSV · JSON · Regular Expressions


