CASE STUDY · PYTHON & FEATURE ENGINEERING

Delhivery Logistics
Data Analysis.

An end-to-end logistics analytics case study focused on data cleaning, analytical grain control, trip-level aggregation, feature engineering and comparison of observed operational performance with routing-system estimates.

PythonPandasFeature EngineeringLogistics AnalyticsEDA
Delhivery logistics analysis visual
BUSINESS PROBLEM

Turn raw logistics records into operationally useful data.

The objective was to understand and process logistics operations data while maintaining the correct analytical grain, consolidating trip records and creating structured features for delivery-performance analysis.

What is the correct analytical grain for the operational data?
How should multiple records belonging to one trip be consolidated?
How can geographic fields be extracted from source and destination data?
Which time-based features are useful for trip analysis?
How does actual performance compare with OSRM estimates?
How can the cleaned dataset support downstream logistics analytics?
DATASET

144,867 operational records across 14,817 trips.

The dataset covers approximately 27 days of operations, from 12 September 2018 to 8 October 2018, and includes trip, route, centre, timestamp, distance, duration and routing-estimate fields.

144,867Operational records
14,817Unique trips
~27 daysObservation period
19Final analytical features
Core data dimensions
Trip identifiers · route schedules · route type · source/destination centres · operational timestamps · actual distance/time · OSRM distance/time · segment-level metrics
ANALYTICAL GRAIN

Get the unit of analysis right before calculating KPIs.

Multiple operational records can belong to the same trip. The project therefore treats trip-level aggregation as a key analytical step before deriving final trip-level metrics.

Raw Records→Trip Identifier→Trip-Level Aggregation→Feature Engineering→Operational Analysis

Avoid duplicate trip counts

Raw operational records should not automatically be treated as independent trips.

Protect distance metrics

Aggregating at the trip level helps avoid inflated or duplicated distance calculations.

Protect time metrics

Trip-level aggregation provides a more appropriate basis for duration comparisons.

Enable route comparison

A controlled grain makes route, source and destination comparisons more meaningful.

DATA PREPARATION

Clean first. Engineer second.

The notebook covers dataset profiling, timestamp conversion, missing-value assessment, duplicate checks, standardisation, trip aggregation and final feature preparation.

293Missing source-name valuesLess than 0.3% of the 144,867 operational records.
261Missing destination-name valuesAlso less than 0.3% of the operational records.
0Duplicate records identifiedNo duplicate records were identified in the analysis.
19Final analytical featuresStructured feature set after preprocessing.
FEATURE ENGINEERING

Convert operational fields into analytical dimensions.

Geographic features

Source and destination strings are parsed into dimensions such as city, place code and state/region.

Time features

Timestamp fields are transformed into year, month, day, trip duration and operational time measures.

Duration features

Duration measures are derived from operational timestamps and compared with scan-to-scan measures for validation.

Routing comparison

Actual distance/time are compared with OSRM distance/time to examine the gap between observed operations and routing estimates.

Segment-level metrics

Segment actual time, OSRM time and OSRM distance provide additional operational context.

Geographic footprint

The dataset contains more than 1,500 source/destination logistics-centre locations for geographic analysis.

KEY FINDINGS

Observed operations versus routing estimates.

The central analytical comparison examines actual operational performance against OSRM routing-system estimates.

417 minAverage actual trip time
214 minAverage OSRM time
234 kmAverage actual distance
285 kmAverage OSRM distance
ACTUAL VS OSRM TIMEAverage actual trip time was approximately 1.95× the OSRM estimate.

The observed gap is a descriptive finding from this dataset. It indicates that routing estimates and actual operational duration should be evaluated as separate measures when analysing delivery performance and capacity.

DISTANCE & OPERATIONAL CONTEXT

Time and distance tell different stories.

The analysis also shows that observed actual distance was lower than the corresponding OSRM estimate.

234 kmAverage actual distanceObserved operational distance in the analytical dataset.
285 kmAverage OSRM distanceRouting-system estimate in the dataset.
~18% lowerActual vs OSRM distanceAverage observed distance was approximately 18% below the OSRM estimate.
1,500+Source/destination locationsBroad geographic footprint for downstream route analysis.
Analytical consideration
Actual and OSRM metrics represent different concepts and should not automatically be treated as interchangeable measures or operational targets.
BUSINESS APPLICATIONS

What the structured dataset enables.

Logistics performance

Analyse delivery duration, route performance, actual distance and actual-vs-estimated time.

Geographic analysis

Compare source regions, destination regions, city/state patterns and route-level differences.

Route analysis

Examine route type, route schedules and trip-level operational performance.

Operational KPIs

Track actual time, estimated time, actual distance, estimated distance and the gaps between them.

ANALYTICAL FRAMEWORK

Understand → Clean → Aggregate → Engineer → Validate → Analyse.

The project is structured as a repeatable data-preparation and operational-analysis workflow.

Raw Logistics Data→Data Profiling→Data Cleaning→Grain Control→Trip Aggregation→Feature Engineering→Operational Insights
TECH STACK

Python for logistics data transformation.

The project uses Python-based data preparation and feature engineering to convert raw operational records into a structured analytical dataset.

PythonPandasNumPyDatetime ManipulationData CleaningFeature EngineeringJupyter / Google ColabGitHub
ANALYTICAL CONSIDERATIONS

Keep operational metrics in context.

Correct grain

Raw operational records should not automatically be treated as independent trips.

Trip identifiers

Trip-level identifiers are important for maintaining the correct analytical grain.

Actual vs OSRM

Actual and routing-system metrics represent different concepts and should be interpreted separately.

Routing estimates

OSRM estimates should not automatically be treated as operational targets.

Geographic fields

Location fields derived from semi-structured strings should be validated before detailed geographic reporting.

Descriptive evidence

The observed time and distance gaps are descriptive findings from this dataset and should not be interpreted as causal evidence.

PROJECT RESOURCES

Explore the logistics
analysis in detail.