Introduction
How do we define what makes two cyclists similar? In this project, I analyzed professional peloton data to build a cyclist recommendation system. Using data cleaning techniques and vector distance metrics, I managed to identify which riders share similar performance profiles, regardless of their team or nationality.
Development Process
The project focused on transforming raw tabular data into comparable performance profiles.
1. Data Collection
The data was obtained from public sources, specifically after an exhaustive search, I decided to use those provided at Cycling Oracle

Example data from Cycling Oracle
This data included a score out of 100 that reflects the performance of each cyclist in different types of races (classics, mountain stages, time trials, etc.), as well as their age and average score.
2. Cleaning and Normalization
The key was to normalize numerical variables (such as victories, UCI points, age, or specialty) so that no characteristic dominated the others.
- Handling null values: Techniques were applied so that missing data would not cause errors in the similarity calculation.
- Feature Engineering: Creation of “climber profile” versus “rouleur” indicators through the analysis of their points.
3. Similarity Measurement: Multimodal Recommendation Engine
To determine the affinity between riders, I implemented a hybrid system that combines three distinct metrics, balanced using specific weights to maximize sporting relevance:
Weighted Cosine Similarity (60%): Instead of treating all skills equally, I applied dynamic weights based on the cyclist’s strengths. If a rider stands out in a discipline (score > 70), the model doubles the weight of that characteristic, avoiding that “weaknesses” (disciplines in which they do not compete) distort the recommendation.
Inverse Euclidean Distance (25%): Used to measure absolute proximity in the vector space, ensuring that riders with similar performance levels numerically get a bonus.
Physical Similarity (15%): I processed height, weight, and age data through robust null imputation, allowing the system to compare morphologies even with incomplete data.
Logical Refinement
In addition to mathematical metrics, I added a filtering logic layer to refine the results:
Age Penalty: I applied a penalty factor for cyclists with a difference greater than 7 years, prioritizing comparable generational cohorts.
Profile Bonus: An additional bonus is granted if both riders share the same dominant profile (e.g., Climber vs. Climber), ensuring that the model prefers specialists in the same discipline.
# Weighted combination of metrics
combined_scores = (0.60 * cosine_scores + 0.25 * euclidean_scores + 0.15 * physical_scores)
# Application of intelligent filters (Age and Profile)
results['SimilarityScore'] = results['SimilarityScore'] - results['AgePenalty']
results['SimilarityScore'] = results['SimilarityScore'] + results['ProfileBonus']
4. Web App Construction
To make the recommendation system accessible, I developed a web application using Flask. The interface allows users to select a cyclist and view their similarity recommendations, along with charts showing profile comparisons.
Result
Beyond the empirical validation of obvious relationships, the engine acts as an analytical scouting tool. By identifying the underlying performance signature in a multidimensional vector space, the model allows for the detection of emerging talents whose statistical profile is analogous to that of WorldTour leaders, facilitating the prospecting of prospects before their consolidation in the elite.
The web application is available for use and can be accessed through the following link: Similar Cyclists

Search engine

Comparative chart

Comparative bar charts

Table of similar cyclists