Machine learning · Team project · CP3403 Data Mining, JCU · 2025
Finding Singapore’s next business hubs
A group data-mining project that turned more than 280,000 business registration records into a ranked shortlist of likely growth locations. It also taught me to be precise about what a label really measures.
Top-ranked locationsRandom Forest · weighted hub score
- Paya Lebar Road postcode 409051995 businesses · p̄ = 0.989score 6.83
- Anson Road postcode 079903670 businesses · p̄ = 0.983score 6.40
- Robinson Road postcode 068914595 businesses · p̄ = 0.990score 6.33
- Circular Road postcode 049422515 businesses · p̄ = 0.986score 6.16
- Venture Drive postcode 608526321 businesses · p̄ = 0.962score 5.56
- Sin Ming Lane postcode 573969235 businesses · p̄ = 0.958score 5.23
- Temasek Boulevard postcode 038987198 businesses · p̄ = 0.966score 5.11
- North Bridge Road postcode 179098210 businesses · p̄ = 0.952score 5.10
- North Bridge Road postcode 179094190 businesses · p̄ = 0.957score 5.03
- Raffles Quay postcode 048581127 businesses · p̄ = 0.960score 4.66
01 Overview
The brief
For CP3403 Data Mining at JCU Singapore, our group set out to identify which areas of Singapore are most likely to become economic and innovation hubs. The only raw material was business registration data: postal code, street name and block, plus each entity’s type, status and industry code (SSIC).
The analysis combined supervised learning, to classify and rank locations, with unsupervised clustering, to describe what different kinds of areas look like.
02 Problem & users
A useful answer is a ranked, explainable shortlist
We wrote for urban planners, investors and policymakers: people deciding where infrastructure, capital or support programmes should go. They need more than a yes/no prediction per street. They need a ranking they can question.
03 My role
My part of the work
This was a group project (Group 34). I worked on feature engineering and model development across the classification and clustering stages: shaping the region-level features, setting up the models and interpreting their outputs.
04 Approach
Decisions that shaped the analysis
Sample before modelling
The raw file had more than 280,000 rows. We drew a 10% random sample (28,000 rows, fixed seed) so models could be iterated on quickly. The trade-off is fewer records for sparsely populated areas.
Define a “region” as postal code + street name
Street names alone merge very different places; the combination gave a stable unit to label, score and cluster.
Try two proxies for “hub”
The SVM used a density label: regions with more registrations than the median. The Random Forest used a growth-oriented label: regions in the top 10% of new registrations in their most recent year.
Rank, don’t just classify
A location with three businesses and a 99% hub probability shouldn’t outrank one with 900. Each region’s score is its average predicted probability × log(1 + number of businesses), which rewards confidence backed by volume without letting size dominate.
Use K-Prototypes for clustering
The data mixes numeric location fields with categorical business attributes. K-Means only handles numbers; K-Prototypes handles both directly.
05 Analysis
From raw records to models
- 280,000+Business registration recordsLocation, entity type, status and industry code
- 28,00010% random sampleFixed seed, for faster iteration
- CleanedIncomplete rows removedInvalid dates and numbers coerced to missing; the block column alone had 360 gaps
- RegionsPostal code + street nameThe unit that is labelled, scored and clustered
- Cleaning: registration dates converted to datetimes (invalid values coerced to missing), block and postal code converted to numbers, text fields trimmed. Rows with missing values were dropped; the block column alone had 360 gaps.
- SVM baseline: block, postal code and registrations-per-region, standardised, with a sigmoid kernel and 3-fold stratified cross-validation.
- Random Forest: 100 trees in a scikit-learn pipeline. Block and postal code were scaled; entity type, company type, SSIC code and status were one-hot encoded (ignoring unseen categories). Evaluated with 3-fold stratified cross-validation, then on a holdout set.
- Clustering: regions aggregated with their business count and most common entity type, company type, SSIC code and status; K-Prototypes with k = 3; PCA to plot the result in two dimensions.
06 Results
What we found
- Random Forest mean accuracy, 3-fold stratified CV
- ~91%
- Random Forest accuracy on the 6,192-record holdout set
- 89%
- SVM mean accuracy (folds: 80.2%, 79.9%, 77.6%)
- ~79%
- SVM ROC AUC
- 0.866
Cross-validated accuracy
- SVM~79%Density label (above-median registrations). Folds: 80.2%, 79.9%, 77.6%.
- Random Forest~91%Growth label (top 10% of latest-year registrations). Mean of 3 stratified folds.
Different labels, so not a like-for-like comparison.
Random Forest · holdout (6,192 records)
| Predicted non-hub | Predicted hub | |
|---|---|---|
| Actual non-hub | 3,476True non-hub | 277False hub |
| Actual hub | 395Missed hub | 2,044True hub |
- Non-hub
- precision 0.90 · recall 0.93 · F1 0.91
- Hub
- precision 0.88 · recall 0.84 · F1 0.86
The ranking at the top of this page is the main output. Its top four, Paya Lebar Road, Anson Road, Robinson Road and Circular Road, were the same under both classifiers, and they are recognisable commercial areas. That agreement is a useful sanity check, not proof.
- Cluster 0Mature commercial hubs515–995registrations in its largest regionsPaya Lebar Rd · Anson Rd · Robinson Rd · Circular RdDense, mostly live companies in high-value sectors.
- Cluster 1Transitional areas81–89registrations in its largest regionsCecil St · Woodlands Sq · Phillip St · Jalan BesarMixed entity types and statuses; moderate activity.
- Cluster 2Low-activity areas21registrations in its largest regionsSmith St · Raffles Blvd · Kim Seng PromenadeLower registration counts; candidates for renewal.
07 Limitations & learning
What I would do differently
- The labels come from the data the models learn from. The SVM’s label is defined by registrations per region, and registrations per region is also one of its features. The accuracy figures therefore show how well each model reproduces our proxy, not how well it forecasts growth.
- Location features can leak the answer. A region is defined by postal code + street name, and the Random Forest also uses postal code and block as features. Because the split was by record rather than by region, businesses from the same region sit in both training and test data, so the model can partly memorise each region’s label. The ~91% is therefore an upper bound, not an estimate of performance on unseen areas.
- 79% vs 91% is not a like-for-like comparison. The two classifiers were trained on different label definitions, so their accuracies answer different questions.
- k = 3 was chosen from business logic. An elbow or silhouette check would make the cluster count easier to defend, and the clusters largely follow registration volume.
08 Links