Thant Thu Aung
All work

Machine learning · Team project · CP3403 Data Mining, JCU · 2025

Finding Singapore’s next business hubs

A group data-mining project that turned more than 280,000 business registration records into a ranked shortlist of likely growth locations. It also taught me to be precise about what a label really measures.

My role
Feature engineering and modelling
Team
Group 34 (group project)
When
Trimester 1, 2025
Stack
Python · pandas · scikit-learn · kmodes

Top-ranked locationsRandom Forest · weighted hub score

  1. Paya Lebar Road postcode 409051995 businesses · p̄ = 0.989score 6.83
  2. Anson Road postcode 079903670 businesses · p̄ = 0.983score 6.40
  3. Robinson Road postcode 068914595 businesses · p̄ = 0.990score 6.33
  4. Circular Road postcode 049422515 businesses · p̄ = 0.986score 6.16
  5. Venture Drive postcode 608526321 businesses · p̄ = 0.962score 5.56
  6. Sin Ming Lane postcode 573969235 businesses · p̄ = 0.958score 5.23
  7. Temasek Boulevard postcode 038987198 businesses · p̄ = 0.966score 5.11
  8. North Bridge Road postcode 179098210 businesses · p̄ = 0.952score 5.10
  9. North Bridge Road postcode 179094190 businesses · p̄ = 0.957score 5.03
  10. Raffles Quay postcode 048581127 businesses · p̄ = 0.960score 4.66
Top ten locations by weighted hub score from the Random Forest model. Score = average predicted hub probability × log(1 + number of businesses). Each region is a postcode on a street, so North Bridge Road appears twice: 179098 and 179094 are separate buildings.

The brief

For CP3403 Data Mining at JCU Singapore, our group set out to identify which areas of Singapore are most likely to become economic and innovation hubs. The only raw material was business registration data: postal code, street name and block, plus each entity’s type, status and industry code (SSIC).

The analysis combined supervised learning, to classify and rank locations, with unsupervised clustering, to describe what different kinds of areas look like.

A useful answer is a ranked, explainable shortlist

We wrote for urban planners, investors and policymakers: people deciding where infrastructure, capital or support programmes should go. They need more than a yes/no prediction per street. They need a ranking they can question.

My part of the work

This was a group project (Group 34). I worked on feature engineering and model development across the classification and clustering stages: shaping the region-level features, setting up the models and interpreting their outputs.

Decisions that shaped the analysis

  1. Sample before modelling

    The raw file had more than 280,000 rows. We drew a 10% random sample (28,000 rows, fixed seed) so models could be iterated on quickly. The trade-off is fewer records for sparsely populated areas.

  2. Define a “region” as postal code + street name

    Street names alone merge very different places; the combination gave a stable unit to label, score and cluster.

  3. Try two proxies for “hub”

    The SVM used a density label: regions with more registrations than the median. The Random Forest used a growth-oriented label: regions in the top 10% of new registrations in their most recent year.

  4. Rank, don’t just classify

    A location with three businesses and a 99% hub probability shouldn’t outrank one with 900. Each region’s score is its average predicted probability × log(1 + number of businesses), which rewards confidence backed by volume without letting size dominate.

  5. Use K-Prototypes for clustering

    The data mixes numeric location fields with categorical business attributes. K-Means only handles numbers; K-Prototypes handles both directly.

From raw records to models

  1. 280,000+Business registration recordsLocation, entity type, status and industry code
  2. 28,00010% random sampleFixed seed, for faster iteration
  3. CleanedIncomplete rows removedInvalid dates and numbers coerced to missing; the block column alone had 360 gaps
  4. RegionsPostal code + street nameThe unit that is labelled, scored and clustered
Preparation steps, with row counts from the report.
  • Cleaning: registration dates converted to datetimes (invalid values coerced to missing), block and postal code converted to numbers, text fields trimmed. Rows with missing values were dropped; the block column alone had 360 gaps.
  • SVM baseline: block, postal code and registrations-per-region, standardised, with a sigmoid kernel and 3-fold stratified cross-validation.
  • Random Forest: 100 trees in a scikit-learn pipeline. Block and postal code were scaled; entity type, company type, SSIC code and status were one-hot encoded (ignoring unseen categories). Evaluated with 3-fold stratified cross-validation, then on a holdout set.
  • Clustering: regions aggregated with their business count and most common entity type, company type, SSIC code and status; K-Prototypes with k = 3; PCA to plot the result in two dimensions.

What we found

Random Forest mean accuracy, 3-fold stratified CV
~91%
Random Forest accuracy on the 6,192-record holdout set
89%
SVM mean accuracy (folds: 80.2%, 79.9%, 77.6%)
~79%
SVM ROC AUC
0.866

Cross-validated accuracy

  • SVM~79%
    Density label (above-median registrations). Folds: 80.2%, 79.9%, 77.6%.
  • Random Forest~91%
    Growth label (top 10% of latest-year registrations). Mean of 3 stratified folds.

Different labels, so not a like-for-like comparison.

Random Forest · holdout (6,192 records)

Random Forest confusion matrix on the holdout set: rows are actual classes, columns are predicted classes.
Predicted non-hubPredicted hub
Actual non-hub3,476True non-hub277False hub
Actual hub395Missed hub2,044True hub
Non-hub
precision 0.90 · recall 0.93 · F1 0.91
Hub
precision 0.88 · recall 0.84 · F1 0.86
Random Forest holdout confusion matrix and per-class precision and recall, from the classification report.

The ranking at the top of this page is the main output. Its top four, Paya Lebar Road, Anson Road, Robinson Road and Circular Road, were the same under both classifiers, and they are recognisable commercial areas. That agreement is a useful sanity check, not proof.

  • Cluster 0Mature commercial hubs515–995registrations in its largest regionsPaya Lebar Rd · Anson Rd · Robinson Rd · Circular RdDense, mostly live companies in high-value sectors.
  • Cluster 1Transitional areas81–89registrations in its largest regionsCecil St · Woodlands Sq · Phillip St · Jalan BesarMixed entity types and statuses; moderate activity.
  • Cluster 2Low-activity areas21registrations in its largest regionsSmith St · Raffles Blvd · Kim Seng PromenadeLower registration counts; candidates for renewal.
Cluster profiles as interpreted in the report, with registration counts for the largest regions in each cluster.

What I would do differently

  • The labels come from the data the models learn from. The SVM’s label is defined by registrations per region, and registrations per region is also one of its features. The accuracy figures therefore show how well each model reproduces our proxy, not how well it forecasts growth.
  • Location features can leak the answer. A region is defined by postal code + street name, and the Random Forest also uses postal code and block as features. Because the split was by record rather than by region, businesses from the same region sit in both training and test data, so the model can partly memorise each region’s label. The ~91% is therefore an upper bound, not an estimate of performance on unseen areas.
  • 79% vs 91% is not a like-for-like comparison. The two classifiers were trained on different label definitions, so their accuracies answer different questions.
  • k = 3 was chosen from business logic. An elbow or silhouette check would make the cluster count easier to defend, and the clusters largely follow registration volume.