Thant Thu Aung

Master of IT in Business · SMU · Singapore

Data analytics and business intelligence

Thant Thu Aung, Teddy

I’m a Master of IT in Business student at SMU, specialising in data science and analytics. I clean and model data, check that the results hold up, and build the tools people use to act on them: a contract dashboard in Power BI and Microsoft Fabric in my internship, machine-learning and statistical analyses at JCU, and an Android app I built on my own.

Open to data & analytics internships

From my work: the internship dashboard (as a layout sketch, no company data), the gym booking app, an ocean-data model and LearnLog.

Highlights

  • 13 weeksas a remote data analytics intern, building a contract dashboard in Power BI and Microsoft Fabric with DAXInternship
  • 280,000+business registration records behind a hub-ranking model (sampled to 28,000 for modelling)Growth hubs
  • 0.959R² for the dissolved-oxygen model I chose from three candidates on AIC and residual checks, not R² aloneOcean statistics
  • 2 sprintsrun as Scrum coordinator for a two-person web app: 8 user stories, 53 person-days, burndown and velocity chartsGym project

Selected work

One internship, two data projects, two software builds.

Each case study covers the question, my part in it, the decisions that mattered, the evidence, and what I would change.

  1. 01Internship · BI & reporting

    Contract Intelligence Dashboard

    Worldtech Electronic Co., Ltd. · Remote · 2026

    A two-page Power BI report on Microsoft Fabric for monitoring a contract portfolio: what is active, what expires soon, and where the contracted value sits.

    My part
    Built both report pages, wrote the DAX measures and expiry logic the model was missing, and redesigned the layout after supervisor review.
    Result
    2-page report tracking active contracts, the expiry pipeline and contracted value, with the DAX expiry logic the data model was missing.
    • Power BI
    • Microsoft Fabric
    • DAX
    • Azure Storage
    Read the case study: Contract Intelligence Dashboard
  2. 02Machine learning

    Singapore Growth Hubs

    Team project · CP3403 Data Mining, JCU · 2025

    Which parts of Singapore look most likely to become business hubs? We used business registration records to classify, rank and cluster locations.

    My part
    Feature engineering and model development across the classification and clustering stages.
    Result
    Ranked shortlist scored by model confidence weighted by business count, so a handful of firms can’t outrank hundreds. Paya Lebar, Anson and Robinson Roads led, and both classifiers agreed on the top four.
    • Python
    • pandas
    • scikit-learn
    • kmodes
    • Matplotlib
    Read the case study: Singapore Growth Hubs
  3. 03Statistical analysis

    Ocean Water Statistics

    Individual coursework · CP2403, JCU · 2025

    Four questions about ocean water quality, each answered with the method that fits the data: ANOVA, chi-squared, simple and multiple regression.

    My part
    Everything: data preparation, method choice, modelling, diagnostics and the written reports.
    Result
    R² = 0.959 for dissolved oxygen, from the model I chose out of three candidates on AIC and residual checks, not on R² alone.
    • Python
    • pandas
    • statsmodels
    • SciPy
    • seaborn
    Read the case study: Ocean Water Statistics
  4. 04Agile delivery · web app

    JCU Gym Management System

    Team of 2 · Advanced Software Engineering, JCU · 2025

    A gym booking system built in two Scrum iterations: members register and book sessions; admins approve accounts, manage sessions and track usage and billing.

    My part
    I ran our Scrum process (planning, task allocation, the project board, burndown and velocity charts) and led the documentation. My teammate wrote most of the code.
    Result
    53 person-days logged and tracked over two sprints (51 were planned), delivering a working Next.js and PostgreSQL app with member and admin consoles.
    • Scrum
    • User stories
    • GitHub Projects
    • Burndown & velocity charts
    • Technical writing
    Read the case study: JCU Gym Management System
  5. 05Android app

    LearnLog

    Solo academic project · 2025

    A study app that keeps tasks, planned sessions and actual focus time in one place, so students can compare what they planned with what they did.

    My part
    Designed, built and tested the whole app on my own: 28 commits over seven weeks.
    Result
    Five modules (Tasks, Planner, Timer, Insights and Notes), with tasks, focus sessions and notes stored locally and a focus timer that keeps running in the background.
    • Kotlin
    • MVVM
    • Hilt
    • Room
    • WorkManager
    • Firebase Auth
    Read the case study: LearnLog

Inside a project · Singapore Growth Hubs

From 280,000 records to a ranked shortlist.

  1. Define what “hub” means

    There’s no column for “future business hub”, so the first decision is the target. We tried two proxies: how busy a location already is, and how quickly new businesses are arriving.

    SVM · density proxy

    hub = registrations in region > median

    Balanced classes; measures how busy a place already is.
    Random Forest · growth proxy

    hub = latest-year new registrations ≥ 90th percentile

    Targets places where new businesses are arriving fastest.
  2. Get the data into shape

    More than 280,000 registration records, sampled to 28,000 for fast iteration. Dates and numbers were coerced and checked, incomplete rows dropped, and each location keyed by postal code and street name.

    1. 280,000+Business registration recordsLocation, entity type, status and industry code
    2. 28,00010% random sampleFixed seed, for faster iteration
    3. CleanedIncomplete rows removedInvalid dates and numbers coerced to missing; the block column alone had 360 gaps
    4. RegionsPostal code + street nameThe unit that is labelled, scored and clustered
  3. Model, then check the model

    An SVM baseline reached about 79% accuracy, and a Random Forest that also used business type and industry reached about 91%. The two used different labels, and location features let the forest partly memorise each region, so I treat 91% as an upper bound.

    Cross-validated accuracy

    • SVM~79%
      Density label (above-median registrations). Folds: 80.2%, 79.9%, 77.6%.
    • Random Forest~91%
      Growth label (top 10% of latest-year registrations). Mean of 3 stratified folds.

    Different labels, so not a like-for-like comparison.

    Random Forest · holdout (6,192 records)

    Random Forest confusion matrix on the holdout set: rows are actual classes, columns are predicted classes.
    Predicted non-hubPredicted hub
    Actual non-hub3,476True non-hub277False hub
    Actual hub395Missed hub2,044True hub
    Non-hub
    precision 0.90 · recall 0.93 · F1 0.91
    Hub
    precision 0.88 · recall 0.84 · F1 0.86
  4. Turn predictions into a shortlist

    Probabilities become a ranking weighted by how many businesses back them up, so a location with three firms can’t outrank one with nine hundred. Paya Lebar Road, Anson Road and Robinson Road came out on top.

    score = p̄(hub) × log(1 + businesses)

    • A region with 3 businessesHypothetical, for comparison · p̄ = 0.990score 1.37
    • Paya Lebar Road995 businesses, ranked first · p̄ = 0.989score 6.83

    Near-identical confidence, five times the score: volume decides.

  5. Read the full case study

01 · Target definitions

SVM · density proxy

hub = registrations in region > median

Balanced classes; measures how busy a place already is.
Random Forest · growth proxy

hub = latest-year new registrations ≥ 90th percentile

Targets places where new businesses are arriving fastest.

Experience & education

Where I’ve done the work.

  1. Data Analytics Intern

    Worldtech Electronic Co., Ltd. · Thailand · Remote

    • Built a two-page Power BI report on Microsoft Fabric for a contract-intelligence project: a Contract Log (KPI cards, filters, full contract table) and a Contract Analytics page (top markets, expiry pipeline, contracted revenue).
    • Worked on the DAX measures and expiry logic the data model was missing, so expiring-contract figures and the analytics summaries came out accurate.
    • Mapped the data model (contract details to contract rates) and the Azure storage flow (raw, processed, archive) to see what the report could support.
    • Reworked the first draft after supervisor review so it followed the company design guide and was easier to scan: corrected the header, improved the KPI layout and filter section, and cleaned up the table’s field names.
    Read the case study: Data Analytics Intern
  2. Master of IT in Business · Data Science & Analytics track

    Singapore Management University · Singapore

    • Current modules: Statistical Thinking for Data Science, Query Processing and Optimisation, Data Analytics Lab, Computational Thinking with Python.
  3. Bachelor of Information Technology

    James Cook University · Singapore

    • Awarded a 25% merit scholarship.
    • Academic projects in data mining, statistical analysis, software engineering, Android development and design thinking. Four are written up on this site.
  4. Secretary & Head of Public Relations

    JCU Myanmar Student Community · Singapore

    • Organised club activities, managed communications and created promotional content.

Skills, and where I’ve used them

Every group links to the project or role that shows it.

Skills by area, with the projects that demonstrate them
AreaTools and methodsEvidence
Statistics & analysisPython, pandas, statsmodels, SciPy, seaborn
Machine learningscikit-learn, SVM, Random Forest, K-Prototypes, PCA
BI & reportingPower BI, Microsoft Fabric, DAX, Power Query, Azure Storage
Databases & data modelsSQL, Room (SQLite), Schema design, Data modelling
Mobile engineeringKotlin, Android, MVVM, Hilt, WorkManager
Delivery & communicationScrum, Sprint planning, User stories, Technical writing, Figma

Also on my résumé: PostgreSQL, React, Java, C++ (basic), Git, Jupyter. Languages: English and Burmese.

About

Hi, I’m Thant Thu Aung (Teddy).

My first degree was in IT, at James Cook University in Singapore, where I took data mining and statistics alongside software engineering. A data analytics internship, building a contract dashboard in Power BI, strengthened my interest in data. Now at Singapore Management University (SMU), I’m specialising in data science and analytics.

I like work where the numbers have to hold up. In my regression work I chose between models on AIC and residual behaviour rather than stopping at R², and my growth-hubs write-up is explicit about what our label did and didn’t measure.

On teams I tend to take on coordination and writing as well as code. I ran our Scrum board and led documentation on the gym system, and at JCU I was Secretary and Head of Public Relations for the Myanmar Student Community, organising events and creating promotional content.

Based in
Singapore
Studying
Master of IT in Business, SMU (to Dec 2027)
Looking for
Data analyst, BI and data science internships
Languages
English, Burmese

Contact

Let’s connect.

Whether it’s a role, a project or a question about my work, email is the fastest way to reach me. I’m happy to walk through anything on this site.

thantthua.2026@mitb.smu.edu.sg