Skip to content
Projects

◆ — Coursework

Coursework, revived.

18 University of Melbourne coursework projects from 2019 to 2024, rebuilt in 2026 as live demos you can try in the browser.

01What revived means

Each project began as an assignment I submitted at the University of Melbourne. In 2026 I brought them back in three steps, kept separate so it is always clear what is original and what is new.

  1. Faithful port

    The original code is ported to TypeScript and checked against the original Python, C, R or JavaScript with parity tests. Known bugs are kept as submitted and documented, with a corrected mode where it helps.

  2. Statistical re-checks

    Results are re-checked with confidence intervals, paired tests and simulation (still in progress for the two newest labs), and the notes record where the original conclusions do not hold.

  3. Optional BYOK AI

    Most labs add an optional AI feature that runs on the visitor's own API key. Its output is checked against the data or scored against simple baselines, and many keep a local log of every call.

Every revival is built withNext.js 16 · React 19 · TypeScript · Tailwind CSS 4 · shadcn/ui · Vitest

  • Dates are the final deadline or submission month, shown only as precisely as the records allow.
  • Team projects name every teammate. Every project links its live demo, plus its guided tour of captioned walkthrough videos where there is one. Source code is linked once its public repository holds the 2026 revival.

02Timeline

Filter projects

Area
Level

Showing 18 of 18 projects

Undergraduate: Bachelor of Science, Data Science, Jul 2019 – Jul 2022

2019, Semester 2

COMP10001 Foundations of Computing

Date:

Live

COMP10001 Playground

My first three programming projects in one interactive site: count an election three ways, guide Falca past a dragon to the treasure with BFS and uniform-cost search, and find the best way to group a hand of cards.

My role
Individual projects. I wrote the originals on Grok Learning and built the 2026 revival.

Highlights

  • The cave project is a line-by-line port of my surviving 2019 file, bugs included, beside a spec-correct mode. The two lost projects were rebuilt from the task and are labelled as rebuilt.
  • Parity tests check the TypeScript against the Python on 230 seeded caves, and property-based tests check invariants against independent oracles.

2020, Semester 1

COMP10002 Foundations of Algorithms

Date:

Live

Message Cleanser

My first C assignment, a five-stage cleanser that strips noise from comma-tokenised messages and keeps only known emoticons, revived as a stage-by-stage visualiser running a faithful TypeScript port.

My role
Individual assignment, written on a skeleton from the teaching staff. I wrote the solution and built the 2026 revival.

Highlights

  • Recovered the complete 2020 solution from a Notability note and confirmed it reproduces both official sample outputs byte for byte.
  • Differential testing on 500 seeded inputs against the compiled C program: all 481 completed runs match byte for byte, and the 19 runs where the C binary crashes are flagged by the port instead.

2020, Semester 2

COMP20008 Elements of Data Processing

Date:

Live

Data Processing Lab

My two Elements of Data Processing projects as one interactive lab: replay a breadth-first crawl of rugby match reports, tune an Abt-Buy product matcher and its blocking keys, and train k-NN and decision tree classifiers on World Bank indicators.

My role
Individual projects. I wrote the original code and reports and built the 2026 revival.

Highlights

  • Every 2020 figure is reproduced exactly, down to each decision tree node, through ports of NumPy's random generator and scikit-learn's tree builder.
  • Repeated 10 × 5-fold cross-validation showed the report's preferred decision tree is not reliably better than 3-NN (95% interval for the gap: -4.4 to +10.2 points).

2021, Semester 1

INFO30005 Web Information Technologies

Date:

Live

Snacks in a Van

A two-portal ordering app for roving snack vans from our INFO30005 group project. Customers find the nearest open vans on a map and order ahead, and vendors run a live order board. Rebuilt as one Next.js app with demo accounts, operations analytics and an A/B test designer.

My role
In Group 4399 I designed the vendor app and built customer ordering, the outstanding-orders list, customer and vendor login, and the optional map and blog features. I rebuilt it in 2026.
Team
Bin Liang, Declan Gannon, Khin Liew and Wei Zhao

Highlights

  • Ported the 2021 business rules (nearest vans, order states, change window, pricing) to TypeScript, with parity tests that run the original JavaScript side by side.
  • Operations analytics include a Kaplan-Meier time-to-fulfil curve and late-discount rates by van, each shown with its uncertainty and sample size.

2021, Semester 2

MAST30034 Applied Data Science

Date:

Live

NYC Taxi 2019

My analysis of every 2019 New York yellow-cab trip, revived as an interactive site: the same cleaning rules re-run on more than 84 million records, joined to weather, permitted events and collisions, with zone maps, route lines and the 2021 regression estimating trip times in the browser.

My role
Individual project. I wrote the original PySpark notebook and built the 2026 revival.

Highlights

  • Re-ran the 2021 cleaning rules as DuckDB SQL on TLC's current 2019 files. Apart from the raw files, which TLC has since re-issued, every row count lands within 0.003% of the notebook's and is logged beside it.
  • The 2021 regression runs in the browser with every term behind a prediction shown, and the 2026 evaluation adds a temporal hold-out, robust standard errors, residual diagnostics and conformal intervals, where a simple lookup table turns out to beat it.

MAST30034 Applied Data Science

Date:

Live

Source Separation Lab

My Applied Data Science assignment on recovering hidden sources from a synthetic fMRI-style dataset, rebuilt as an interactive lab. Generate the data, then recover the time courses and spatial maps with least squares, ridge, lasso and principal component regression, all recomputed in the browser.

My role
Individual assignment. I wrote the original analysis, notebook and report, and built the 2026 revival.

Highlights

  • A bit-for-bit port of NumPy's legacy random generator regenerates the 2021 noise, so every reported number is recomputed exactly in the browser.
  • Re-examined the lasso penalty in a paired design on 50 seeded datasets against a comparator fixed in advance: the submitted ρ = 0.625 gives about 10% higher MSE than ρ = 0.60 and is worse on all 50.

COMP30022 IT Project

Date:

Live

4399 CRM

A mobile-first personal CRM from our COMP30022 IT Project: keep track of contacts, log geo-tagged meetings, view them on a map and calendar, and swap contacts by QR code. Rebuilt as a single Next.js app with a 'Try as guest' sandbox, an Insights page and an optional AI assistant.

My role
Scrum Master for Team 4399 (Group 49) and a front-end developer on contacts, meeting records and the map. I rebuilt it in 2026.
Team
Bin Liang, Hongji (Harrison) Huang, Wei Zhao and Yixiao Tian

Highlights

  • Ported the original search, sort, validation and contact-sync logic, with parity tests that load the 2021 functions straight from the coursework folder.
  • Closed security gaps found during the port: queries are scoped to the signed-in owner, a constant reset bypass is gone and sessions use httpOnly cookies.

2022, Semester 1

COMP30024 Artificial Intelligence

Date:

Live

Cachex Arena

Our COMP30024 agent for Cachex, a Hex-like connection game with captures and a steal rule, revived as a browser arena. Play the minimax agent, watch AI vs AI, step through the A* solver, and run a seeded tournament that reports win rates with confidence intervals.

My role
Pair project with Wei Zhao as team _4399: we wrote the A* solver, the minimax agent and the Part A report together. I rebuilt it as Cachex Arena in 2026.
Team
Wei Zhao

Highlights

  • Ported the Python agent and A* solver line by line to TypeScript running in Web Workers. Parity tests match the original on 184 A* runs, 40 refereed games and 179 minimax searches.
  • A seeded 1,200-game round robin with Wilson and bootstrap intervals showed the dynamic-depth search rarely went past one ply, and a one-line bug stopped alpha-beta from narrowing its window. Both are documented and the agent is kept as submitted.

Master's: Master of Data Science, Feb 2023 – Jul 2024

2023, Semester 1

COMP90051 Statistical Machine Learning

Date:

Live

Human or Machine?

Team 27's detectors for a class Kaggle competition: decide whether a text was written by a person or a model when every word has been replaced by a number. Revived as a lab where the original weights score texts live in the browser, with a class-imbalance sandbox and a leak-free re-evaluation.

My role
Team 27 with Yifei Du and Huihui He. I worked on the MPI preprocessing and the CNN experiments, and shared the bag-of-words features with Yifei and the SMOTE and under-sampling comparison with both; Yifei built the SVM and SGD models and Huihui the logistic regression and word embeddings. I built the 2026 revival.
Team
Yifei Du and Huihui He

Highlights

  • Exported the six original detectors' weights once and score them in the browser; all 6,000 test predictions (six models by 1,000 texts) match the 2023 files.
  • An interactive detector blends human and machine token profiles into a specimen and shows which token ids push the score each way.

COMP90024 Cluster and Cloud Computing

Date:

Live

Spartan Tweet Cruncher

Our pair assignment that processed 9.09 million geotagged tweets with MPI on the University's Spartan HPC, revived with the original results, the scaling story, and the same algorithm running across Web Workers in your browser.

My role
Pair assignment with Wei Zhao: we wrote the MPI program and the report together. I built the 2026 revival.
Team
Wei Zhao

Highlights

  • The original split the 18.74 GB file into byte ranges per MPI rank; wall-clock time fell from 11:01 on one core to 1:41 on eight (one run per layout).
  • While porting, found a chunk-boundary bug that can count a tweet twice (about 1 in 500 boundaries). The port keeps it, a test pins it and the lab explains it.

COMP90024 Cluster and Cloud Computing

Date:

Live

Social Sense

Team 57's cloud analytics project, which scored 2.4 million geotagged tweets and 1.7 million Mastodon toots for sentiment and compared them with official income and crime data, revived as a read-only site with maps, linked charts and a reproducible data pipeline.

My role
Built the React front end for Team 57 and worked on data processing and API tools. I built the 2026 revival.
Team
Xuan Wang, Wei Zhao, Zongchao Xie and Runqiu Fei

Highlights

  • Re-ran the team's own Python on the surviving inputs with 18 parity checks against the 2023 dashboard, producing a 1.7 MB read-only SQLite database.
  • Ported NLTK's tokeniser, WordNet lemmatiser and VADER to TypeScript; they match the original Python with no mismatches on 44,156 real toots.

COMP90051 Statistical Machine Learning

Date:

Live

Bandit Lab

My multi-armed bandit project, revived as a browser lab: successive-elimination bandits that cope with late or lost feedback, Thompson sampling and Doubly-Adaptive Thompson Sampling, evaluated in 2023 by replaying 300,000 recorded news clicks. Now they race on a seeded synthetic log whose delays you can reshape, up to seven algorithms side by side.

My role
Individual project. I wrote the original notebook and built the 2026 revival.

Highlights

  • Ported the five algorithms line for line to TypeScript running in a Web Worker, with NumPy's PCG64 generator, ziggurat normals and pairwise sums reproduced exactly, so the browser prints the notebook's results to the last digit.
  • Where the notebook had a single run, each algorithm now replays over 20 seeded repetitions with bootstrap bands, paired comparisons against Thompson sampling, effect sizes and delay-sensitivity heatmaps.

MAST90139 Statistical Modelling for Data Science

Date:

Live

GLM Playground

Three assignments on generalised linear models, revived as six interactive case studies: logistic, log-linear, multinomial, proportional-odds and GEE models refitted live in the browser and checked against the R code I submitted in 2023.

My role
Individual assignments. I wrote the original R Markdown analyses and built the 2026 revival.

Highlights

  • Reimplemented glm, nnet::multinom, MASS::polr and geepack::geeglm in about 1,300 lines of dependency-free TypeScript; 49 parity checks hold them to R's output, typically within 1e-9.
  • Each case study lets visitors change the inputs and recomputes estimates, intervals and tests on the spot, from beetle dose-response to a three-way table that shows Simpson's paradox.

2023, Semester 2

MAST90083 Computational Statistics & Data Science

Date:

Live

CompStats Playground

Two assignments on ridge and lasso, autoregressive model selection, support vector machines and the bootstrap, ported line by line from R. Drag λ, rerun a Monte Carlo study or retrain an SVM, and the site still prints the numbers in the original PDFs.

My role
Individual assignments (Assignments 1 and 3). I wrote the original R Markdown and built the 2026 revival.

Highlights

  • Ported glmnet's coordinate descent, LIBSVM's SMO solver, cv.glmnet and e1071::tune to TypeScript, with R's Mersenne-Twister seeding replayed bit for bit, so set.seed(10) draws the same folds and series as in 2023.
  • Vitest holds the ports to reference outputs exported from R and to the printed PDFs, from 1e-10 relative error up to exact equality of counts.

MAST90125 Bayesian Statistical Learning

Date:

Live

Four Roads to a Posterior

My Bayesian assignment fitting a logistic regression to a small dose-response trial four ways (Laplace, Metropolis-Hastings, HMC and expectation propagation), revived as an interactive explainer that replays the 2023 chains draw for draw.

My role
Individual assignment. I wrote the original R Markdown analysis and built the 2026 revival.

Highlights

  • Ported R's random number generator, mvtnorm and QUADPACK integration to TypeScript, so the browser matches the original R output to about 1e-11.
  • Found and documented five bugs in the 2023 data encoding and likelihood. An 'As submitted' toggle keeps the original behaviour beside the corrected analysis.

MAST90138 Multivariate Statistics for Data Science

Date:

Live

Multivariate Lab

Two assignments on covariance, PCA and high-dimensional classification, revived as three labs: a covariance and eigen playground, a wheat-seeds PCA explorer, and the story of how 365 days of rainfall tell northern and southern Australian weather stations apart.

My role
Individual assignments (Assignments 1 and 3). I wrote the original R Markdown and built the 2026 revival, which I am still extending.

Highlights

  • Ported a Jacobi eigen-solver, R's glm.fit with LINPACK's pivoting QR, kernel PLS and rpart's tree growing to TypeScript, checked against R in the test suite, so the browser recomputes the wheat PCA, every logistic fit and the PLS tree.
  • Wherever the 2023 work had a bug, the lab offers an 'As submitted' and 'Corrected' switch, and the covariance playground corrects one of the original answers.

2024, Semester 1

COMP90042 Natural Language Processing

Date:

Live

NLP Playground

My two NLP assignments rebuilt as three browser demos: a hashtag segmenter, a tweet geolocator using Naive Bayes and logistic regression, and Hangman played by n-gram language models. The 2026 upgrade adds intervals, paired tests, calibration checks and an optional LLM evaluation.

My role
Individual assignments. I wrote both original notebooks and built the 2026 revival.

Highlights

  • Ported NLTK's TweetTokenizer, the WordNet lemmatiser and both classifiers to TypeScript; the ports reproduce every test-set prediction from the submitted notebooks.
  • Re-analysis showed the two classifiers cannot be separated on 142 test tweets (exact McNemar p = 1.00), and that Naive Bayes is badly overconfident.

COMP90042 Natural Language Processing

Date:

Live

Climate Claim Checker

Our group's two-stage fact-checker for climate-science claims: TF-IDF retrieval over about 1.2 million Wikipedia passages, then a Transformer trained from scratch labels each claim. Revived so visitors can check a claim of their own or browse the 154 development claims, with the original pipeline re-run faithfully, weak results included.

My role
Group project with Wei Zhao and Xuan Wang. I worked on the system design, preprocessing, TF-IDF evidence retrieval, the report and the presentation; Wei built the Transformer and LSTM classifiers, and Xuan tested and debugged the retrieval, led the literature review and co-wrote the report and presentation. I built the 2026 revival, which I am still extending.
Team
Wei Zhao and Xuan Wang

Highlights

  • Ported every step of the 2024 notebook to TypeScript and checked it against the original Python, down to NLTK's Porter stemmer and the order in which NumPy breaks ties. Re-running the retrieval rule over all 1.19 million passages reproduces the reported development evidence score to every printed digit.
  • Re-running the code in 2026 showed that the Transformer encoder mixed claims within a batch instead of words within a claim, so on a single claim it reduces to a per-word lookup table, which makes every prediction on the site fully explainable.

03Skills matrix

Each skill lists the projects that show it, oldest first. Choose a skill to filter the timeline above.

In every projectFaithful ports checked against the original

Statistics

  • COMP10001 Playground · Message Cleanser · Data Processing Lab · Snacks in a Van · NYC Taxi 2019 · Source Separation Lab · 4399 CRM · Cachex Arena · Human or Machine? · Spartan Tweet Cruncher · Social Sense · Bandit Lab · CompStats Playground · NLP Playground

  • Message Cleanser · Data Processing Lab · Snacks in a Van · Source Separation Lab · 4399 CRM · Cachex Arena · Human or Machine? · Social Sense · Bandit Lab · CompStats Playground · NLP Playground

  • COMP10001 Playground · Snacks in a Van · Source Separation Lab · Spartan Tweet Cruncher · Bandit Lab · GLM Playground · CompStats Playground

  • Snacks in a Van · Cachex Arena · Spartan Tweet Cruncher · Bandit Lab · NLP Playground

  • NYC Taxi 2019 · Source Separation Lab · Social Sense · GLM Playground · CompStats Playground

  • Bandit Lab · Four Roads to a Posterior

  • Social Sense

  • Snacks in a Van

AI & machine learning

  • Data Processing Lab · Human or Machine? · CompStats Playground · Multivariate Lab · NLP Playground · Climate Claim Checker

  • Data Processing Lab · NYC Taxi 2019 · Source Separation Lab · Human or Machine? · Bandit Lab · GLM Playground · CompStats Playground · Four Roads to a Posterior · Multivariate Lab · NLP Playground · Climate Claim Checker

  • Data Processing Lab · Source Separation Lab · Multivariate Lab

  • Human or Machine? · Social Sense · NLP Playground · Climate Claim Checker

Algorithms

  • COMP10001 Playground · Cachex Arena

  • Message Cleanser · Data Processing Lab · Spartan Tweet Cruncher · NLP Playground · Climate Claim Checker

Testing & reproducibility

  • COMP10001 Playground · Message Cleanser

  • Data Processing Lab · Source Separation Lab · Bandit Lab · CompStats Playground · Four Roads to a Posterior

Data engineering

  • Spartan Tweet Cruncher · Social Sense

  • Data Processing Lab · NYC Taxi 2019 · Social Sense

  • Data Processing Lab

  • Snacks in a Van · NYC Taxi 2019 · 4399 CRM · Social Sense · Climate Claim Checker

Web

  • Snacks in a Van · 4399 CRM · Social Sense

  • COMP10001 Playground · Message Cleanser · Snacks in a Van · 4399 CRM

  • Snacks in a Van · NYC Taxi 2019 · 4399 CRM · Social Sense

GenAI evaluation

  • COMP10001 Playground · Message Cleanser · Data Processing Lab · 4399 CRM · Cachex Arena · Spartan Tweet Cruncher · Social Sense · Bandit Lab · NLP Playground

  • Snacks in a Van · Source Separation Lab · Cachex Arena · Spartan Tweet Cruncher · Social Sense · GLM Playground · Four Roads to a Posterior

  • COMP10001 Playground · Message Cleanser · Data Processing Lab · Snacks in a Van · NYC Taxi 2019 · Source Separation Lab · 4399 CRM · Cachex Arena · Spartan Tweet Cruncher · Social Sense · Bandit Lab · GLM Playground · CompStats Playground · Four Roads to a Posterior · NLP Playground