03 — Work
20 Selected Projects
Built across data engineering, cloud infrastructure, mobile, and web platforms.
Signal
Governance Layer for AI-Assisted Government Data
A governed data product that puts AI governance on the request path — tamper-evident hash-chained audit logs, auto-generated DTA and EU AI Act compliance artefacts, and faithfulness-checked LLM narratives. The live reference implementation analyses South Australian and NYC crime statistics with Mann-Kendall trend tests, Sen-slope forecasting, and z-score anomaly review.
Live deployment · 128 tests · Tamper-evident audit + DTA v2.0 governance set · Open-core product
Stack
Domain
AI / ML · Open Source
ClassBro
Tutoring and Mentoring Across Computer Science
Tutoring and mentoring university students across computer science subjects through ClassBro, from programming fundamentals to advanced topics in machine learning, AI, databases, systems and cloud computing. Working one-on-one with learners to bridge concepts and code.
10 subject areas · Subjects spanning programming, data science, AI, systems and more
Stack
Status
Active
Domain
Education
Ranking Radar
University Rankings Visualisation & Open Dataset
Brings four ranking systems (QS, Times Higher Education, and two U.S. News rankings) together for 3,790 universities across 40 years. A no-hard-coding pipeline pulls each ranking live and recovers QS history from archived edition IDs, published as an open, FAIR-licensed dataset.
3,790 universities · 4 ranking systems · live-sourced FAIR open dataset
Stack
Status
Live
Domain
Open Source · Research
Meal-Prep Studio Order SystemActive
Members, Prepaid Cards, Orders & Finance App
Built as sole developer for my mum's meal-prep studio: one system for members, prepaid meal cards, daily orders, kitchen and delivery, and finance. A single Expo codebase builds for iOS, Android and the web, on a Hono API on Vercel with Turso (libSQL) and Drizzle. It was built to replace group-chat sign-ups that were tallied by hand and re-keyed into an Excel workbook.
Sole developer · In production since April 2026 · One codebase for iOS, Android and web
Stack
Status
In production · Private
Domain
Personal
SAPOL EPSB AnalyticsActive
Professional Standards Reporting & Tooling
As ASO7 Senior Data Analyst in the Intelligence & Probity Unit of SAPOL's Ethical and Professional Standards Branch, produces the quarterly Use of Force and Vehicle Pursuit statistical reports, led an end-to-end review of the complaint administration workflow, and analysed a financial year of expiation notices. Also built a Python client and web console for the complaint-management system APIs, covering more than 1,100 endpoints.
Quarterly executive reporting · Admin workflow review · 1,100+ complaints API endpoints in one client
Stack
Status
In production
Domain
Government
EM Algorithm, Explained
An Interactive Expectation-Maximisation Lab
A maths-complete explainer of the expectation-maximisation algorithm that I wrote in September 2025, framed around sparse movie ratings, with a notebook that fits a two-component Gaussian mixture from scratch. In 2026 I rebuilt it as a web lab where you can step through the worked example by hand, run EM on the notebook's 200 ratings, and trigger label switching, stopping too early, local maxima and variance collapse. The TypeScript port reproduces the notebook's printed results, and the slips in my own hand-worked numbers are corrected beside the originals.
Notebook results reproduced to within 10⁻⁶ · Four pitfalls you can trigger · Corrections shown beside the originals
Stack
Status
Revived in 2026
- Revived
- Upgraded
- Guided tour
- Merged
Domain
AI / ML · Open Source
SA Gaming Machine Statistics
Public Gaming-Machine Figures, FY 2009/10 to FY 2024/25
South Australia publishes its gaming-machine statistics as separate PDF releases. In September 2025 I transcribed about 110 of them, from FY 2009/10 to FY 2024/25, into one workbook with a Power BI report. In 2026 I rebuilt it as a website on statewide revenue and tax, council areas, licences and manufacturers, with every figure checked against its PDF and the gaps in the series shown. It is a personal project on public releases from Consumer and Business Services, not affiliated with CBS, the Government of South Australia or my employer, and it takes no position for or against gambling. If gambling is affecting you or someone close to you, call the Gambling Help Line on 1800 858 858.
4,406 of 4,409 figures matched to their PDF · Yearly totals corrected · Gambling Help Line 1800 858 858
Stack
Status
Revived in 2026
- Revived
- Upgraded
- Guided tour
- Merged
Domain
Open Source
MapivaActive
Map-First Social Discovery App
Co-founded a Melbourne startup building a map-first app for discovering people and events nearby. As Dev Lead, owns the technical architecture (an Expo React Native client on a Django, Rust and PostgreSQL backend, shipped through GitHub Actions) and leads code review for a part-time engineering team.
Co-founder · Dev Lead · Beta planned for early 2027
Stack
Status
Beta 2027
Domain
Startup
SA Address Generator
Mock South Australian Addresses for Software Testing
A personal command-line tool I wrote in August 2025 that makes mock South Australian addresses for software testing, each with a suburb, postcode, council and remoteness level, and looks up real addresses through the Mapbox API. Porting it to the web in 2026 showed that its remoteness and socio-economic weights were never applied and its socio-economic column was empty. The revival keeps that uniform mode as built, adds the promised weighting from ABS remoteness and SEIFA data, and checks each sample against its target mix. Every output is labelled as synthetic test data.
20 of 20 recorded Python runs replayed exactly · Sample mix checked against its target · Labelled synthetic test data
Stack
Status
Revived in 2026
Domain
Open Source
US Political Data
Presidential Debate & Campaign Document Scraper
Two notebooks I wrote in August 2025 that collect public pages from The American Presidency Project at UC Santa Barbara, with a thread pool, retries and resumable runs: 7,556 campaign documents from 2016 to 2024 and 179 debate transcripts from 1960 to 2024. In 2026 I revived the collection as Campaign Text Lab, a descriptive reading room that reads the original CSV files without scraping again, with distinctive-word comparisons, debate talk shares and term timelines. Every speaker goes through the same code, and nothing is scored or predicted.
7,556 campaign documents · 179 debate transcripts, 1960 to 2024 · Descriptive only, nothing predicted
Stack
Status
Revived in 2026
Domain
Research · Open Source
CBS Intelligence
Regulatory Analytics from the Ground Up
Built the first intelligence analytics capability within the CBS Prevention Team — integrating ABS, SA Health, ACCC, and DataSA data into unified dashboards and GIS maps used by the Minister's Office.
Minister's Office reporting · SOPs institutionalised
Stack
Status
Completed
Domain
Government
Lecture Revision Lab
Revision Material That Cites Its Sources
A retrieval-augmented generation prototype I built in October 2024 to turn course notes into revision material, with DSPy, sentence-transformer embeddings and SQLite. In 2026 I rebuilt it to run in the browser. It finds the passages that answer a question and drafts revision agendas, worked examples and practice questions that cite those passages, and nothing is exported until a person has reviewed each item. Generation runs on the visitor's own API key, and the demo uses an openly licensed OpenStax statistics textbook instead of my private notes.
Retrieval matches the 2024 Python on 10 of 10 sample queries · Every item cites its passages · Reviewed before export
Stack
Status
Revived in 2026
Domain
AI / ML · Personal
Moodist
Mental Health Mobile Application
A clinician-facing and patient-facing mental health app for the University of Melbourne's Department of Psychiatry, built as sole developer. Rebuilt it from Uniapp to Expo React Native with a clinician dashboard on a Flask backend, kept hosting under $500 a month by running each service in its own Docker container on AWS LightSail, and kept it GDPR-aligned before handing it to a professional team for production.
Sole developer · Hosting under $500/month · Cross-platform iOS + Android
Stack
Status
Handed to production team
Domain
Research
Genomics Metadata Multiplexing
Sample-sheet tooling for plate-based single-cell sequencing
Worked on Genomics Metadata Multiplexing (GMM), an R Shiny tool that builds CEL-Seq2 sample sheets for plate-based single-cell sequencing. It merges flow-cytometry sort metadata from FACS index-sort files with the plate layout and primer indexes, so nobody has to stitch them together by hand. Wrote the Shiny side of the app, added test inputs with expected outputs, and contributed a column clean-up step to WEHIGenomicsRnD/celseq-sample-sheet-generator.
One sample sheet instead of a hand merge · Contribution merged into celseq-sample-sheet-generator
Stack
Status
Completed
Domain
Biotech · Research
Psyckitchen Echo
Memory-Aware Companion Chatbot Prototype
A June 2024 prototype of Echo, a companion chatbot for Psyckitchen built to remember a person's profile and past conversations. A Flask backend on Azure App Service kept both in MongoDB and sent them to GPT-4o with every new message. A Next.js test page sat alongside it, with a chat box and a mock wearable display. The prototype has since ended.
Prototype, since ended · 8 to 14 June 2024 · Flask and MongoDB backend with a Next.js test page
Stack
Status
Prototype · Ended
Domain
AI / ML
Climate Fact-Checker
Automated Fact-Checking for Climate Claims
A two-stage fact-checker for climate-science claims, built for COMP90042 with Wei Zhao and Xuan Wang. TF-IDF retrieval pulls evidence from about 1.2 million Wikipedia passages, then a Transformer trained from scratch labels each claim as supported, refuted, disputed or not enough information. I worked on the system design, the preprocessing and the TF-IDF retrieval, and Wei built the classifiers. I revived it in 2026, with the original pipeline re-run as it was and the weak results shown.
TF-IDF retrieval over 1.19 million passages · Reported evidence score reproduced in 2026 · Weak spots stated plainly
Stack
Status
Completed
Domain
Climate Research
HEX
Digital Content Library Builder
Built and refined interactive digital educational content aimed at bridging the gap between education and professional success. Conducted market research on leveraging advanced technologies to improve course engagement and learner outcomes for students across Australia and beyond.
80+ students supported · Improved course engagement through technology-driven content innovation
Stack
Status
Domain
Startup
Self-Driving Databases
Workload-Driven Optimisation & Index Selection
A survey of self-driving databases for COMP90050, written by Group 40 with Xiaoyi Liu, Runqiu Fei and Qingxuan Yang. Runqiu and I covered index selection, from heuristics to bandits, and Xiaoyi and Qingxuan covered workload-driven optimisation. I revived it in 2026 as Self-Driving DB Lab, where the index advisors the survey described race on a live SQLite database in the browser.
Group 40 survey · Five index advisors implemented in 2026 · Revived as a browser lab
Stack
Status
Completed
Domain
AI / ML
Twitter HPC Analysis
Big Data on SPARTAN
Processed a large-scale Twitter dataset on the University of Melbourne's SPARTAN HPC cluster using MPI and Python, identifying tweet distribution across Australian cities and top contributors.
HPC parallel processing · City-level tweet distribution insights
Stack
Status
Completed
Domain
Cloud / HPC
ENSO Climate Risk
Climate and Staple Food Prices, a Public-Data Revival
This was our MAST90106 and MAST90107 industry capstone, hosted by CSIRO in 2023. I worked in Group 36 with Ritwik Giri, Jiaqi Hu, Xiangyi (Emma) He and Jihang (Jonathan) Yu. In 2026 I rebuilt it from scratch as a clean-room revival on public World Bank, NOAA and FRED data, with no CSIRO or team data, code or documents. It asks whether lagged climate indices such as El Niño help forecast next month's prices for wheat, maize, rice, soybeans, palm oil and sugar, comparing forecasts with and without the climate input on months the models never saw. None of the eight commodities forecast better with it.
Team of five · 6 climate indices and 8 commodities tested · Public data only (World Bank, NOAA, FRED)
Stack
Status
Revived in 2026
Domain
Climate Research
40 Reps, Re-checked
A Machine Learning Notebook Gallery, Audited
Forty projects I worked through for practice in late 2022, one notebook each, following Aman Kharwal's 180 Data Science and Machine Learning Projects with Python series on thecleverprogrammer.com. The project ideas and methods come from his tutorials. In 2026 I turned them into a study log that records which notebooks ran, which broke and what each printed, plus eight audits where a little more rigour changes the conclusion, such as stock forecasts that lose to a flat forecast once the split respects time.
40 notebooks logged, 33 ran end to end · 8 audits against baselines and time-ordered splits · Tutorial source credited
Stack
Status
Revived in 2026
Domain
AI / ML · Personal
Cachex AI
Game-Playing AI Agent
A pair project with Wei Zhao as team _4399 for COMP30024. Cachex is a two-player connection game like Hex, with captures and a steal rule. We wrote an A* solver for the search task and a minimax agent with alpha-beta pruning to play the full game. I revived it in 2026 as Cachex Arena, where you can play the agent or run a seeded tournament.
A* solver · Minimax agent with alpha-beta pruning · Pair project, revived as a playable arena
Stack
Status
Completed
Domain
AI / ML
Project Cradle
Invite-Only Member Hub
A small community you can only join with an invite code from a member, where each member holds at most two unused codes. Jiahong Zheng and I built it in early 2022 to learn a full stack, with an Express and MongoDB API behind a Next.js front end. The database is long gone, so in 2026 I revived it as one Next.js app on SQLite (Turso), with the 2022 invite rules ported and checked by parity tests, an invite tree, demo logins and a guided tour.
Pair project with Jiahong Zheng · 2022 invite rules ported with parity tests · Demo logins and a guided tour
Stack
Status
Revived in 2026
Domain
Personal
HPLC Pipeline
Lab-Data Quality, an Industry Capstone with CSL
This was our MAST30034 industry capstone with CSL. I worked in Group 07 (Team 4399) with Yuchen (Cynthia) Cai, Rongshun Li, Zhiliang (Tommy) Tang and Quzihan (Martin) Wu, and was the team's Scrum Master. We parsed free-text HPLC lab-result spreadsheets into tidy records in the spirit of the FAIR principles, rebuilt control charts for four assays to find out-of-control runs, and compared PCA, t-SNE and UMAP with K-Means and DBSCAN as an aid for spotting unusual runs. I revived it in 2026 as HPLC QC Lab, which runs the team's methods on synthetic data, with no CSL data, code or documents.
Team of five · About 99.7% of lab-result records parsed in 2021 · Revived on synthetic data only
Stack
Status
Completed
Domain
Biotech
PCRM
Personal Customer Relationship Management
A mobile-first personal CRM built by Team 4399 for the COMP30022 IT Project, with Bin Liang, Hongji (Harrison) Huang, Wei Zhao and Yixiao Tian. It has a React front end, an Express REST API and a MongoDB database. I was the team's Scrum Master and worked on the contacts, meeting records and map screens. I rebuilt it in 2026 as 4399 CRM, a single Next.js app with a guest sandbox.
Five-person Scrum team · Full-stack web app · Revived in 2026 as 4399 CRM
Stack
Status
Completed
Domain
Research
NYC Taxi Analysis
Quantitative Analysis with Spark
My individual project for MAST30034 Applied Data Science. I cleaned every 2019 New York yellow-cab trip in PySpark, more than 84 million records, and fitted an elastic-net regression to estimate trip times. I revived it in 2026 as NYC Taxi 2019, with zone maps, weather, events and collisions, and the 2021 regression running in the browser.
84 million trips cleaned in PySpark · Elastic-net regression of trip times · Revived as a live demo
Stack
Status
Completed
Domain
Research
rin.contactActive
Personal Portfolio, v5
Five major iterations of a personal website — currently rebuilt in v5 with a Nothing OS-inspired minimal design language. A living record of technical and professional development since 2020.
5 major versions · rin.contact · Open source
Stack
Status
Live · v5.27
Domain
Personal
Virtual internships (Forage)
Twelve Company-Designed Job Simulations
I completed twelve self-paced job simulations on Forage, each designed by the company that set the brief. The companies were KPMG, BCG, British Airways, Quantium, Tata, Cognizant, GE Aviation, Accenture, Red Bull, PwC (two) and Standard Bank. The work covered customer analytics, churn and booking models, dashboards and data joins.
12 job simulations · Oct 2022 – Aug 2023 · analytics, ML and visualisation
Stack
Status
Completed
Domain
Personal
29 projects
In the pipeline
Planned and in progress
Work I am writing up or rebuilding next. Once one ships, it gets a live demo on a card in the main list.
Public Services: Open Source Data
In progress · as of Oct 2026
Personal · Open Source · Curated SA/AU government datasets
Curated collection of South Australian and Australian government datasets, cleaned and documented for public use. Makes hard-to-find public data discoverable and ready to use.
RepositoryGmail Labeler
In progress · as of Oct 2026
Personal · Productivity · Gmail triage with auto-labels
A Gmail triage tool that auto-labels incoming mail by sender, subject patterns, and content, helping you reach inbox zero faster. Runs on your own Gmail account with your own labeling rules.
Repository
Social Media Analytics
Australia on the Cloud
Harvested and analysed Twitter and Mastodon data alongside ABS SUDO spatial data to produce a Social Sense Dashboard illuminating Australian sentiment, social trends, and regional behavioural patterns.
Cross-platform social analytics · Spatial + social data fusion
Stack
Status
Completed
Domain
Cloud / HPC