00 / profileAI · data science · Hong Kong

Hi, I'm Matthew Thom.

I build AI projects that help explore research, evaluate models and understand data.

My work spans analytics, model evaluation and applied AI, with a focus on practical experiments and clear results.

about.sh~/matthew

whoami

Matthew Thom · data science / AI / Hong Kong

cat interests.txt

AI research / model evaluation / data science

open github

github.com/m21hm9

01 / knowledge graph

My digital brain.

A map of the tools, experiences and ideas behind my work. Switch between 2D and 3D to explore how they connect.

graph.view / interactive map68 nodes · 91 connections
Loading graph...

input: drag to explore · switch layouts · click the background to reset

02 / skills.index

Technical toolkit.

5 collections / education

The languages, tools and methods I use across AI and data science projects.

skills.json / toolkitread only
01

Languages

  • Python
  • Java
  • C++
  • SQL
  • JavaScript
  • TypeScript
02

AI/ML

  • PyTorch
  • TensorFlow
  • Scikit-learn
  • LangChain / LangGraph
03

Data & Cloud

  • Azure Databricks
  • PySpark
  • PostgreSQL
  • Supabase
  • MySQL
  • DBeaver
04

Web & DevOps

  • React
  • Next.js
  • Node.js
  • Git
  • SourceTree
  • Docker
  • Linux
  • CI/CD
  • Jira
  • Vercel
05

Tools

  • Figma
  • LaTeX
  • VS Code
  • PyCharm
  • Google Colab
  • Jupyter
  • HuggingFace Hub
education.md

Education

City University of Hong Kong

Bachelor of Science in Data Science, Minor in Psychology

Aug. 2024 – June 2028

Hong Kong

03 / work.log

Experience.

Selected roles and the work behind them. Open a role to see the details.

  1. log/01 · June 2026 – August 2026

    Summer Intern — Fintech

    The Bank of East Asia · Hong Kong & Shenzhen (Qianhai)

    • Built an end-to-end pipeline with Cantonese-English audio and TTS finance data to fine-tune Qwen-1.7B ASR.
    • Designed context-aware SFT with controlled context injection to preserve base model alignment.
    • Tuned hyperparameters with Optuna, substantially reducing WER and notably improving CER compared to the base Qwen-1.7B model.
  2. log/02 · June. 2025 – May. 2026

    Data Analysis Assistant

    AS Watson Group · Hong Kong

    • Performed QA testing and validation on AI models with trillion-scale datasets.
    • Built automated end-to-end testing frameworks using Databricks and PySpark.
    • Developed real-time dashboard that cut debugging time by 70-80%.
    • Achieved 100% story points completion every sprint.
  3. log/03 · Sept. 2025 – May. 2026

    AI Consultant

    Versed Digital Technology Limited · Hong Kong

    • Backed by Cyberport & VCs.
    • Created DCF financial models and pitch decks for valuation and fundraising.
    • Researched ViT and VLM models, achieving 99.9% accuracy and 30% lower human cost.
    • Built technical validation partnerships with VCs and universities.

04 / research.index

Projects and contributions.

The first four repositories pinned to my GitHub, spanning open-source molecular AI, research agents and machine learning experiments.

All repositories
research/ index.md04 entries

01 / 04

research/01.md

Open source / molecular AI

torch-molecule

Contributing to an open-source Python package that makes molecular prediction, generation and representation models easier to use.

// research question

How can molecular AI models become easier to use in practical research?

// my contribution

  • Integrated pretrained molecular generators, including NovoMolGen, MolGen, Molexar and SAFE-GPT.
  • Contributed dataset splitting modules for molecular machine-learning workflows.
20K+lifetime PyPI downloads for the package
#Python#PyTorch#Molecular-AI#Open-source
View project

02 / 04

research/02.md

AI agents / research tools

Autonomous Research Agent v0

A LangGraph research assistant that breaks a broad question into targeted searches and assembles a sourced report.

// research question

Can an AI agent turn an open-ended topic into a sourced research report?

// what I built

  • Uses DeepSeek to generate queries and Tavily to search them in parallel.
  • Reflects on research completeness before producing a Markdown report through a Streamlit interface.
#Python#LangGraph#DeepSeek-LLM#Tavily#Streamlit
View project

03 / 04

research/03.md

Machine learning / data pipelines

Time-Series Anomaly Detection

A modular pipeline for finding unusual behavior in univariate and multivariate time-series data.

// research question

How can different models detect unusual patterns in time-series data?

// what I built

  • Combines preprocessing, sliding windows and detectors such as Isolation Forest, One-Class SVM and autoencoders.
  • Includes threshold selection, evaluation metrics and visualizations for inspecting detection quality.
#Python#Scikit-learn#TensorFlow#Jupyter
View project

04 / 04

research/04.md

AI research / experiments

Reversible Semantic Anchor (RSA)

An experimental approach to preserving meaning when information moves between AI agents through a compact latent representation.

// research question

Can a compact latent anchor preserve intent across multiple AI handoffs?

// what I built

  • Built a six-layer Transformer encoder and decoder reconstruction pipeline.
  • Added training, drift evaluation and stress-test workflows for multi-step handoffs.
#Python#PyTorch#Transformers#Hugging-Face
View project

06 / contact.sh

open channel / collaboration

Get in touch.

Have a question or want to collaborate? I'd love to hear from you.

$ choose a channel →