Skip to content

Ali Abbas al-Muhammadi

Full-Stack, Mobile & AI Engineer · Search & Retrieval

ali@alimuhammadi.com0451 128 256Melbourne, Australiaalimuhammadi.comgithub.com/aliabbas-muhammadilinkedin.com/in/aliabbas-muhammadi

Full-stack, mobile & AI engineer: final-year Monash IT & Commerce student (graduating 2026) who builds and operates production web and mobile apps, and trains and evaluates ML models for search and retrieval systems. Seeking new-grad / intern software & AI roles, Melbourne or remote.

Projects

Shia LibrarySearch & retrieval platform

2025 – present
  • Built and operate a live multilingual search platform over a ~315k-record corpus: diacritic-insensitive Arabic search (custom SQL normalization + GIN trigram indexes), hybrid keyword + vector retrieval (pgvector), and a citation-grounded "Ask" (in beta) scoring 91% of claims grounded, 100% out-of-corpus abstention (30-question eval).
  • Built and shipped its offline-first React Native (Expo) app to the App Store and Google Play from one TypeScript codebase: crash-safe SQLite/MMKV storage, forward-only migrations, and custom native modules in Swift, Kotlin, and Objective-C++ for background downloads and reader-scroll coordination under each OS's limits.
  • Ship both surfaces from one Supabase/Postgres backend with row-level security on every public table, 4,900+ automated tests across web and mobile, and a gated, probe-verified release pipeline that publishes to both public app stores.

Fed by a crash-safe ingestion + LLM-translation pipeline (220 books processed, 41 published, ~90% prompt-cache savings).

Next.js · React Native · TypeScript · PostgreSQL / pgvector · Claude · Swift / Kotlin / Obj-C++

PolarityMachine learning & evaluation

2026
  • Built and fine-tuned a 184M-parameter DeBERTa-v3 cross-encoder for semantic-cache equivalence after historical evaluation showed embedding similarity could rank a meaning-changing query above valid paraphrases.
  • Diagnosed complete training collapse through metric validation, a controlled fp32 experiment, and model-pipeline tracing; identified a Transformers-version-sensitive failure and recovered custom low-FP step-pAUC 0.716 versus 0.048 for the NLI baseline.
  • Stress-tested the frozen validation-selected threshold on an author-created 75-pair target-domain set: recalled 31/31 equivalent pairs with 2/44 false accepts and 0/22 negation false accepts, then documented weak MRPC transfer at ROC-AUC 0.640 instead of claiming general robustness.

Released source-only with 416 passing tests across Python 3.12 and 3.13; model weights remain private.

Python · PyTorch · Hugging Face Transformers · DeBERTa-v3 · scikit-learn · pytest / mypy

Citation-Grounded RAGAI / retrieval systems

2026
  • Built a reference RAG engine with hybrid retrieval (BM25 + dense + RRF) and Claude-native citations tied to verifiable source spans: recall@5 0.91, faithfulness 1.00 (Claude judge) on a 22-question golden set.
  • Designed abstention as a feature (3/3 out-of-corpus questions declined) with a CI gate that fails the build on retrieval regressions.

The public, keyless demo and evaluation harness for the retrieval-and-citation engine used by Shia Library's beta Ask system.

Next.js · TypeScript · Claude · OpenAI · BM25 + RRF · Vercel

JudgelabAI evaluation & methodology

2026
  • Built an open-source (MIT), Python-first lab measuring LLM-as-a-judge reliability: reproduced the MT-Bench GPT-4-vs-human agreement and added the chance-corrected statistics the original omitted, including Cohen's kappa 0.767 on decisive cases (0.505 including ties) with bootstrap confidence intervals, benchmarked against human-human agreement of Krippendorff's alpha 0.485.
  • Made the benchmark fully keyless and reproducible, computed from a committed CC-BY-4.0 snapshot and re-derived byte-for-byte in CI as a drift gate, with 127 tests, strict typing, and every statistic cross-checked against scikit-learn and SciPy.

Python · NumPy / SciPy · pytest · mypy · GitHub Actions

Skills

Languages
TypeScript, JavaScript, SQL, Python, Java, Swift, Kotlin, Bash
Frontend
React, Next.js (App Router), Tailwind / MUI
Mobile
React Native (Expo), Swift / Kotlin / Objective-C++ modules, accessibility
Backend & data
Node.js, PostgreSQL, Supabase / pgvector, full-text search, row-level security
AI, ML & retrieval
PyTorch, Hugging Face Transformers, supervised fine-tuning, model evaluation & calibration, RAG, hybrid retrieval, LLM judge reliability
Infra & practices
Vercel, CI/CD (GitHub Actions), Vitest / Playwright / Maestro, Sentry, Git, secret scanning

Experience

Digital Testing Officer, NAATI

Jun 2025 – Present
  • Support delivery of high-stakes online language examinations, triaging connectivity, browser, and platform failures in real time; mined ticketing data for recurring failure patterns and drove the resulting workflow change, resolving 80%+ of issues at first contact.

Client Engagement Officer (ATO), ProbeCX

Jan 2023 – Jun 2025
  • Resolved complex technical and account issues across Australian Taxation Office systems while sustaining 95% CSAT over 18 months and mentoring new staff.

Education

Bachelor of Information Technology & Commerce (Double Degree), Monash University

2023 – 2026

Officer: Monash Association of Coding · Computing & Commerce Association.