Work
Built, measured, and stress-tested
A production knowledge platform still leads the work. Around it are independent projects in model training, evaluation, retrieval, and AI infrastructure, each scoped by evidence rather than a demo claim.
On the roadmap
What I'm building next
These aren't shipped yet — listed honestly. The first items deepen the platform; alongside them I'm branching into independent tools that solve broader problems. Depth and range, not scattered demos.
- Finish the corpus: work the book-by-book validation queue that gates publication, and re-ingest + embed the excluded large book so semantic recall covers the whole library.
- If weight release becomes worthwhile, retrain from a provenance-clean data recipe under a fully captured environment and review the resulting checkpoint rights separately from the code license.
- Ship daily-hadith push notifications (designed, not yet built) and grow the Maestro end-to-end suite in CI, now that both public store listings are live.
- Deep-link each citation chip straight to its source page in the reader, now that the production corpus path is live.
- Push the learned-threshold work further: the leave-one-out experiment shows cheap features cap at recall 0.32, so the path is a vCache-style per-prompt boundary on a richer labelled set — can it close the gap to the judge without the judge call?
- Wire the owner-gated live-judge path (OpenAI / Anthropic adapters behind cost caps) and run the bias battery on fresh orderings: position, verbosity, self-preference, and test-retest probes.
- Extract the ingestion core into a config-driven, source-agnostic open-source framework with a typed SDK and CLI.
- More broadly — distributed-systems and real-user products in new domains, to build range beyond a single platform.