Projects

TokForge

Status: Live

Private AI that runs on your phone. No account, no cloud.

Links

TokForge runs large language models, image generation, voice and document search directly on Android phones and iPhones. Nothing leaves the device. It works on a $150 phone and on a flagship, because the app picks the fastest part of each phone (processor, graphics chip or AI chip) for each model.

What I built

  • A native engine in Kotlin, Swift and C++ on top of llama.cpp and a fork of Alibaba’s MNN runtime that I maintain.
  • GPU inference on Qualcomm, Samsung and Arm graphics chips. On a flagship phone it generated text 1.4 to 2.4 times faster than the leading competitor’s app.
  • Streaming for mixture-of-experts models, so a 30-billion-parameter model runs on a 16 GB phone that could not load it before.
  • Image generation on the Qualcomm AI chip without root access: a picture in 6 to 10 seconds.
  • Document search and memory (RAG) that stays on the phone, with a reusable cache that cuts follow-up response time from seconds to under a tenth of a second.
  • An auto-tuner that benchmarks each phone and saves the best settings, plus a public speed leaderboard.
  • A test fleet of 12 Android phones and 4 Apple devices, and 11,000+ automated tests per release.

Built with

Kotlin, C++, Swift, llama.cpp, MNN, Vulkan, OpenCL, Qualcomm QNN, Core ML, MLX, TFLite

SurvivalOps

Status: In development

A survival and medical guide that works with no signal.

Links

SurvivalOps is an Android app that answers survival, medical and practical questions from a library stored on the phone, and shows where each answer came from. It runs a 4-billion-parameter model on the device and needs no network at all. The library holds more than 1,400 documents and 176,000 searchable passages, plus offline encyclopedias and topographic maps with GPS.

What I built

  • Hybrid search (keyword plus meaning-based) with a refusal rule: if the library does not cover a topic well enough, the app says so instead of guessing.
  • Pre-computed answer cards for the most common emergencies, so the most important answers come back in about a second with no model involved.
  • A 279-question graded test set with a hand-checked AI judge, used to measure every change.
  • Signed content packs (Ed25519, SHA-256) so a tampered pack cannot install.

Built with

Kotlin, C++, Python, MNN, sqlite-vec, Kiwix

Folio

Status: Pre-alpha

An AI game master for long stories that remembers what happened.

Links

Not public yet

Folio is a phone-first app for interactive stories with an AI game master. The hard problem is memory: a story can run to hundreds of thousands of words, and the AI has to keep its facts straight. Folio keeps a running ledger of what happened, searches past scenes by meaning and by keyword, and checks the AI’s drafts against the record.

What I built

  • A memory layer that raised continuity recall from 66% to 86% on a 152-question benchmark. Simply loading the whole story into the model scored 40%.
  • Serving for a 321-billion-parameter model on two GPUs at 150 to 190 tokens per second.
  • A character-portrait pipeline with two safety tiers and a blind review gate.
  • Speech input and a cloned narration voice.

Built with

Python, FastAPI, PostgreSQL, pgvector, Preact, SGLang

Elevated Marketing hosting platform

Status: Live

Hosting for 18 client websites across three data centers, with automatic failover.

Links

The platform behind Elevated Marketing’s client sites. Two servers in Michigan work as a pair: if the main one fails, the backup takes over on its own within about three minutes. Copies of every site and database also live in Texas and Ohio, and backups go to storage that ransomware cannot alter.

What I built

  • A zero-downtime migration of all 19 sites to a new data center.
  • Fenced automatic failover, proven by pulling the power in a live drill (database over in 69 seconds).
  • A private network linking six sites, with a direct tunnel that cut Ohio-to-Michigan latency from 60 ms to 21 ms.
  • An AI tool that reads server error logs, explains the problem on a Grafana dashboard, and fixes common issues on its own within safety limits.
  • Attack blocking that actually blocks: a firewall rule fix took automatic bans from zero to working within minutes, against 1,600+ logged attacks.

Built with

Linux, VMware ESXi, VyOS, WireGuard, Cloudflare, MySQL replication, Veeam, MinIO, Prometheus, Grafana, vLLM

Niku POS

Status: Pilot

A restaurant point-of-sale system built in four weeks.

Links

Not public yet

A point-of-sale system for a sushi restaurant, with card-terminal payments, kitchen display, receipt printers and a cash drawer. Every sale is stored as a sequence of events, so the system can always rebuild its state and nothing is lost if a device goes down. I set the plan and the quality gate and directed AI coding agents to build it: 1,077 commits in under four weeks.

What I built

  • Event-sourced design on PostgreSQL with a 66-step automated test gate.
  • Helcim and Stripe card-terminal payments with reconciliation.
  • Printer recovery that finds a printer again when its network address changes.

Built with

TypeScript, Fastify, PostgreSQL, Helcim, Stripe Terminal

Home AI lab

Status: Running

A self-hosted AI platform that serves models up to 321 billion parameters.

Links

Not public yet

The server behind my AI work: three NVIDIA GPUs and 377 GB of memory, running a custom model router that swaps models in seconds instead of minutes and heals itself when a backend hangs. I use it to serve models, train LoRA adapters, score retrieval experiments, and run a voice assistant for the house.

What I built

  • Speculative decoding tuned across four serving engines: 185 tokens per second on a 125-billion-parameter model, 3.4 times faster than the baseline.
  • LoRA training on two 96 GB GPUs for a 321-billion-parameter model, by streaming expert weights from system memory.
  • A home voice assistant (wake word, Whisper speech-to-text, a local model, a cloned voice) tied into Home Assistant.

Built with

vLLM, SGLang, llama.cpp, ExLlamaV3, PyTorch, PEFT, Home Assistant

Other projects

XMRAnd Mobile

Status: Open source

A CPU-mining app for Android and iPhone, built to learn how far phone processors can be pushed.

Links

XMRAnd Mobile is an open-source fork of the xmrig miner for Android and iOS. The hard part was getting a just-in-time compiler to run on a stock, non-jailbroken iPhone. It also pauses itself whenever TokForge is running, so the two never fight for the processor.

What I built

  • RandomX just-in-time compilation on stock iPhones: about 2,000 hashes per second on an iPhone 15 Pro Max and about 2,900 on an iPhone 17 Pro Max.
  • Automatic pause and resume whenever TokForge is in use.
  • A personal-signing and setup flow so it installs without a jailbreak.

Built with

C, C++, Objective-C, Swift, xmrig, RandomX