TokForge
Status: Live
Private AI that runs on your phone. No account, no cloud.
Links
TokForge runs large language models, image generation, voice and document search directly on Android phones and iPhones. Nothing leaves the device. It works on a $150 phone and on a flagship, because the app picks the fastest part of each phone (processor, graphics chip or AI chip) for each model.
What I built
- A native engine in Kotlin, Swift and C++ on top of llama.cpp and a fork of Alibaba’s MNN runtime that I maintain.
- GPU inference on Qualcomm, Samsung and Arm graphics chips. On a flagship phone it generated text 1.4 to 2.4 times faster than the leading competitor’s app.
- Streaming for mixture-of-experts models, so a 30-billion-parameter model runs on a 16 GB phone that could not load it before.
- Image generation on the Qualcomm AI chip without root access: a picture in 6 to 10 seconds.
- Document search and memory (RAG) that stays on the phone, with a reusable cache that cuts follow-up response time from seconds to under a tenth of a second.
- An auto-tuner that benchmarks each phone and saves the best settings, plus a public speed leaderboard.
- A test fleet of 12 Android phones and 4 Apple devices, and 11,000+ automated tests per release.
Built with
Kotlin, C++, Swift, llama.cpp, MNN, Vulkan, OpenCL, Qualcomm QNN, Core ML, MLX, TFLite