
NobodyWho
Run text, vision and speech models locally on any device
About NobodyWho
NobodyWho is a library for running AI models on the device rather than calling an API, and its reach is the unusual part. Most local-inference projects target one runtime and one platform. This one exposes text, vision and speech models across React Native, Flutter, Kotlin, Swift, Python and Godot, which means the same approach covers a mobile app, a desktop tool and a game without adopting a different stack for each.
The pitch is stated plainly on the site: no API keys, no usage fees. For a developer that removes an entire category of problem, namely metering, rate limits, key rotation and a cost that scales with how successful your app becomes. It also removes the privacy conversation, because there is no cloud logging by design and nothing to log. Models come from Hugging Face in GGUF format, so the catalogue is thousands of models rather than a curated few, and optimised kernels for Metal, CUDA and Vulkan mean it uses whatever acceleration the hardware offers. Speech is included on both sides, text to speech and speech to text, so a voice interface can run without a network. It is published under EUPL 1.2, a copyleft licence approved by the European Commission, which is worth reading before shipping commercially.
Add the library for whichever platform you are building on, then point it at a GGUF model from Hugging Face. Because the format is standard, swapping models is a matter of changing the file rather than rewriting integration code, and you can size the model to the device instead of accepting one default. Inference runs locally using optimised kernels for Metal on Apple hardware, CUDA on NVIDIA and Vulkan more broadly, so a phone, a laptop and a smartwatch each use the acceleration they have. Text, vision and speech all run through the same local path, meaning a multimodal feature does not require adding a cloud dependency for one part of it. The application works offline as a consequence rather than as a special mode, and there is no key management, no per-request billing and no rate limiting to design around.
- •Six Platforms - React Native, Flutter, Kotlin, Swift, Python and Godot from one project
- •Text, Vision and Speech - Multimodal locally, including both speech-to-text and text-to-speech
- •Thousands of Models - Direct GGUF support, so most of Hugging Face is available
- •Hardware Acceleration - Optimised kernels for Metal, CUDA and Vulkan
- •No Keys, No Usage Fees - Cost does not scale with usage because there is no service to bill you
- •Private by Construction - Processing is local, with no cloud logging and nothing transmitted
- •Offline Capable - Works with no connectivity, since nothing is remote
- •Runs on Small Hardware - Smartwatches and phones as well as desktops
- •Open Source - Published under EUPL 1.2 with the full source on GitHub
Developers shipping AI features in apps where a per-request cost, a network dependency or a privacy commitment makes a cloud API unworkable. Mobile developers get the most direct benefit, since on-device inference removes both the latency and the running cost of calling out for every interaction. Game developers are an explicit audience given the Godot support, where an API call mid-frame is not viable. It suits anyone building for offline or intermittent connectivity, and anyone who has watched an inference bill grow with user numbers. Two things to weigh: local inference is bounded by the device, so model size is a real constraint on older hardware, and EUPL 1.2 is a copyleft licence rather than a permissive one, so check the obligations before building a closed product on it.
Pricing
No pricing page found
Free and open source under EUPL 1.2 with no API keys and no usage fees - runs entirely on the user device, so the only cost is the hardware you already have
Checked 2026-09-18
















