Data as of Aug 25, 2026 · Based on 38 AI responses from ChatGPT Search and Google AI Mode · See how Parse measures this
Your brand can be here too.
Yes. There’s a growing cluster of startups around small-model customization + private/on-device deployment. The most relevant ones I found are:
| Need | Most interesting |
|---|---|
| Fine-tune → ship to phones | Forgelet, Ertas |
| Small models / SLM platform | SmaLLM, Liquid AI |
| Robotics / embedded edge | Prysm Labs, Sutura |
| Private enterprise/on-prem | Passion Labs, Icosa |
| Custom engineering for mobile | Callstack |
One distinction is important: most of these still fine-tune in the cloud and deploy privately afterward. True training/fine-tuning on the device itself is a much smaller category. For example, the open-source MobileFineTuner project is explicitly working on native mobile LoRA/full fine-tuning, while MIT's PockEngine tackles sparse, resource-efficient on-device training.
If you're evaluating this as a startup/company landscape, I can also map these into funding, founders, customers, model sizes, deployment targets, and competitive differentiation.
A growing ecosystem of startups and specialized software tools focuses on making it fast, cheap, and memory-efficient to fine-tune small language models (SLMs) and tiny models (sub-1B to 9B parameters) for private or on-device deployment.
The standout startups and open-core tools in this space include:
(Note: While not independent venture-backed startups, heavyweights like Hugging Face with trl/torchtune and Google AI Edge with LiteRT and on-device quantization pipelines provide the foundational infrastructure that many of these specialized workflows plug into).
If you'd like to narrow this down, tell me:
I can help recommend the best toolchain or workflow for your specific project.
Yes. There’s a growing niche around small/specialist models that can be fine-tuned and then run entirely on a customer’s hardware. The startups I’d put on a shortlist are:
llama.cpp, Ollama, or LM Studio; it explicitly targets phone/local deployment. | Company | Best fit | Main differentiator |
|---|---|---|
| Ertas | App/product teams | No-code-ish fine-tuning → deployable model |
| Forgelet | Mobile/on-device apps | Tiny models + mobile exports |
| Icosa | Private knowledge work | Custom local models + local workspace |
| Microtensor | Model efficiency | Distillation into specialist models |
| SLM in a Box | Large enterprises | Full SLM lifecycle/on-prem |
| VDF AI | Regulated enterprises | Governed, on-prem fine-tuning |
| Sutura | Edge/audio | Runtime optimization + audio models |
| smelt.works | Mobile/desktop specialists | Very small task-specific LMs |
One important distinction: Ertas/Forgelet/smelt.works are closer to “fine-tune a model and ship it inside your product,” whereas SLM in a Box/VDF/Icosa are closer to “build a private model capability for an enterprise.” Microtensor is somewhat different again because its core bet is distillation/compression into specialist models rather than merely conventional LoRA fine-tuning.
If you're looking at this as a startup/competitive landscape, I’d pay particular attention to Ertas, Forgelet, Microtensor, and Icosa—they’re attacking the interesting intersection of fine-tuning + model compression + deployment on hardware the customer controls.
Several specialized startups, tools, and platforms help teams fine-tune small language models (1B–8B parameters) for private, secure, or on-device deployment.
Key Startups and Platforms
Comparison of Fine-Tuning Approaches for Private Deployment
| Tool / Platform | Primary Focus | Best Suited For |
|---|---|---|
| Unsloth AI | Ultra-fast, low-memory model training | Developers optimizing open-source 1B–8B models locally or via cloud notebooks |
| Databricks / MosaicML | Full-lifecycle private enterprise training | Secure, compliant corporate infrastructure and large-scale fine-tuning |
| Baseten | High-speed private model hosting & deployment | Production environments requiring ultra-low latency and data privacy |
If you'd like, let me know:
I can give you a more tailored recommendation on the right toolchain to use.
Yes. The market is getting quite interesting, but I’d separate true small-model/on-device specialists from broader private-LLM platforms.
| Startup | What they do | Best fit |
|---|---|---|
| Forgelet | Turns your data/spec into a fine-tuned 270M–3B model, exporting to GGUF, LiteRT, or Core ML for fully offline use. forgelet.tech | Phones, offline apps, embedded AI |
| Tunerbench | Fine-tuning platform explicitly built for small edge models, with QLoRA, quantization-aware workflows, Core ML export and air-gapped operation. www.tunerbench.compredibase.comforgelet.tech | Teams shipping edge AI |
| Ertas AI | No-code fine-tuning studio for open models, with GGUF export for local/on-device deployment. www.ertas.ai | Product teams without ML specialists |
| Smelt.works | Fine-tunes compact 1–3B mobile and 7–13B desktop models, emphasizing 100% on-device inference. smelt.works | Consumer/mobile apps |
| SmaLLM | Captures production data, fine-tunes/distills specialist SLMs, then deploys them in private cloud, edge, mobile or browsers. www.small-lm.com | Enterprise SLM lifecycle |
| SLMinabox | Builds task-specific small language models, with on-prem/private-cloud deployment and model ownership. www.slminabox.ai | Enterprise/private SLMs |
| Icosa Computing | Builds domain-specific models from company documents and runs them locally; its Zeno product is specifically aimed at private local AI. icosa.co | Teams wanting an owned local AI stack |
VDF AI is interesting if "private" matters more than literally running on a phone: it handles dataset preparation, private training, evaluation and deployment entirely on the customer's infrastructure.
InsightDLM similarly focuses on domain-specific SLMs/LLMs trained and served inside a company's network, including private cloud and air-gapped environments.
Predibase is the more established/general-purpose option: it supports fine-tuning specialized SLMs and flexible deployment "on our cloud or yours," although it's less specifically focused on consumer-device inference.
I'd put the startups into roughly four buckets:
1. "Fine-tune → ship to phone"
Forgelet, Tunerbench, Smelt.works, Ertas
2. "Build proprietary SLMs for enterprises"
SmaLLM, SLMinabox, InsightDLM, VDF AI
3. "Private AI appliance / local runtime"
Icosa, Cascadia, Sutura
4. "General fine-tuning infrastructure"
Predibase and similar platforms
The most differentiated niche appears to be the first one: taking a team from proprietary data → small fine-tuned model → Core ML/LiteRT/GGUF/edge deployment, rather than merely offering another cloud fine-tuning API. Tunerbench and Forgelet are especially close to that description.
If you're looking at this from an investment/competitive-intelligence perspective, I can also give you a table of ~20 companies with funding, founders, launch date, model sizes, target hardware, customers, and whether they're actually training models vs. just deploying them.
A growing sector of AI infrastructure and tooling startups focuses on adapting open-weight small language models (SLMs) and compact foundation models for private, localized, or on-device deployment . Rather than relying on massive, general-purpose cloud LLMs, these platforms specialize in parameter-efficient fine-tuning (PEFT), quantization, and hardware-aware edge optimization.
Specialized On-Device & Edge AI Platforms
Local & Enterprise-Grade Fine-Tuning Tools
If you are looking to narrow this down, let me know:
Several specialized startups and platforms help product and engineering teams fine-tune small language models (SLMs) for private, on-premise, or device-native deployment . Prominent options include Liquid AI , which provides device-native foundation models and an SDK to specialize and bake models directly into local runtimes; Predibase , offering a low-overhead declarative platform for fine-tuning and serving efficient open-source models; Together AI , which supplies managed decentralized infrastructure for parameter-efficient fine-tuning (LoRA/QLoRA) on open weights; and RunPod , which delivers on-demand, cost-effective GPU clusters purpose-built for fast SLM training and knowledge distillation.
Core Platforms & Tools
If you'd like, let me know:
I can help narrow down the ideal fine-tuning stack and quantization approach for your project.
Yes. The market is still fairly early, but there are several startups worth watching. I’d split them into purpose-built edge/on-device tuning and private/self-hosted fine-tuning:
| Startup | What they do | Best fit |
|---|---|---|
| Tunerbench | Fine-tunes small models specifically for edge devices, with QLoRA, quantization-aware workflows, device evaluation, and offline export. It targets models under ~200 MB and edge GPUs/NPUs. www.tunerbench.com | On-device LLM/assistant teams |
| Forgelet | Generates training data, fine-tunes small models, and exports to GGUF, LiteRT, or Core ML. Targets roughly 270M–3B parameter models and emphasizes fully offline execution. www.forgelet.tech | Mobile/offline apps |
| Liquid AI | Provides the LEAP stack around its small LFM models: fine-tuning, model bundling, and deployment through an on-device SDK. Its models range from hundreds of millions to ~1.5B parameters. leap.liquid.ai | Teams wanting an integrated edge stack |
| ModelNova | Edge-AI development environment covering model selection, transfer-learning fine-tuning, quantization, compilation, benchmarking, and deployment to embedded hardware. modelnova.com | Embedded/industrial AI |
| Sutura | Focuses on private on-device AI, including model optimization and fine-tuning for Whisper, TTS and audio models, plus INT8/INT4 optimization. sutura.computer | Voice/audio + edge |
| KIS.ai | Builds private, fine-tuned models that run on customers' own infrastructure; particularly focused on specialized enterprise workflows rather than generic LLMs. kis.ai | Private enterprise SLMs |
| Alphient | Its “Model Factory” creates small language models specialized to a company's workflows using its knowledge graph, documents and approved agent traces, then serves them privately. alphient.aikis.ai | Enterprise workflow specialists |
| InfoPlatform.ai | Fine-tunes open-weight models on company data while giving customers ownership of the data, weights and inference stack, including in-house deployment. infoplatform.aikis.ai | Private/on-prem enterprise AI |
For true on-device deployment: Tunerbench, Forgelet, and Liquid AI are the most directly aligned. Tunerbench and Forgelet are particularly interesting because they're attacking the fine-tuning → small model → device artifact workflow rather than simply providing another generic GPU training platform.
For private enterprise deployment: KIS.ai, Alphient, and InfoPlatform.ai are closer to the “your data → specialized small model → your infrastructure” proposition.
One important distinction: Together AI, Fireworks, Modal, etc. are relevant to fine-tuning/open-model infrastructure, but they're primarily cloud infrastructure rather than companies specifically optimizing small models for local/on-device execution. Together AI, for example, is increasingly positioning itself around open-model training and inference infrastructure.
If you're evaluating this as a startup/market landscape, I can also map these companies by funding, founding year, model sizes, target hardware (iPhone/Android/Jetson/PC), fine-tuning method, deployment model, and likely competitive moat.
Yes. This is becoming a distinct startup category, although it’s still quite early. The companies I’d put on a shortlist are:
| Startup | What it does | Best fit |
|---|---|---|
| Forgelet | Turns a task description + data into a fine-tuned small model, then exports it for phones, servers, or Raspberry Pi. Supports roughly 270M–3B models and exports to GGUF, LiteRT, and Core ML. | Best match for app teams that don't want an ML/MLOps team. forgelet.techwww.tunerbench.com |
| Tunerbench | Edge-focused fine-tuning/evaluation platform using LoRA/QLoRA, quantization, pruning and device-specific constraints. Supports offline/air-gapped workflows and production exports. | Best match for teams shipping embedded/on-device AI where latency, memory and compliance matter. www.tunerbench.com |
| Smelt.works | Specializes in fine-tuning compact models for mobile and desktop, with 1–3B models for mobile and 7–13B for desktop. Uses LoRA or full-weight tuning and targets zero-cloud inference. | Best match for highly specialized mobile assistants/features. smelt.works |
| distil labs | Fine-tuning platform focused on small language models, including systematic selection/evaluation of SLMs and deployment on private hardware. | Best match if you care about choosing the right SLM and squeezing maximum quality out of it, rather than strictly mobile. www.distillabs.ai |
| Callstack | More of an engineering services company than a pure platform: fine-tunes with LoRA and deploys private models on iOS/Android using ExecuTorch, MediaPipe, llama.cpp and MLX. | Best match if you need someone to actually build the on-device stack. www.callstack.com |
There are really three markets emerging:
The particularly interesting part is the move toward sub-3B models + PEFT + quantization + hardware-specific export. Recent research shows that parameter-efficient approaches such as LoRA can make adaptation dramatically more practical under edge memory/compute constraints.
If you're thinking about this as a startup/investment landscape, I'd pay especially close attention to Forgelet and Tunerbench. They are attacking the gap between generic fine-tuning platforms and the very different requirements of shipping a model that actually has to fit inside a phone, laptop, robot, or industrial device. Both are also early enough that the category isn't settled yet.
One caveat: "private deployment" and "private fine-tuning" aren't necessarily the same thing. Forgelet, for example, fine-tunes in managed cloud infrastructure but lets you run the resulting model entirely inside your own perimeter; Tunerbench advertises an offline control plane and air-gapped enterprise deployment.
If your target is specifically enterprise teams fine-tuning 0.5B–7B models on confidential company data and deploying them on-prem/edge, I can also map out the 10–15 startups competing in that narrower segment, including funding, founders, model sizes, customers, and how they differ.
Several specialized startups and open-source platforms help engineering teams fine-tune small language models (1B–8B parameters) for private Virtual Private Cloud (VPC), on-premises, or edge/on-device deployments . Key players in this space focus on parameter-efficient fine-tuning (PEFT), memory reduction, and local hardware optimization:
If you'd like to narrow this down, let me know:
I can help outline a step-by-step fine-tuning workflow using the right tool for your target environment.