
MODULE · Local LLM inference

In a few clicks you bring LLM inference inside your perimeter: on-premise, private cloud or air-gapped. Zero egress, zero dependency on external APIs: every ARI module relies on ARIrun from minute one.
Scales by:
GPU class and number of instances *
·
from €270/month
How it works
01
Pick the model from the catalog and activate it
Select the model from the supported Hugging Face library. One click and ARIrun puts it into service on your hardware.
02
The whole suite, with maximum privacy
ARIchat, ARIknow, ARIdb, ARIdoc and ARIflow hook into ARIrun with no integrations to write. The architecture stays the same: you only change where inference runs.
03
Zero data outside the perimeter
From here on, every AI call in the suite runs inside your perimeter. Monitoring included, zero traffic to external servers.
The whole suite
No separate integration to maintain: each module connects on its own, and no extra data leaves your perimeter.
ARIknow
Embeddings and retrieval computed locally, RAG that never leaves the perimeter.
ARIdb
NL→SQL generated by models running on your infrastructure.
ARIdoc
Extraction and semantic OCR processed internally, documents never sent out.
ARIflow
Every LLM node in your workflows points to ARIrun, private end-to-end orchestration.
ARIchat
Multi-user chat on local models, included for free, zero extra setup.
Included for free
ARIchat is ARI's multi-user chat UI: included at zero cost with any ARIrun purchase. Persistent sessions, user management, routing to the models configured on your runtime. Everything runs locally: no token leaves your infrastructure.
Features with ARIflow
When you build workflows with ARIflow, every node that calls an LLM can point to ARIrun. Orchestration, automated tests and deployment happen without a single token leaving your infrastructure.
Local inference in workflows
Every ARIflow step that invokes a model uses ARIrun as backend. No external API calls in production workflows.
Model switch with no downtime
Update the model used by a workflow with one click. ARIrun handles the gradual rollout: the old model keeps responding until the new one is ready.
Testing across multiple models
ARIflow's testing system can compare the same workflow across different models in parallel. Choose the best model for each task with real data.
Unified monitoring
ARIrun metrics (latency, throughput) are visible directly in the ARIflow dashboard. A single view over workflow and inference-layer performance.
Modules can be purchased individually. ARIflow is optional and expands each module's capabilities.
Pricing
Prices per on-premise GPU instance: ARIrun software on your hardware. Also available in SaaS mode on BlackBytes cloud (hardware included), with pricing on request. ARIchat is included for free with ARIrun.
✓ ARIchat included for free
✓ Also available as SaaS (on request)
✓ Monthly or annual billing
Class S
€270
/instance / month
24 GB VRAM · Gemma 4 26B, Mistral Small 24B, Qwen3-35B
✓
Models ~26–35B parameters (Q4)
✓
~3–5 req/min single-user throughput
✓
Ideal for privacy-first chat and KB QA
✓
ARIchat included for free
✓
OpenAI-compatible endpoint
✓
Monitoring and audit trail included
Class M
€530
/instance / month
48 GB VRAM · Llama 3.3 70B, Qwen3-72B, Mistral Small 4 119B
✓
Models ~70–120B parameters (Q4)
✓
~100 concurrent active users
✓
Quality comparable to GPT-4o-mini
✓
Configurable multi-model routing
✓
Autoscaling and load balancing
✓
Support for custom fine-tuned models
Class L
€1,100
/instance / month
80 GB VRAM · Llama 4, Qwen3-72B FP8, DeepSeek V4
✓
Frontier-adjacent models (FP8 / MoE)
✓
~250+ concurrent users
✓
Quality comparable to Claude Haiku / Gemini Flash
✓
Certified air-gapped environments
✓
Premium SLAs available
Class XL
€2,700
/instance / month
180+ GB VRAM · multi-node cluster deploy
✓
Frontier models on a dedicated cluster
✓
Multi-node deploy and horizontal scaling
✓
Maximum throughput and concurrency
✓
Certified air-gapped environments
✓
Premium SLAs available
* The prices shown are per single on-premise instance (ARIrun software on your GPU) and scale by class (S / M / L / XL) and by number of active instances. The tier is calculated on the peak number of active instances during the billing period. Multi-GPU bundle discount: -10% from the 2nd instance, -15% from the 4th instance onward (same class). Also available in SaaS mode (BlackBytes cloud, hardware included): pricing on request.
Deployment
On-Premise
Installed on your servers. Data never leaves your perimeter.
Air-Gapped
Fully isolated from the network. For high-security environments.
Upon request
SaaS
Hosted by us. Zero infrastructure to manage.
Details on privacy, compliance and architecture → Deploy & Privacy
Bundle discount
Each additional module beyond the first gives you 5% off the total: up to 15% with four modules. Add ARIflow and another 10% kicks in.
The maximum discount is 25%: four modules plus ARIflow.
Example: ARIdb + ARIknow + ARIflow → 15% discount. €23,400/year become €19,890.
ARIchat is included for free and does not count toward the bundle.
1 module
0%
2 modules
5%
3 modules
10%
4 modules
15%
+ ARIflow
+10%
ARIchat, ARIknow, ARIdb and ARIdoc can be tried for free: same suite, same privacy.
Still have doubts?
Book a quick consultationCookies and tracking
We use technical cookies and — with your consent — analytics and marketing cookies to measure traffic and improve your experience. You can accept, refuse or change your choice at any time.
More info in the Cookies Policy →