Local Model Self-Hosting and Recommendations
Recommendations on hardware and infrastructure for self-hosting models for AgentOS and the Agent Service Bus (ASB).
Table of Contents
Recommended “Good Production” Build 2
Why this configuration works well 3
Lower-cost option for gpt-oss:120b (Budget Build) 3
Recommended Production Hardware for Full Kimi K2 (Serious Enterprise Build) 4
More realistic recommendation for Kimi 5
Best balance of cost/performance 6
Inference Engine Recommendations 7
Database WAL Write Capacity by Hardware Tier 9
What the ASB Does Under Load 9
Tier 1 - Small Team / Lab (up to ~50 agents) 9
Option A: Mini PC (Best Value) 9
Option B: Raspberry Pi 5 (Ultra-Low Power, Air-Gapped Deployments) 10
Tier 2 - Department / Mid-Size Deployment (50-500 agents) 11
Tier 3 - Enterprise / Data Center Rack (500+ agents, HA, Compliance) 12
ASB Hardware Capacity Analysis 14
Note: The following information assumes hosting with 10 or more concurrent users. For 10+ concurrent users, the biggest factor isn’t just fitting the model into VRAM - it’s maintaining enough throughput and KV-cache headroom so latency doesn’t become painful once several people are generating at the same time.
This document primarily focuses on the gpt-oss:120b model as it has been tested and confirmed as working sufficiently with AgentoS. Kimi K2 has also been included since it has been mentioned as being evaluated by some customers. Additional models and server-sizing recommendations will be added as they get tested by Lucus Labs for use with AgentOS.
Model: gpt-oss:120b
VRAM requirements
Bare minimum usable VRAM: ~60-80 GB
Comfortable production VRAM: ~96-160 GB
10 active users: multi-GPU strongly recommended
Community deployments commonly recommend 4x RTX 3090s or enterprise GPUs for this class of workload.
Recommended “Good Production” Build
Why this configuration works well
VRAM headroom
gpt-oss:120b at Q4 needs roughly:
60-73 GB for weights
Plus KV cache
Plus batching/concurrency overhead
With 192 GB total VRAM you can comfortably support:
10 users
Larger context windows (including gpt-oss:120b’s 132k token ctx)
Batching
Agent workflows
RAG pipelines
Expected performance
Approximate:
80-180 tokens/sec aggregate
~10-25 tokens/sec per active user
Depends heavily on:
Context size
Batching
Quantization
Prompt length
Lower-cost option for gpt-oss:120b (Budget Build)
This is the most common “prosumer lab” approach and acceptable for starting out and testing ideas & applications. Many users mention 4x 3090 setups as the “economical route” for 120B-class models.
Tradeoffs:
More power draw (~1500W GPU-only)
More heat/noise
PCIe/riser complexity
Lower reliability than enterprise GPUs
Model: Kimi K2
The full Moonshot AI Kimi K2 model is described as:
~1 trillion parameter MoE model
~32B active parameters
~500-620+ GB VRAM requirement (even when quantized)
This is in a completely different infrastructure tier than the gpt-oss:120b model.
Recommended Production Hardware for Full Kimi K2 (Serious Enterprise Build)
Why this is required
Kimi K2 Q4 quantization alone is estimated around:
500-620 GB VRAM
Before large KV cache overhead
For 10 concurrent users:
You ned huge KV cache memory
Fast interconnects
Aggressive batching
Tensor parallelism
this is effectively:
A “mini AI datacenter”
Likely a $180K-$350K server depending on GPU pricing
More realistic recommendation for Kimi
If your goals is:
Internal coding assistant
Enterprise chatbot
AI agents
RAG
Code review
Automation
… then the full Kimi K2 is usually not economically sensible to self-host.
Instead:
Best Practical Option
Run:
Kimi K2 Distill 32B
or gpt-oss:120b
on:
2x RTX 6000 Ada
or 4x 3090
You’ll get:
Dramatically lower latency
Massively lower power costs
Simple infrastructure
Far easier scaling
Much better reliability
The distilled Kimi variants are specifically intended for practical local deployment.
Recommendation by Budget
Best balance of cost/performance
2x RTX 6000 Ada 96GB
256GB of ECC RAM
EPYC CPU
vLLM (instead of Ollama)
gpt-oss:120b
That gives you:
Excellent coding performance
Large context support
Strong agent workflows
Reasonable power usage
room for future models
without crossing into hyperscale infrastructure territory.
Inference Engine Recommendations
AgentOS has been tested and confirmed primarily against the Ollama inference server. However, for production and 10+ concurrent users, vLLM is recommended over Ollama because vLLM is much better for concurrency. vLLM also exposes additional customization parameters that Ollama doesn’t provide.
vLLM vs. Ollama
Ollama is easier for local/dev/testing usage:
Ollama advantages:
Simple setup and configuration
Ollama-provided model library
Automatic quantizations
vLLM advantages:
Production-grade serving
Multi-user scale
Enterprise inference
Higher GPU utilization
vLLM includes:
Paged attention
Continuous batching
Tensor parallelism
KV cache optimization
Advanced scheduler
Meaning:
Far higher throughput
Much better concurrent performance
Dramatically lower latency under load
Ollama can struggle under load, especially on:
10+ concurrent users
70B+ models
120B+ models
MoE models
Agent Service Bus (ASB)
The ASB is a lightweight application, so its hardware requirements are genuinely modest.
ASB Architecture Overview
Transaction Anatomy
A single user-facing AI task generates the following ASB HTTP calls:
For billing purposes, 1 transaction = 1 complete tasks/send > agent/result round trip. Internal polls and status checks are infrastructure overhead and are not counted toward subscription transaction limits.
Concurrency Model
The ASB uses a broker pull model: each registered agent processes one task at a time. The maximum number of tasks in flight at any instant equals the number of actively working agents. This makes agent count the primary lever for concurrent user capacity - not the ASB’s HTTP throughput, which comfortable exceeds what any reasonable number of agents generates.
Database WAL Write Capacity by Hardware Tier
Since every transaction produces exactly 2 DB WAL writes, the sustainable write rate sets a hard ceiling on transaction throughput (though in practice this ceiling is far higher than inference can fill).
What the ASB Does Under Load
CPU: Almost none. The ASB runs an event-loop router, not a compute engine. JSON parsing + DB writes are the bottleneck.
Memory: Very low. The in-memory registry + session store for 100 agents is well under 100 MB.
Disk I/O: DB WAL writes on every task/telemetry event. NVMe matters more than capacity.
Network: Lots of small HTTP requests (agent polls every few seconds each). Low bandwidth, high connection count.
Workload Reality Check
Tier 1 - Small Team / Lab (up to ~50 agents)
Option A: Mini PC (Best Value)
Intel NUC 13 Essential or Beelink EQ12 / Mini S12 Pro
Recommended config: Beelink EQ12 with 16 GB RAM + 500 GB NVMe - ~$200 total, no assembly required
Add a UPS (APC Back-UPS 600VA, ~$80) to survive power blips
Total Tier 1A Cost: ~$280-$400
Option B: Raspberry Pi 5 (Ultra-Low Power, Air-Gapped Deployments)
Total Tier 1B cost: ~$140 - good for air-gapped factory floors, retail POS backrooms, or edge deployments where power is constrained.
Tier 2 - Department / Mid-Size Deployment (50-500 agents)
Small Form-Factor Server
Dell OptiPlex 7000 MFF or HP EliteDesk 800 G9 Mini
These are enterprise mini-desktops with server-grade reliability, out-of-band management, and 3-5 year warranties.
Recommended add-ons:
Tier 3 - Enterprise / Data Center Rack (500+ agents, HA, Compliance)
1U Rack Server
Dell PowerEdge R250 / R350 or HPE ProLiant DL20 Gen10+
Honest note: A $2,000 rack server is almost certainly overkill for the ASB alone. The justification is co-location with other services, compliance requirements (ECC RAM, IPMI), or a customer IT policy that mandates rack-mounted servers.
Supporting Infrastructure (if not already present):
Total Tier 3 cost: ~$2,500-$4,500
ASB Hardware Capacity Analysis
The table below summarizes capacity across all hardware tiers. Figures assume the gpt-oss:120b model on appropriately-sized inference hardware, with agents polling at 500 ms and task durations of 10-60 seconds. Monthly transaction figures assume 40% utilization - a realistic business-hours production load.
Note: Max open connections is the number of simultaneous open TCP sockets to the ASB (agents polling + clients waiting). The OS file descriptor limit must be raised to ulimit -n 65535 for Tier 2 and above. This is a configuration step, not a hardware limitation.