Login

New user? Create an account
Forgot your password? Get it back!

Sign up

Already a member? Login

2FA

2-Factor Authentication is enabled for this account
Login with a different user

Email Confirmation

Send confirmation email to the following
Login with a different user
Help Center / Self-Hosting

Local Model Self-Hosting and Recommendations

Recommendations on hardware and infrastructure for self-hosting models for AgentOS and the Agent Service Bus (ASB).

Table of Contents

Model: gpt-oss:120b 2

VRAM requirements 2

Recommended “Good Production” Build 2

Why this configuration works well 3

VRAM headroom 3

Expected performance 3

Lower-cost option for gpt-oss:120b (Budget Build) 3

Model: Kimi K2 4

Recommended Production Hardware for Full Kimi K2 (Serious Enterprise Build) 4

Why this is required 5

More realistic recommendation for Kimi 5

Best Practical Option 5

Recommendation by Budget 6

Best balance of cost/performance 6

Inference Engine Recommendations 7

vLLM vs. Ollama 7

Agent Service Bus (ASB) 8

ASB Architecture Overview 8

Transaction Anatomy 8

Concurrency Model 8

Database WAL Write Capacity by Hardware Tier 9

What the ASB Does Under Load 9

Workload Reality Check 9

Tier 1 - Small Team / Lab (up to ~50 agents) 9

Option A: Mini PC (Best Value) 9

Option B: Raspberry Pi 5 (Ultra-Low Power, Air-Gapped Deployments) 10

Tier 2 - Department / Mid-Size Deployment (50-500 agents) 11

Tier 3 - Enterprise / Data Center Rack (500+ agents, HA, Compliance) 12

ASB Hardware Capacity Analysis 14

Pricing Summary Table 14



Note: The following information assumes hosting with 10 or more concurrent users. For 10+ concurrent users, the biggest factor isn’t just fitting the model into VRAM - it’s maintaining enough throughput and KV-cache headroom so latency doesn’t become painful once several people are generating at the same time.


This document primarily focuses on the gpt-oss:120b model as it has been tested and confirmed as working sufficiently with AgentoS. Kimi K2 has also been included since it has been mentioned as being evaluated by some customers. Additional models and server-sizing recommendations will be added as they get tested by Lucus Labs for use with AgentOS.





Model: gpt-oss:120b

VRAM requirements


  • Bare minimum usable VRAM: ~60-80 GB

  • Comfortable production VRAM: ~96-160 GB

  • 10 active users: multi-GPU strongly recommended


Community deployments commonly recommend 4x RTX 3090s or enterprise GPUs for this class of workload.


Recommended “Good Production” Build


Component

Recommendation

CPU

AMD EPYC 9374F or EPYC 9354

RAM

256 GB DDR5 ECC

GPUs

2x NVIDIA RTX 6000 Ada Generation (96 GB each)

Total VRAM

192 GB

Storage

2x 4 TB NVMe Gen4 SSD (RAID1 optional)

PSU

2400W redundant

Networking

10 GbE

OS

Ubuntu Server 24.04


Why this configuration works well

VRAM headroom

gpt-oss:120b at Q4 needs roughly:


  • 60-73 GB for weights

  • Plus KV cache

  • Plus batching/concurrency overhead


With 192 GB total VRAM you can comfortably support:


  • 10 users

  • Larger context windows (including gpt-oss:120b’s 132k token ctx)

  • Batching

  • Agent workflows

  • RAG pipelines


Expected performance

Approximate:


  • 80-180 tokens/sec aggregate

  • ~10-25 tokens/sec per active user


Depends heavily on:


  • Context size

  • Batching

  • Quantization

  • Prompt length


Lower-cost option for gpt-oss:120b (Budget Build)


Component

Recommendation

CPU

AMD Threadripper Pro

RAM

192 GB ECC

GPUs

4x NVIDIA GeForce RTX 3090

Total VRAM

96 GB

Storage

2 TB NVMe

PSU

2000W


This is the most common “prosumer lab” approach and acceptable for starting out and testing ideas & applications. Many users mention 4x 3090 setups as the “economical route” for 120B-class models.


Tradeoffs:


  • More power draw (~1500W GPU-only)

  • More heat/noise

  • PCIe/riser complexity

  • Lower reliability than enterprise GPUs





Model: Kimi K2

The full Moonshot AI Kimi K2 model is described as:


  • ~1 trillion parameter MoE model

  • ~32B active parameters

  • ~500-620+ GB VRAM requirement (even when quantized)


This is in a completely different infrastructure tier than the gpt-oss:120b model.


Recommended Production Hardware for Full Kimi K2 (Serious Enterprise Build)


Component

Recommendation

CPU

Dual AMD EPYC 9755

RAM

1 - 2 TB DDR5 ECC

GPUs

8x NVIDIA H200 SXM

Total VRAM

~1.1 TB

Interconnect

NVLink / NVSwitch

Storage

8 TB+ Gen5 NVMe

Networking

100 GbE

Chassis

DGX-class server


Why this is required

Kimi K2 Q4 quantization alone is estimated around:


  • 500-620 GB VRAM

  • Before large KV cache overhead


For 10 concurrent users:


  • You ned huge KV cache memory

  • Fast interconnects

  • Aggressive batching

  • Tensor parallelism


this is effectively:


  • A “mini AI datacenter”

  • Likely a $180K-$350K server depending on GPU pricing


More realistic recommendation for Kimi

If your goals is:


  • Internal coding assistant

  • Enterprise chatbot

  • AI agents

  • RAG

  • Code review

  • Automation


… then the full Kimi K2 is usually not economically sensible to self-host.


Instead:


Best Practical Option

Run:


  • Kimi K2 Distill 32B

  • or gpt-oss:120b


on:


  • 2x RTX 6000 Ada

  • or 4x 3090


You’ll get:


  • Dramatically lower latency

  • Massively lower power costs

  • Simple infrastructure

  • Far easier scaling

  • Much better reliability


The distilled Kimi variants are specifically intended for practical local deployment.





Recommendation by Budget


Budget

Recommendation

<$10k

4x RTX 3090 + gpt-oss:120b

$15k - 30k

2x RTX 6000 Ada + gpt-oss:120b

$40k - 80k

4x H100/H200 for serious enterprise serving

$200k

Full Kimi K2 cluster


Best balance of cost/performance


  • 2x RTX 6000 Ada 96GB

  • 256GB of ECC RAM

  • EPYC CPU

  • vLLM (instead of Ollama)

  • gpt-oss:120b


That gives you:


  • Excellent coding performance

  • Large context support

  • Strong agent workflows

  • Reasonable power usage

  • room for future models


without crossing into hyperscale infrastructure territory.





Inference Engine Recommendations

AgentOS has been tested and confirmed primarily against the Ollama inference server. However, for production and 10+ concurrent users, vLLM is recommended over Ollama because vLLM is much better for concurrency. vLLM also exposes additional customization parameters that Ollama doesn’t provide.


vLLM vs. Ollama

Ollama is easier for local/dev/testing usage:


Ollama advantages:


  • Simple setup and configuration

  • Ollama-provided model library

  • Automatic quantizations


vLLM advantages:


  • Production-grade serving

  • Multi-user scale

  • Enterprise inference

  • Higher GPU utilization


vLLM includes:


  • Paged attention

  • Continuous batching

  • Tensor parallelism

  • KV cache optimization

  • Advanced scheduler


Meaning:


  • Far higher throughput

  • Much better concurrent performance

  • Dramatically lower latency under load


Ollama can struggle under load, especially on:


  • 10+ concurrent users

  • 70B+ models

  • 120B+ models

  • MoE models




Agent Service Bus (ASB)

The ASB is a lightweight application, so its hardware requirements are genuinely modest.


ASB Architecture Overview

Transaction Anatomy

A single user-facing AI task generates the following ASB HTTP calls:


Call

Description

task/send (x1)

Client submits the task. Writes: 1 DB WAL row (task creation).

agent/poll (x2 - 8)

Agent pulls pending work. Poll interval 500 ms; no DB write on empty queue; immediate re-poll after completing work.

agent/result (x1)

Agent posts result. Writes: 1 DB WAL row (task state update)

tasks/get (x0 - 3)

Client optionally checks status. Read-only, no DB write.

Total HTTP calls

~5 - 13 per user transaction

DB writes

Exactly 2 per transaction (tasks/send + agent/result)


For billing purposes, 1 transaction = 1 complete tasks/send > agent/result round trip. Internal polls and status checks are infrastructure overhead and are not counted toward subscription transaction limits.


Concurrency Model

The ASB uses a broker pull model: each registered agent processes one task at a time. The maximum number of tasks in flight at any instant equals the number of actively working agents. This makes agent count the primary lever for concurrent user capacity - not the ASB’s HTTP throughput, which comfortable exceeds what any reasonable number of agents generates.


Primary Capacity Constraint

Concurrent users in flight = number of registered, healthy agents actively working

Total registered users the system can serve = concurrent users in flight ÷ task utilization rate

Example: 50 agents, 30-second avg response, user submits 1 task/minute = 50 ÷ (30s/60s) = 100 total users can be served without queueing


Database WAL Write Capacity by Hardware Tier

Since every transaction produces exactly 2 DB WAL writes, the sustainable write rate sets a hard ceiling on transaction throughput (though in practice this ceiling is far higher than inference can fill).


Hardware

Estimated DB WAL Writes/sec

RPi 5 + USB 3.0 SSD

3,000 - 8,000

Mini PC M.2 NVMe (N100)

20,000 - 50,000

SFF Server 12th-gen + NVMe

40,000 - 100,000

1U Rack + NVMe U.2 RAID

80,000 - 200,000


What the ASB Does Under Load


  • CPU: Almost none. The ASB runs an event-loop router, not a compute engine. JSON parsing + DB writes are the bottleneck.

  • Memory: Very low. The in-memory registry + session store for 100 agents is well under 100 MB.

  • Disk I/O: DB WAL writes on every task/telemetry event. NVMe matters more than capacity.

  • Network: Lots of small HTTP requests (agent polls every few seconds each). Low bandwidth, high connection count.


Workload Reality Check


Resource

What the ASB Uses

CPU

< 5% sustained on a modern core

RAM

256 MB typical, 1 GB comfortable ceiling

Disk I/O

DB WAL: small sequential writes, NVMe preferred but not required

Network

~1-5 KB per agent poll request; 100 agents = negligible bandwidth

Power

The binary itself draws near zero - the hardware around it dominates


Tier 1 - Small Team / Lab (up to ~50 agents)

Option A: Mini PC (Best Value)

Intel NUC 13 Essential or Beelink EQ12 / Mini S12 Pro


Spec

Value

CPU

Intel N100 / Core i3-1315U (4-6 cores)

RAM

8 - 16 GB DDR4/DDR5

Storage

256 GB NVMe M.2 SSD

Network

2.5 GbE onboard

Power draw

10 - 25W idle

Form factor

~4” x 4” x 2”

Max registered agents

50

Concurrent users

25 - 50

Max open connections

500 - 1,500

ASB HTTP throughput

800 - 1,500 req/sec

Transactions / month

~500k - 2m at 40% utilization

Best for

Small clinic, pilot deployment, departmental proof-of-concept


Product

Price (approx.)

Beelink EQ12 (N100, 16 GB, 500 GB)

$180 - $220

Beelink SER5 Max (Ryzen 5 5560U, 16 GB, 500 GB)

$280 - $320

Intel NUC 13 Essential (i3-1315U, barebones)

$200 + RAM/SSD ~$80


Recommended config: Beelink EQ12 with 16 GB RAM + 500 GB NVMe - ~$200 total, no assembly required


Add a UPS (APC Back-UPS 600VA, ~$80) to survive power blips


Total Tier 1A Cost: ~$280-$400


Option B: Raspberry Pi 5 (Ultra-Low Power, Air-Gapped Deployments)


Spec

Value

CPU

Arm Cortex-A76 quad-core 2.4 GHz

RAM

8 GB LPDDR4X

Storage

Samsung 256 GB microSD or USB 3 SSD

Network

Gigabit Ethernet

Power draw

5-12W

Form factor

Credit card-size

Max registered agents

20 - 30 (practical limit before USB SSD I/O becomes tight)

Concurrent users

10 - 25

Max open connections

200 - 500 (raise OD fd limit to 1024+)

ASB HTTP throughput

300 - 600 req/sec

Transactions / month

~200k - 800k at 40% utilization

Best for

Air-gapped factory floors, edge deployments, proof-of-concept with cost constraints



Item

Price

Raspberry Pi 5 8 GB

$80

Official case + active cooler

$15

256 GB Samsung USB-A SSD (for ASB data persistence)

$30

CanaKit power supply

$12


Total Tier 1B cost: ~$140 - good for air-gapped factory floors, retail POS backrooms, or edge deployments where power is constrained.


Tier 2 - Department / Mid-Size Deployment (50-500 agents)

Small Form-Factor Server


Dell OptiPlex 7000 MFF or HP EliteDesk 800 G9 Mini


These are enterprise mini-desktops with server-grade reliability, out-of-band management, and 3-5 year warranties.


Spec

Value

CPU

Core i5-12500T / i7-12700T

RAM

16 - 32 GB DDR5

Storage

512 GB NVMe (os) + 1 TB NVMe (data volume for ASB)

Network

Intel I219 GbE + optional 2.5 GbE add-in

Power draw

35-65W under load

Form factor

1.4L - fits in a network closet

Max registered agents

500

Concurrent users

150 - 500

Max open connections

2,000 - 10,000 (ulimit -n 65,535 recommended)

ASB HTTP throughput

2,000 - 5,000 req/sec

Transactions / month

~3m - 15m at 40% utilization

Best for

Department deployments, community hospitals, small call centers (25 - 100 seats)


Configuration

Price

Dell OptiPlex 7010 MFF (i5-12500, 16 GB, 512 GB NVMe)

$650 - $800 new

HP EliteDesk 800 G9 Mini (i7-12700T, 16 GB, 512 GB)

$750 - $900 new

Refurbished Gen 8/9 (i5, 16 GB, SSD) - Dell Renewed

$300 - $450


Recommended add-ons:


Item

Price

APC Back-UPS Pro 1500VA (covers server + switch)

$200

Unmanaged 8-port GbE switch (if needed)

$30 - $60

External USB drive for DB backups (2 TB)

$60


Tier 3 - Enterprise / Data Center Rack (500+ agents, HA, Compliance)

1U Rack Server


Dell PowerEdge R250 / R350 or HPE ProLiant DL20 Gen10+


Spec

Value

CPU

Xeon E-2300 series (4-8 cores)

RAM

32-64 GB ECC DDR4

Storage

2x 480 GB SSD RAID-1 (OS) + 2x 1.92 TB NVMe U.2 (data)

Network

Dual-port Intel GbE + optional 10 GbE

iDRA / iLO

Out-of-band management, remote console

Power draw

Redundant 450W PSUs

Form factor

1U rack

Max registered agents

1,000+ (ASB has no hard limit; OS and network are the ceiling)

Concurrent users

500 - 2000

Max open connections

10,000 - 50,000 (ulimit -n 65,535 required)

ASB HTTP throughput

5,000 - 15,000 req/sec

Transactions / month

~10m - 100m at 40% utilization

Best for

Large hospitals, multi-site health networks, enterprise call centers (500+ seats), HA/compliance deployments



Configuration

Price

Dell PowerEdge R250 (Xeon E-2314, 16 GB, 2x 480 GB SATA SSD)

$1,800 - $2,400

HPE ProLiant DL20 Gen10+ (Xeon E-2314, 16 GB, 1x 960 GB SSD)

$2,000 - $2,800

Dell R350 (Xeon E-2336, 32 GB, RAID controller)

$2,800 - $3,500


Honest note: A $2,000 rack server is almost certainly overkill for the ASB alone. The justification is co-location with other services, compliance requirements (ECC RAM, IPMI), or a customer IT policy that mandates rack-mounted servers.


Supporting Infrastructure (if not already present):


Item

Price

APC Smart-UPS 1500VA rack-mount

$500 - $700

12U wall-mount rack enclosure

$200 - $400

Cat6 patch cables + keystone panel

$50 - $100

Managed 8-port GbE PoE switch (Ubiquiti US-8-60W)

$110


Total Tier 3 cost: ~$2,500-$4,500


ASB Hardware Capacity Analysis

The table below summarizes capacity across all hardware tiers. Figures assume the gpt-oss:120b model on appropriately-sized inference hardware, with agents polling at 500 ms and task durations of 10-60 seconds. Monthly transaction figures assume 40% utilization - a realistic business-hours production load.


Tier

Hardware

Max Agents

Concurrent Users

Max Open Connections

ASB HTTP Req/sec

Transactions / Month

1B - Pi 5

Raspberry Pi 5 + USB SSD

20 - 30

10 - 25

200 - 500

300 - 600

200k - 800k

1A - Mini PC

Beelink EQ12 + NVMe

50

25 - 50

500 - 1,500

800 - 1,500

500k - 2m

2 - SFF Server

Dell OptiPlex / HP EliteDesk

500

150 - 500

2,000 - 10,000

2,000 - 5,000

3m - 15m

2R - Refurb

Gen 8/9 OptiPlex / EliteDesk

300 - 400

100 - 350

1,500 - 7,500

1,500 - 3,500

2m - 10m

3 - Rack

Dell R250 / HPE DL20 Gen10+

1000+

500 - 2,000

10,000 - 50,000

5,000 - 1,5000

10m - 100m


Note: Max open connections is the number of simultaneous open TCP sockets to the ASB (agents polling + clients waiting). The OS file descriptor limit must be raised to ulimit -n 65535 for Tier 2 and above. This is a configuration step, not a hardware limitation.


Pricing Summary Table


Tier

Hardware

Agents

Power

Total Cost

1A - Mini PC

Beelink EQ12 + UPS

1 - 50

15W

~$280 - $400

1B - Pi 5

Raspberry Pi 5 8GB + SSD

1 - 50

8W

~$140

2 - SFF Server

Dell OptiPlex / HP EliteDesk + UPS

50 - 500

50W

~$800 - $1,200

2R - Refurb

Gen 8 OptiPlex/EliteDesk refurb + UPS

50 - 500

50W

~$500 - $750

3 - Rack 1U

Dell R250 / HPE DL20

500+

120W

~$2,500 - $4,500