Overview
This page includes a sample of different sizing needed to meet a number of users/data expected
Please note that the primary parameter to consider is the number of concurrent users. Any GPU-based sizing can support a large number of users; however, if many users submit requests simultaneously, some may experience delays while the GPU processes previous prompts
Single Unit
You’ll need two virtual or physical servers. One Windows server and another Linux server to host the docker containers. The docker containers may be collocated on an existing Linux Docker host.
Component | Small | Medium | Large | Large Plus Storage Optimised | X-Large |
|---|---|---|---|---|---|
Number of Licensed Users Supported | 150 Data Analysis not supported | 600 | 2000 | 2000 | 3000 |
Number of concurrent users | 5 | 20 | 40 | 40 | 100 |
Amount of data supported | 500 GB | 700 GB | 1.4 TB | 30 TB | 30 TB |
CPU | 12 Cores i7 /Xeon | 16 cores i7 /Xeon | 32 cores i7 /Xeon | 64 Cores i7 /Xeon | 64 cores i7 /Xeon |
Memory | 48 GB DDR4+ | 64 GB DDR4+ | 128 GB DDR4+ | 256 GB DDR4+ | 256 GB DDR4+ |
NVMe SSD | 256 GB | 1 TB | 4 TB | 8 TB (Source Files not stored) | 8 TB (Source Files not stored) |
1 x RTX 4090 24GB 1 x RTX 4070 Ti 12 GB | 4 x RTX 4090 24GB 1 x RTX 4070 Ti | 8 x RTX 4090 24GB 2 x RTX 4070 Ti | 10 x RTX 4090 24GB | ||
2 x L4 24 GB | 1 x H100NVL GPU | 2 x H100NVL GPU | 2 x H100NVL GPU | 3 x H100NVL GPU | |
LLM Supported Options | Runs 1x GPT OSS 20B 1 Embedding model | 1x GPT OSS 120B -- 1 Embedding model | 2 x Gemma 27B OR 2 x GPT OSS 120B -- Embedding model supported | 2 x Gemma 27B OR 2 x GPT OSS 120B -- Embedding model supported | 3 x Gemma 27B OR 3 x GPT OSS 120B Embedding model supported |
Enterprise Deployments
Enterprise 1 | Enterprise 2 | |
|---|---|---|
Number of Licensed Users Supported | 6,000 | 10,000 |
Number of concurrent users | 500 | 1,000 |
Amount of data supported | 60 TB | 100TB |
Windows Servers | 2 x Active - Active (may be VMs) Specs per server: 32 GB RAM 16 Cores 200 GB SSD/NVMe | 2 x Active - Active (may be VMs) Specs per server: 64 GB RAM 32 Cores 200 GB SSD/NVMe |
Linux Servers | 2x Active - Active (may be on one host with multiple instances of containers or separate hosts) Total resources required (may be split between 2 hosts): 512GB RAM 128 Cores 16 TB SSD/NVMe 4 x H100 / H200 | 2x Active - Active (may be on one host with multiple instances of containers or separate hosts) Total resources required (may be split between 2 hosts): 768 GB – 1 TB DDR4+ RAM 192 Cores 24 TB SSD/NVMe 6 x H100 / H200 |
LLM GPU Compatibility Table
Model | Minimum VRAM (Single GPU) | Recommended Single-GPU Options | Multi-GPU (Tensor Parallel) Options | Notes / Requirements |
|---|---|---|---|---|
OSS 20B | 16–24 GB | RTX 3090 / 4090 (24 GB), RTX A5000 (24 GB), A10 (24 GB), A40 (48 GB), A100 (40/80 GB), H100 |
| GPU must support CUDA 12.x, Tensor Cores, and Compute Capability ≥ 7.0. NVLink recommended for TP. |
OSS 120B | 80 GB+ | A100 80 GB, H100 80 GB, H200, MI300X |
| Designed for 80 GB-class GPUs. PCIe-only setups not supported. Requires CUDA 12.x & CC ≥ 7.0. |
Gemma 3 27B | 48 GB | RTX 6000 Ada (48 GB), A40 (48 GB), A100 80 GB, H100 |
| Gemma 27B requires ~54 GB in FP16; ≥48 GB VRAM recommended for full-precision. Requires CUDA 12.x & Compute Capability ≥ 7.0. |
General GPU Requirements
Compute Capability: ≥ 7.0 (Volta or newer)
Tensor Cores: Required (Volta/Turing/Ampere/Ada/Hopper)
CUDA Version: CUDA 12.x required for vLLM / SGLang / Triton / FlashAttention
VRAM Pooling: VRAM does not combine across GPUs; multi-GPU requires TP
Interconnect: NVLink strongly recommended for multi-GPU TP on 20B+ models
Not Supported: Maxwell (M10/M60/M40), Pascal (P100/P40/GTX 10xx), GPUs with <12 GB VRAM
High Availability
The following components need to be considered when designing a highly available solution
Dashboard (user interface) + BG Service - Two Windows servers required. May be load balanced with Active - Active or Active - Passive configurations.
Gateway - Requires two Linux hosts. Paired with one Dashboard/ BG Service each.
LLMs. At Least 2 embedding GPUs and 2 LLM GPUs. Active - Active configuration recommended to maximise routine GPU utilisation for speed.
Databases - Postgres and MS SQL databases will need to be configured to run HA replicas.