Modular / Configurator
Any model.
One clear setup.
Choose a model, a serving runtime, and the settings that matter. Leave with a configuration you can inspect and adapt.
01
Models 12
Choose a name. The reference is filled in for you.
Selected reference
Qwen/Qwen3-32BUse another model Paste a link or reference
Source
Architecture
02
Engines 6 + custom
Choose a serving runtime.
Weight format / quantization
Variables: {{model}}, {{port}}, {{host}}, {{context}}, {{quant}}. Quote values in your template as needed.
03
Configuration up to 15
Set capacity first. Open a group when the workload asks for it.
Context window
512262K
1 device64 devices
Scheduling
vLLM · 6 controlsTokens per batchAuto leaves the server default
Sequences per stepAuto leaves the server default
Chunked prefillSplit long prompts across scheduler steps
Eager executionUseful when debugging; may reduce throughput
Precision and cache Two more decisions
Compute dtype
KV cache dtype
FP8 cache needs hardware and checkpoint support; verify quality for your workload.
Advanced settings Network, memory, and flags
Host
Port
102465535
50%99%
Optional
Trust model codeAllow model repository code where the server supports it
Prefix cacheReuse computation for shared prompt prefixes