modular.
ConfiguratorEXPLORE / 0.4.1

Modular / Configurator

Any model.
One clear setup.

Choose a model, a serving runtime, and the settings that matter. Leave with a configuration you can inspect and adapt.

01

Models 12

Choose a name. The reference is filled in for you.

Selected reference Qwen/Qwen3-32B
Use another model Paste a link or reference
Source
Architecture
02

Engines 6 + custom

Choose a serving runtime.

Weight format / quantization
03

Configuration up to 15

Set capacity first. Open a group when the workload asks for it.

Context window
1
1 device64 devices

Scheduling

vLLM · 6 controls
Tokens per batchAuto leaves the server default
Sequences per stepAuto leaves the server default
Chunked prefillSplit long prompts across scheduler steps
Eager executionUseful when debugging; may reduce throughput
Precision and cache Two more decisions
Compute dtype
KV cache dtype

FP8 cache needs hardware and checkpoint support; verify quality for your workload.

Advanced settings Network, memory, and flags
Host
Port
90%
50%99%
Optional
Trust model codeAllow model repository code where the server supports it
Prefix cacheReuse computation for shared prompt prefixes