Skip to content

Inference · Python

LLMLoad

Open Source

LLMLoad is a launcher and load-testing utility for current llama.cpp builds. It makes it easy to run multiple models on mixed GPUs — Intel Arc Pro, ROCm, CUDA — with one command, then load-test token throughput against the running server.

github.com/AStarStarship/… →

Specification

LanguagePython + shell
BackendsVulkan (Intel Arc), ROCm, SYCL, CUDA, hybrid
Upstreamllama.cpp (fast-forward-only updates)

Design & implementation

One-command lifecycle

UPDATE, BUILD, and LS are subcommands. Launch uses positional arguments for port, backend, devices, model, and runtime settings. Record the printed revision when comparing runs; upstream llama.cpp changes can still regress.

Mixed-GPU builds

BUILD HYBRID produces a llama.cpp that exposes Arc cards through Vulkan and NVIDIA cards through CUDA in the same binary. Device numbers come from LS — never assume PCI ordering.

Token-throughput load testing

llmload.py hammers the running server with configurable concurrency, request count, warmup, and max-tokens — real throughput numbers, not vibes.

Context-budget testing

Configure the context budget for the model and hardware, then measure memory use and token throughput. Long-context capacity depends on the selected build and available memory.

Related