It works like a GPU instance on EC2. You SSH into your own container and run what you want, and your API is served at this address over HTTPS.
The machine is reserved for Risika only. It is not shared with other customers.
Once we have added your SSH public key:
ssh -p 2601 root@193.181.215.114
Host key fingerprint (ED25519): SHA256:o3qDRo9yFJoirUhFjN7p/LUEXYVTnm11qK2G1Vu2tfA. To use just ssh herd, add this to ~/.ssh/config:
Host herd
HostName 193.181.215.114
Port 2601
User root
The Mistral Small 3.2 weights are already downloaded. To serve it in FP8 across both GPUs:
/workspace/examples/mistral-small.sh
The first start takes about 3 minutes while vLLM compiles. When it's ready, the status pill at the top of this page turns green. The very first request after a start is slow (up to a minute). After that, expect around 40 tokens/s. Run it inside tmux to keep it running after you log out, or see "Keep it running" below.
Every request needs the API key we sent you separately, as a bearer token. The endpoint is OpenAI-compatible:
import os
from openai import OpenAI
client = OpenAI(
base_url="https://risika.herdinfrastructure.com/v1",
api_key=os.environ["HERD_API_KEY"],
)
r = client.chat.completions.create(
model="mistralai/Mistral-Small-3.2-24B-Instruct-2506",
messages=[{"role": "user", "content": "Hello!"}],
)
print(r.choices[0].message.content)
curl https://risika.herdinfrastructure.com/v1/chat/completions \
-H "Authorization: Bearer $HERD_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "mistralai/Mistral-Small-3.2-24B-Instruct-2506",
"messages": [{"role": "user", "content": "Hello!"}]}'
Anything you serve on port 8000 in your container is available here, not only vLLM. Authentication is handled by our gateway, so don't add --api-key to vLLM.
If /workspace/start.sh exists and is executable, it runs every time the container starts, so your API comes back on its own after a restart. Its output goes to /workspace/start.log.
cp /workspace/examples/mistral-small.sh /workspace/start.sh
| Location | Survives a restart or update? |
|---|---|
/workspace: code, models, venvs, data | Yes |
pip install / apt install outside /workspace | No, it's reset when we update the container |
Install your own packages into a venv on /workspace, so they survive updates:
python3 -m venv --system-site-packages /workspace/venv
source /workspace/venv/bin/activate
Each card has 24 GB. A 24B model in bf16 (≈48 GB) doesn't fit, even across both cards, once vLLM needs room for its cache. Use FP8 across both cards (--quantization fp8 --tensor-parallel-size 2), or a 4-bit AWQ/GPTQ build on one card.
For now, add --disable-custom-all-reduce when you use --tensor-parallel-size 2. The example script already does this. NCCL_P2P_DISABLE=1 is set for you. We'll let you know when a host update removes the need for both.
export HF_TOKEN=… before downloading. Downloads go to /workspace/.cache/huggingface.
The container has outbound internet access, so pip, Hugging Face and your own services work. Internal and private networks are blocked by design. Traffic to this endpoint is encrypted with TLS.
If you'd rather deploy your own image with self-service redeploy, tell us which registry it's in (ECR, GHCR or Docker Hub) and we'll set it up.