Herd Infrastructure × Risika
Checking API…

Your dedicated GPU machine

It works like a GPU instance on EC2. You SSH into your own container and run what you want, and your API is served at this address over HTTPS.

What you have

GPUs
2× RTX PRO 4000 Blackwell
GPU memory
2 × 24 GB
CPU / RAM
56 threads · 52 GB
Persistent storage
~1.5 TB in /workspace
Preinstalled
CUDA · PyTorch · vLLM 0.24
Ready to run
Mistral Small 3.2 (24B)

The machine is reserved for Risika only. It is not shared with other customers.

Get started

  1. Log in

    Once we have added your SSH public key:

    ssh -p 2601 root@193.181.215.114

    Host key fingerprint (ED25519): SHA256:o3qDRo9yFJoirUhFjN7p/LUEXYVTnm11qK2G1Vu2tfA. To use just ssh herd, add this to ~/.ssh/config:

    Host herd
      HostName 193.181.215.114
      Port 2601
      User root
  2. Start a model

    The Mistral Small 3.2 weights are already downloaded. To serve it in FP8 across both GPUs:

    /workspace/examples/mistral-small.sh

    The first start takes about 3 minutes while vLLM compiles. When it's ready, the status pill at the top of this page turns green. The very first request after a start is slow (up to a minute). After that, expect around 40 tokens/s. Run it inside tmux to keep it running after you log out, or see "Keep it running" below.

  3. Call your API from Risika

    Every request needs the API key we sent you separately, as a bearer token. The endpoint is OpenAI-compatible:

    import os
    from openai import OpenAI
    
    client = OpenAI(
        base_url="https://risika.herdinfrastructure.com/v1",
        api_key=os.environ["HERD_API_KEY"],
    )
    r = client.chat.completions.create(
        model="mistralai/Mistral-Small-3.2-24B-Instruct-2506",
        messages=[{"role": "user", "content": "Hello!"}],
    )
    print(r.choices[0].message.content)

    Anything you serve on port 8000 in your container is available here, not only vLLM. Authentication is handled by our gateway, so don't add --api-key to vLLM.

Keep it running

If /workspace/start.sh exists and is executable, it runs every time the container starts, so your API comes back on its own after a restart. Its output goes to /workspace/start.log.

cp /workspace/examples/mistral-small.sh /workspace/start.sh

What persists

LocationSurvives a restart or update?
/workspace: code, models, venvs, dataYes
pip install / apt install outside /workspaceNo, it's reset when we update the container

Install your own packages into a venv on /workspace, so they survive updates:

python3 -m venv --system-site-packages /workspace/venv
source /workspace/venv/bin/activate

Good to know

Which models fit?

Each card has 24 GB. A 24B model in bf16 (≈48 GB) doesn't fit, even across both cards, once vLLM needs room for its cache. Use FP8 across both cards (--quantization fp8 --tensor-parallel-size 2), or a 4-bit AWQ/GPTQ build on one card.

Running on both GPUs

For now, add --disable-custom-all-reduce when you use --tensor-parallel-size 2. The example script already does this. NCCL_P2P_DISABLE=1 is set for you. We'll let you know when a host update removes the need for both.

Gated Hugging Face models

export HF_TOKEN=… before downloading. Downloads go to /workspace/.cache/huggingface.

Network

The container has outbound internet access, so pip, Hugging Face and your own services work. Internal and private networks are blocked by design. Traffic to this endpoint is encrypted with TLS.

Your own Docker image

If you'd rather deploy your own image with self-service redeploy, tell us which registry it's in (ECR, GHCR or Docker Hub) and we'll set it up.