CosmicAC Logo
Quick start

Create a vLLM Managed Inference Job with the CLI

Create a vLLM Managed Inference Job with the CLI and verify that it's running.

Create a Managed Inference Job with the CosmicAC CLI to serve a language model with vLLM behind an OpenAI-compatible chat endpoint.

Before you create the job, make sure that the model has a recommended configuration. CosmicAC uses this configuration to prefill the serving values that the model needs to deploy.

Prerequisites

Before you start, make sure that you have the following.

Steps

Create the job

CosmicAC provides recommended configurations for a set of models. For the list, see Recommended configuration values. To serve another model that vLLM supports, register a vLLM model or add a recommended configuration before you create the job.

The model needs some of its recommended serving values to deploy. If you change those values, the job might not start.

Create the job interactively by answering prompts, or pass the job configuration as flags.

Start the interactive job setup.

cosmicac jobs create

Select Managed Inference (vLLM) as the job type, and then answer the prompts.

Configure the following fields.

  • Job name: a name that identifies the job.
  • Tags: comma-separated labels for the job.
  • Location: the region where the job runs.
  • GPU type: the GPU to use. The CLI lists each GPU type with the number of free GPUs on one node and in the selected location.
  • GPU count: the number of GPUs for one replica. One of 1, 2, 4, 8, or 16. See GPU configuration.
  • Model: the model to serve. The CLI lists the models that have a recommended configuration.
  • Runtime image: the vLLM serving image. The CLI uses the runtime image from the model's recommended configuration.
  • Data type: the numeric precision the model runs at.
  • Quantisation: the method that compresses the model weights.
  • Tensor parallel: the number of GPUs to split the model across.
  • GPU memory utilization: the fraction of GPU memory to use. Interactive mode shows the options as percentages.
  • Max model length: the maximum context length.
  • Max concurrent sequences: the maximum number of requests handled at the same time.
  • Reasoning parser: the parser that separates thinking tokens from the final response. If the model needs no parser, select Default.
  • Video & image input: whether the model accepts multimodal input.
  • Endpoint name: the name used in the endpoint URL. Use lowercase letters, numbers, and hyphens only. Interactive mode checks that the name is available and shows the endpoint URL.
  • Replicas: the number of model copies to run. CosmicAC doesn't autoscale replicas.
  • Require Authorization header: whether callers must send an API key. See Create an API key.
  • Root disk size: the root disk size in GB. The minimum is the Root disk (GB) value in the model's recommended configuration.
  • Environment variables: optional variables for the inference agent, including vllm serve options. Interactive mode asks whether to add any, and then prompts for each name and value. See vLLM serving options.
  • Notifications: the job lifecycle events to report. All four events are on by default. Interactive mode lets you select them with checkboxes.

An event reaches your webhook only if it's also turned on in Settings > Notifications. See What controls delivery.

In interactive mode, the CLI shows a job summary and asks you to confirm the job. If a team is active, the prompt shows the team ID. Enter yes to create the job. The CLI then shows the job ID.

To serve a speech-to-text model instead, see Create a Parakeet Managed Inference Job with the CLI.

For every field and its CLI flag, see vLLM Managed Inference Job configuration.

Verify the job

List your jobs.

cosmicac jobs list

Check that the new job appears in the list with its ID, name, tags, and status. Wait for the job to reach running. The endpoint accepts requests after the job reaches this status.

Help and troubleshooting

Job stuck in Creating or Starting

If a job stays in Creating or Starting, check the status of its KubeVirt virtual machine instance (VMI).

  1. Find the job's container ID.

    cosmicac jobs detail <jobId>

    The output lists the Container ID for each container.

  2. Find the VMI for the container.

    CosmicAC creates one VMI for each container and names it <container-id>-n0. A multi-node job has one VMI per node.

    From a machine with kubectl access to your Kubernetes cluster, run the following command.

    kubectl get vmi -n <namespace>

    Replace <namespace> with the namespace configured in K8S_NAMESPACE.

  3. Check the VMI status.

Next steps

On this page