Create a Parakeet Managed Inference Job in the web interface
Serve the Parakeet speech-to-text model behind an OpenAI-compatible transcription endpoint from the CosmicAC web interface.
Create a Parakeet Managed Inference Job to serve the speech-to-text model behind an OpenAI-compatible transcription endpoint. The form has six sections, and you click Continue to move from one section to the next. For a description of each field in the form, see Parakeet Managed Inference Job configuration.
Prerequisites
Before you start, make sure that you have the following.
- A running CosmicAC deployment. See Set up CosmicAC.
- Access to the CosmicAC web interface.
Steps
Open the new job form
In the left navigation, click Jobs, and then click New Job.
Select the job type
In the What kind of job? section, select Managed Inference, and then click Continue.
Enter the basics
In the Basics section, enter a Job name. Use lowercase letters and hyphens.
In Tags, add at least one tag. To add a tag, type it, and then press Enter or type a comma.
Click Continue.
Select the model
In the Model to serve section, select nvidia/parakeet-tdt-0.6b-v3. CosmicAC prefills the Parakeet configuration from the model's recommended configuration. The configuration has the following fields.
- Chunk duration: the seconds of audio in each chunk, with a minimum of 10.
- Chunk overlap: the seconds of overlap between adjacent chunks, with a minimum of 5. The overlap must be less than the chunk duration.
- Max file size: the maximum upload size in MB, with a minimum of 1024.
A Parakeet recommended configuration stores no runtime image, data type, quantization, tensor parallel, or reasoning parser, so the form doesn't show those fields. CosmicAC supplies the speech-to-text runtime image from the job type.
Name the endpoint
Under Endpoint, enter an Endpoint name. The name must be unique across the active Managed Inference Jobs in the deployment. A failed job keeps its name until you delete the job.
A Parakeet job serves /v1/audio/transcriptions. See Audio transcriptions.
Set the root disk
Under Instance resources, set the Root disk (GB). Select a preset or enter a value. Increase it for a large checkpoint that exceeds the cluster default.
Require an API key
Under API key required, select Require Authorization header, and then click Continue. To create an API key that authenticates requests to the endpoint, see Create an API key in the web interface.
Select the hardware
In the Hardware section, select a Location first. The GPU list stays empty until you select one.
Select a GPU from the ones available in that location. Each card shows the GPU's VRAM, CPU, and RAM. Set the GPU count, and then set the CUDA / driver. Below the count, Recommended for this model shows the model's GPU count. On one node is the most free GPUs on any single node, and In this region is the total free across the location.
Set Replicas to 1, 2, or 4. CosmicAC doesn't autoscale replicas, and a job keeps the count it starts with. A Parakeet replica always runs on one node. For how requests reach the replicas, see Replicas.
Click Continue.
Select the notification events
In the Notifications section, turn on each job lifecycle event that you want this job to report. CosmicAC turns all four on by default.
- job.failed: the job transitions to Failed, and the event carries the failure reason.
- job.degraded: healthy replicas drop below the count you set, and the endpoint stays live.
- job.recovered: the job returns to Active from Degraded or Failed.
- job.restart_storm: any replica restarts three times within 10 minutes.
These preferences cover this job alone. An event you turn on here reaches your webhook only if it's also turned on in Settings > Notifications, which also controls the model health and usage window events for the whole deployment. See Set up webhook notifications.
Click Continue.
Review and create the job
In the Review & launch section, check that it reports Ready to create, and then click Create job. If it reports issues instead, click Edit on the row that names the problem, fix it, and then return to this section.
Open the endpoint
In the Provisioning status dialog, click Open job. When the job is running, click the Endpoint tab to view the endpoint URL. To send audio to the endpoint, see Transcribe audio.
Help and troubleshooting
Job stuck in Creating or Starting
If a job stays in Creating or Starting, check the status of its KubeVirt virtual machine instance (VMI).
-
Find the job's container ID. Click Jobs in the left navigation, and then click the job. The Containers section lists the Container ID for each container.
-
Find the VMI for the container.
CosmicAC creates one VMI for each container and names it
<container-id>-n0. A multi-node job has one VMI per node.From a machine with
kubectlaccess to your Kubernetes cluster, run the following command.kubectl get vmi -n <namespace>Replace
<namespace>with the namespace configured inK8S_NAMESPACE. -
Check the VMI status.
-
If the VMI is not Running, inspect its events.
kubectl describe vmi <container-id>-n0 -n <namespace> -
If the VMI is Running but the job stays in Creating or Starting, cosmicac-wrk-agent-inference cannot reach cosmicac-wrk-server-k8s-nvidia. These two components connect directly, and some cluster network configurations can block the connection.
To route the connection through a relay, see Set up a relay for CosmicAC.
-