GPU containers
CUDA and local models run in containers, not on the host. On the NVIDIA image a container can request the GPU by name.
Give a container the GPU
If this prints your card, it works:
podman run --rm --device nvidia.com/gpu=all registry.fedoraproject.org/fedora nvidia-smi
For example, ollama:
podman run -d --device nvidia.com/gpu=all -p 11434:11434 \
-v ollama:/root/.ollama docker.io/ollama/ollama
The image provides the driver and libcuda; the container brings the CUDA runtime.
As a quadlet:
[Container]
Image=docker.io/ollama/ollama
AddDevice=nvidia.com/gpu=all
PublishPort=127.0.0.1:11434:11434
Volume=ollama:/root/.ollama
NVIDIA image only. On the standard image, pass /dev/dri for Vulkan.
What makes it work
Podman uses CDI specs, which map nvidia.com/gpu=all to device nodes and driver
libraries. Two units maintain it:
-
nvidia-cdi-refresh.servicewrites the spec for the running driver to/var/run/cdi/nvidia.yamleach boot (cleared at shutdown, so never stale). -
pulsar-gpu-containers.serviceenables the SELinux booleancontainer_use_xserver_deviceseach boot. Without it,nvidia-smifails with “Failed to initialize NVML: Insufficient Permissions”. It only affects containers given a GPU with--device.
Don’t write your own spec to /etc/cdi
nvidia-ctk cdi generate in /etc/cdi survives driver
updates and goes stale. The boot unit already writes a fresh one.
When it does not
$ pulsar doctor gpu-containers
ok gpu-containers CDI spec matches driver 615.71.09
podman run --rm --device nvidia.com/gpu=all IMAGE
| doctor says | Do this |
|---|---|
| no CDI spec, so containers cannot be given the GPU | sudo systemctl restart nvidia-cdi-refresh.service, then read journalctl -b -u nvidia-cdi-refresh.service |
| stale CDI spec for driver … | Delete the spec it names. It is one somebody wrote by hand into /etc/cdi, and it went stale at the last driver update. |
| CDI spec matches …, but SELinux blocks containers from the GPU | systemctl status pulsar-gpu-containers.service: the unit that turns the boolean on did not run. |
No NVIDIA module loaded at all: see the nvidia check in
Troubleshooting.
A model for your agents
pulsar agent model on runs llama.cpp’s server as a rootless quadlet (CUDA, Vulkan
or CPU) on 127.0.0.1 behind a key, for opencode and aider.
Coding agents.
Written with help from AI and reviewed by a person before publishing. Spotted a mistake? Let us know.