LLMs on a GP102/Pascal 3-Way SLI Rig

Hassle-free local inference through Vulkan.
Nvidia keeps nudging Pascal further out the door, but there’s an easier way to keep these cards doing useful work than fighting the CUDA toolchain on a rolling distro.

Three Cards, One Dead Standard

The rig is a leftover from when multi-GPU gaming was still a thing people did: three GP102-based GeForce cards on an X299 board, bridged for 3-way SLI, fed by a 1600W supply that has spent most of the last few years being wildly overqualified for the job.

Three GeForce GTX cards in a SLI setup

Three-way GP102 setup.

SLI itself is thoroughly dead — no driver profile is coming to make three cards render one game any more. But that has nothing to do with whether the chips can still compute. Each GP102 is a perfectly capable device with a useful amount of VRAM attached, and three of them idling next to an oversized PSU is a waste when running models locally has become the normal way to try anything out. So: ollama, and how hard can it be.

Nvidia Shows Pascal the Door

Harder than it should be, because Pascal is now actively being removed from the stack — twice over, at two different layers.

The first wall is the driver. Nvidia’s 590 branch drops support for Pascal and everything older, as part of the same move that makes the open-source kernel modules the default. Pascal predates the GSP firmware those modules rely on, so it cannot follow. On Arch this is not a gentle deprecation: the nvidia, nvidia-dkms and nvidia-lts packages roll forward to 590, and if you let that happen on a Pascal box you get a broken graphical environment. The prescribed fix is to uninstall those and install nvidia-580xx-dkms from the AUR instead.

The good news is that 580 is an LTS branch, and Nvidia has committed to supporting pre-Turing hardware on it into mid-2028. So the card keeps working — you just pin it and stop upgrading that one thing.

CUDA 13 Can’t Build for It Either

The second wall is the toolkit, and it is the one that actually bites. CUDA 13.0 removed offline compilation for everything below compute capability 7.5 — which sweeps up Maxwell, Pascal and Volta together. GP102 is sm_61. nvcc from CUDA 13 will not emit device code for it at all; there is no flag to talk it round, because the target support is gone rather than merely deprecated.

Worth being precise about what this does and does not break. It is a build-time wall, not a runtime one: binaries compiled with CUDA 12.9 or earlier keep running fine on any hardware the installed driver supports. Nothing already working stops working. But the moment you want to compile something yourself — which is exactly what you are doing if you want an ollama build that targets your cards — the last toolkit that can do it is 12.9.

The Toolchain Rabbit Hole

Pinning an old toolkit sounds simple. On a rolling-release distro it starts a chain.

CUDA 12.9 supports GCC up to 14.x and no further. Arch has long since moved past that, so the system compiler is rejected outright — not with a subtle miscompile, but with a hard version check baked into CUDA’s host_config.h:

unsupported GNU version! gcc versions later than 14 are not supported!

So you install gcc14 alongside your real compiler and point nvcc at it with -ccbin, or wave the check away with -allow-unsupported-compiler and hope. Getting either of those threaded through a package build’s CMake configuration is its own small adventure.

Clear that, and the next layer surfaces: an old CUDA toolkit’s headers against a current libstdc++ don’t agree on exception specifications any more, and the build dies on mismatched declarations. Patching noexcept(true) onto the offending declarations in CUDA’s own headers gets you moving again.

Add it up and the “working” configuration is a pinned driver, a pinned toolkit, a second compiler installed solely to appease the first one, and hand-edited vendor headers — a stack of four things that each break independently on the next update. That is not a setup, that is a pet.

Vulkan Sidesteps All of It

None of it is necessary, because CUDA is not the only way to get compute out of these cards.

llama.cpp’s Vulkan backend — which ollama can be built against — talks to the GPU through the Vulkan compute API, and its shaders are compiled at runtime by the driver. There is no offline device-code generation step, so there is no nvcc, no architecture flag, no host compiler ceiling and no vendor headers to patch. The 580xx driver you are already pinned to for display purposes ships a perfectly good Vulkan implementation. What got cut for Pascal was the CUDA build toolchain, not the silicon’s ability to run compute shaders.

In practice this means swapping vulkan in for the cuda_v* entries in OLLAMA_LLAMA_BACKENDS when building the package, and that is essentially the whole intervention.

The trade-off is real but narrow. Vulkan and CUDA land in roughly the same place on token generation, which is the number you feel when a model is answering you. Prompt processing is where CUDA still has a clear lead, so long-context work and big prompt ingests are noticeably slower. For interactive use on hardware this old, that is a good trade.

Worth It?

For three cards that were otherwise going to sit in a cupboard, comfortably. The honest framing is that Pascal has entered its long tail: the driver is pinned until 2028, the CUDA path is closed to new builds, and nothing about that is going to improve. Vulkan is the option that doesn’t fight it — it asks nothing of the toolchain, so there is nothing in the setup to break the next time something upstream moves.