FICHA · AUR

python-llama-cpp-cuda

Python bindings for llama.cpp

  • CUDA llama.cpp bindings
  • LIBRARY
  • AI
  • HARDWARE
  • Dependency only
official+codex · reviewed · Jun 3, 2026 description in en

Description

Local llama.cpp inference can use CUDA acceleration through Python bindings on NVIDIA GPUs. AI developers use this variant for faster chat, embeddings, and local model experiments. GPU memory, driver compatibility, model provenance, and generated output need validation.

Permissions

Permissions not analysed for this source yet.