llama.cpp
C/C++ LLM inference engine and GGUF format that runs quantized models efficiently on CPUs, Apple Silicon and GPUs, with an OpenAI-compatible server.
llama.cpp reviews
- No reviews yet.
See something outdated? Suggest an update
C/C++ LLM inference engine and GGUF format that runs quantized models efficiently on CPUs, Apple Silicon and GPUs, with an OpenAI-compatible server.