<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Llama-Cpp on scriptable.com</title><link>https://scriptable.com/categories/llama-cpp/</link><description>Recent content in Llama-Cpp on scriptable.com</description><generator>Hugo</generator><language>en</language><lastBuildDate>Mon, 10 Aug 2026 16:20:12 -0400</lastBuildDate><atom:link href="https://scriptable.com/categories/llama-cpp/index.xml" rel="self" type="application/rss+xml"/><item><title>Serve a Local OpenAI-Compatible Endpoint with llama.cpp on macOS</title><link>https://scriptable.com/posts/llama-cpp/llama-server-openai-endpoint-macos/</link><pubDate>Mon, 10 Aug 2026 16:20:12 -0400</pubDate><guid>https://scriptable.com/posts/llama-cpp/llama-server-openai-endpoint-macos/</guid><description>One command turns a downloaded GGUF file into an HTTP server that speaks the OpenAI API. Code written against openai.OpenAI runs against it with one line changed — the base_url — and nothing leaves…</description></item><item><title>Getting Started with llama.cpp on macOS</title><link>https://scriptable.com/posts/llama-cpp/getting-started-llama-cpp-macos/</link><pubDate>Fri, 07 Aug 2026 15:31:03 -0400</pubDate><guid>https://scriptable.com/posts/llama-cpp/getting-started-llama-cpp-macos/</guid><description>llama.cpp runs large language models directly on your machine, with no Python runtime, no server process you did not start, and no account. On Apple Silicon it uses Metal for GPU work and the unified…</description></item></channel></rss>