Red-Team an Open-Weight LLM on Your Mac with NVIDIA garak and Ollama

This tutorial ends with NVIDIA’s garak scanner attacking an open-weight model running on your own Mac, and with you able to answer a question most teams answer by guesswork: did this change to my prompt make the model easier or harder to hijack? You will scan Llama 3.2 3B (3 billion parameters), served locally by Ollama, with hundreds of prompt-injection attacks, read the individual attacks that worked, then write two defensive system prompts and measure each one against the same attacks. The first one makes the model more vulnerable. The second helps against one kind of injection and is indistinguishable from noise against the other. You will finish with a make gate target that fails a build when a prompt change pushes the attack success rate past a threshold you choose. Nothing leaves the machine: no API keys, no hosted model, no account. ...

38 min

Run vLLM on Apple Silicon with the Metal Plugin

vLLM is the inference server many production LLM deployments sit behind, and until recently it had no answer on a Mac beyond a source build of its CPU backend. The vllm-metal plugin changed that: it keeps vLLM’s engine, scheduler and OpenAI-compatible API, and swaps the compute layer for MLX, Apple’s array framework, running on the GPU through Metal. By the end of this article you will have that server running on your Mac, answering on the same OpenAI-compatible endpoint you would deploy on a CUDA box, sized so it leaves the machine usable. You will then measure the number that decides whether vLLM is the right server for your workload at all: how much aggregate throughput continuous batching buys you, and how much per-request latency it costs. Continuous batching is vLLM’s core trick, running many requests through the model together and admitting new ones as others finish. ...

33 min

The Model Is Read-Only: Build a Glass-Box LLM Harness in Python

A language model is a function from text to text. You send it a list of messages, it returns one message, and nothing else happens. It cannot read your disk, check the clock, remember your previous question, or run the tool it just asked for. Every capability an AI product appears to have belongs to the program wrapped around the model. That program is the harness: the code that builds each request, executes the tools the model asks for, and decides what the model is shown in the first place. ...

60 min

Build a User-Scoped MCP Server on macOS

An MCP server that adds two numbers has no opinion about who is calling it. One that posts to a social account has to answer a question before it can do anything at all: whose account? There are two workable answers. You can build a multi-tenant server that authenticates every caller and looks up the account they own, which is the shape Add GitHub OAuth to a FastMCP Server builds. Or you can build one process that acts as exactly one account, reads that account out of its own environment, and runs on the same machine as the client calling it. ...

48 min

Build an AI Agent in Python on a Local Model with llama.cpp

An agent is a program that lets a language model decide which of your functions to run, runs them, and hands back the results until the model says it is done. There is no framework in this article and no API key. By the end you will have a coding agent on your own machine that lists, reads and writes files inside a directory you choose, asks permission before running a shell command, and stops instead of spinning when the model gets stuck. The core agent is a few hundred lines of Python, plus a model-free pytest suite, for a 795 MiB model. ...

40 min

Install Hermes Agent on Two Macs with Ansible

Hermes Agent installs from a shell script in a couple of minutes. Installing it the same way twice, a year apart, on a machine you have since forgotten the details of, is the harder problem — and it is the one Ansible solves. This article builds one role that installs Hermes on two Macs that differ in the ways that actually matter: devbot5 — the Mac you are typing on minime — a headless Mac mini Connection local, no SSH at all ssh Runs as your login account a dedicated service account launchd job LaunchAgent in your home LaunchDaemon in /Library API bound to 127.0.0.1 0.0.0.0 Those last two rows are what this article is about. A headless Mac has no one logged in, so there is no GUI session for a LaunchAgent to live in and the job has to be a LaunchDaemon that starts at boot and drops privileges. A laptop joins hotel and coffee-shop networks, so binding an agent’s API to every interface there would publish a shell to whoever else is on that LAN. ...

44 min

Getting Started with Hermes Agent on macOS, from CLI to Slack

Hermes Agent is Nous Research’s open-source (MIT) agent runtime: one agent with persistent memory that you can reach from a terminal or from a chat platform, backed by whichever model provider you point it at. Unlike an agent library, it ships as an installed program with its own config directory, a messaging gateway, and a command approval layer, so most of the work in getting started is deciding what it is allowed to do rather than writing code. ...

36 min

Round-Robin an MCP Server Behind nginx with Redis-Backed Sessions on macOS

Add Per-Plan Rate Limiting to a FastMCP Server on macOS closes its troubleshooting with “back it with Redis (shared, atomic counters) if you run several instances behind a load balancer”, and Add Observability to a FastMCP Server on macOS adds a /health endpoint “for load balancers”. Neither article puts a load balancer in front of anything. This one does, and the first thing that happens is that the server stops working. ...

38 min

Trigger Synthetic Orders from an MCP Server on macOS

Generate Synthetic JSON Requests to Test an API on macOS built two generators and a CLI that fires batches at an order-intake API. This article puts the same generators behind an MCP server, so an agent can preview a payload, check the target, and trigger a batch by asking for one. Wrapping a generator in tools is the easy half. The half worth attention is the control surface: which decisions the caller gets to make and which ones the server keeps. A model that can pick the destination of a traffic generator is a server-side request forgery primitive with a friendly name, and a model that can pick the batch size can turn one sentence into fifty thousand POSTs. Here the target comes from the environment and the batch size is capped, so the tools stay useful without handing over either decision. ...

32 min

Build MCP Prompts That Trigger Multi-Step Workflows on macOS

Most teams have a procedure that only lives in someone’s head. Cutting release notes, triaging a breaking change, prepping an on-call handoff: five steps, done slightly differently every time, and badly the week the person who knows them is on vacation. MCP gives you three primitives to fix that, and the interesting one is the least used. Tools are called by the model. Resources are read for context. Prompts are chosen by a person: named, parameterized templates a client surfaces as a slash command. That makes a prompt the natural home for a procedure — the user picks it, fills in one argument, and the model runs the same five steps in the same order every time. ...

26 min