Red-Team an Open-Weight LLM on Your Mac with NVIDIA garak and Ollama

This tutorial ends with NVIDIA’s garak scanner attacking an open-weight model running on your own Mac, and with you able to answer a question most teams answer by guesswork: did this change to my prompt make the model easier or harder to hijack? You will scan Llama 3.2 3B (3 billion parameters), served locally by Ollama, with hundreds of prompt-injection attacks, read the individual attacks that worked, then write two defensive system prompts and measure each one against the same attacks. The first one makes the model more vulnerable. The second helps against one kind of injection and is indistinguishable from noise against the other. You will finish with a make gate target that fails a build when a prompt change pushes the attack success rate past a threshold you choose. Nothing leaves the machine: no API keys, no hosted model, no account. ...

38 min