Posts

Test Before You Trust

By Matthew Hunter |  Aug 31, 2026  | embeddings, memstore, go-embedding, claude-code

My workstation is an AMD Strix Halo. It has an NPU, and the NPU can run an embedding model – embed-gemma-300m-FLM, served by Lemonade alongside the same model on the integrated GPU. Same weights, two runtimes. Moving embeddings to the NPU is more an experiment in using available hardware than anything else. I’ve experimented with Whisper transcription on the NPU successfully with (dicta)[]; it was acceptable, slower than the iGPU, slightly worse, but functional enough that I use it and the GPU stays free for more demanding tasks. This seemed like a natural follow up when I saw it in (Lemonade)[]’s list of NPU models.

Continue Reading...

Fast Is Slow, Slow Is Smooth, Smooth Is Fast

By Matthew Hunter |  Aug 30, 2026  | ai, performance, concurrency, go, whisper

My D&D transcription pipeline got about a third faster. I did it by taking away twenty-four of the thirty-two threads it was using, and by not using the fastest GPU in the house.

Neither of those was the plan. The plan was to throw a 5090 at it.

Fast is slow

I’ve written before about how the session reports get made: record the Discord audio, transcribe it with speaker labels, hand the transcript to an LLM, edit the result into narrative prose. That post described a pair of bash scripts wrapping WhisperX. Since then the scripts have been replaced by a single Go binary that does the same job without PyTorch, without a Hugging Face token, and without a Python environment to rot.

Continue Reading...

AAISM: ISACA's AI Security Management Credential

By Matthew Hunter |  Jul 23, 2026  | isaca, ai

The AAISM is ISACA’s Advanced in AI Security Management credential, and new is the word that explains most of what follows. New exam, new training materials, new subject matter that is changing faster than any of them can keep up with. I passed it in July 2026. The credential is worth holding and the training was worth buying – but all of it still has the rough edges of a first edition, and the honest way to review it is as one.

Continue Reading...

The Capture Card That Wouldn't Capture

By Matthew Hunter |  Jun 29, 2026  | linux, kernel, v4l2, claude-code

The card is a clone. Its PCI vendor ID is 0x8888 – not a registered vendor, just four eights, the fingerprint of hardware built to look like something it isn’t. lspci calls it a “Silicon Magic AVMatrix VC12 4K HDMI Capture.” It’s a cheap 4K HDMI capture card in the lineage of a Magewell Pro Capture, and it had sat dead in one of my machines since I bought it, because the only Linux driver I could find for it didn’t work.

Continue Reading...

Hard Lemonade: Three Fixes to Get Local AI Pouring on AMD

By Matthew Hunter |  Jun 23, 2026  | ai, lemonade, olla, amd, rocm, open-source, golang

Running a local LLM server is the easy part. Getting three separate pieces of infrastructure to agree that a model is downloaded, reachable, and worth waiting for is where the afternoon goes. Over the past two weeks I shipped three fixes across two open-source projects to get AMD’s Lemonade serving models behind the Olla proxy on my Strix Halo box. None of them was hard in the algorithmic sense – the diffs are a struct field, a config key, and a prepended path. They all came out of the same goal: point Olla at Lemonade on a Radeon and get a chat completion back.

Continue Reading...

Smoke: Black-Box Route Testing the Router Gates Itself

By Matthew Hunter |  Jun 6, 2026  | go, testing, ci, http, architecture

Every unit test was green. The page returned 502 anyway.

The route was a treasure generator. Its store had a thorough test suite, all passing, because the test built the store the way the test knew to build it – with the database pool wired in. Production built it differently: a copy-paste in the route setup left the pool out, the store carried a nil handle, and the first query dereferenced nil. The handler panicked, the connection dropped, the reverse proxy turned that into a 502. No test caught it, because no test exercised the wired route against a running server. The tests checked the parts. Nothing checked that the assembled thing served.

Continue Reading...

The Layers That Didn't Hold

By Matthew Hunter |  Jun 1, 2026  | ai, security, prompt-injection, architecture, rss

A few weeks ago I wrote that defense in depth for AI agents means layers, not walls: screen untrusted content before the model acts on it, sanitize what comes back out, and never trust the data flowing through. Clean theory. Then I went back and read the code in Herald that was supposed to implement those layers.

Several of them didn’t hold.

Herald is my feed reader. It pulls RSS and Atom from across the internet, runs each article through a local security model before anything else touches it, scores the survivors for relevance, and announces the interesting ones. Every feed item is untrusted content aimed at a model. That’s the whole premise of the defense-in-depth piece, and it’s exactly the threat I built Herald to study. What follows is the v0.2.0 hardening pass – the bugs the theory missed, and a couple of ideas that worked.

Continue Reading...

The LiteLLM Supply Chain Attack: A Homelab Postmortem

By Matthew Hunter |  May 15, 2026  | ai, security, supply-chain, litellm, homelab, postmortem

On March 24, 2026, the LiteLLM PyPI package was compromised. Versions 1.82.7 and 1.82.8, published by an account labeled TeamPCP, contained malicious code. I had LiteLLM running in my homelab as a routing layer between local AI clients and several model backends. This post is the postmortem: what I was running, what the exposure actually was, why I removed LiteLLM rather than just upgrading, and what the incident clarified about supply chain risk in homelab AI infrastructure.

Continue Reading...