Bleeding Llama: When AI Model Files Become Memory Leaks
Technical analysis of CVE-2026-7482, a critical unauthenticated heap OOB read vulnerability in Ollama's GGUF processing that leaks process memory through quantization
Hello! My name is Matt Suiche. I work on AI Security at Tolmo, and I also experiment with side projects in AI Safety (Weightless, etc.) and Emulation & Operating System Research (WASM PSX, WASM NanoKrnl, etc.). I recently discussed cyberwar in the age of AI, Iran’s cyber capabilities, and how AI is reshaping hacking on Bloomberg’s Odd Lots and the National Security Lab podcast.
Previously, I founded OnDB Inc., a data infrastructure startup for the agentic economy, and co-founded CloudVolumes (acquired by VMware in 2014) and Comae Technologies (acquired by Magnet Forensics in 2022), where I later served as Head of Detection Engineering. I also founded the cybersecurity community project OPCDE.
My path into technology started in reverse engineering as a teenager, and has since spanned memory forensics, operating systems, virtualization, blockchain, and now AI infrastructure.
Technical analysis of CVE-2026-7482, a critical unauthenticated heap OOB read vulnerability in Ollama's GGUF processing that leaks process memory through quantization
Security, not model capability, is what's blocking Agentic AI in the enterprise, where the real market is. MCP misuse, the Lovable and Mercor breaches, the Vercel/Context incident, Delve's SOC 2 problems: the AI ecosystem is failing at …
Building a complete generative techno engine and a reference-track analyzer in pure NumPy — from raw oscillators to 8 arranged-track genre presets. Covers DSP fundamentals, envelopes, LFOs, filters, sidechain, generative patterns, …