<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Grammar-Constrained Decoding on Matt Suiche</title><link>https://www.msuiche.com/tags/grammar-constrained-decoding/</link><description>Recent content in Grammar-Constrained Decoding on Matt Suiche</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Wed, 09 Sep 2026 00:00:00 +0200</lastBuildDate><atom:link href="https://www.msuiche.com/tags/grammar-constrained-decoding/index.xml" rel="self" type="application/rss+xml"/><item><title>Membership vs Mass: Grammar-Constrained Decoding, Forced Abliteration, and One Honest Negative Result</title><link>https://www.msuiche.com/posts/membership-vs-mass-grammar-constrained-decoding-forced-abliteration-and-one-honest-negative-result/</link><pubDate>Wed, 09 Sep 2026 00:00:00 +0200</pubDate><guid>https://www.msuiche.com/posts/membership-vs-mass-grammar-constrained-decoding-forced-abliteration-and-one-honest-negative-result/</guid><description>&lt;p&gt;&lt;a href="https://tantalus.io/" target="_blank" rel="noopener"&gt;Vince&amp;rsquo;s Tantalus arena&lt;/a&gt; is the first public
deployment of &lt;strong&gt;grammar-constrained decoding&lt;/strong&gt; as a security boundary: the
model sits behind a GBNF grammar, and the grammar decides which token ids
may exist next, per position, for the entire generation. Round 1 asks you
to smuggle an instruction past it. Round 2 inverts the game: you speak
only in tool calls, one character at a time, through the model&amp;rsquo;s own
next-token suggestions. Both rounds are the same object seen from two
sides: a token-id filter applied at the sampler, compiled from a grammar,
never touching the weights.&lt;/p&gt;</description></item></channel></rss>